The problem
A larger collection job may not finish while your agent is waiting for a single reply. If it starts again instead of checking the original job, it can repeat work and create unnecessary costs.
How NativePort helps
NativePort connects your agent to Apify to start a collection job and return a reference it can use later. Your agent can check whether that job has finished and download the results when they are ready.
For example, it can collect text from up to twenty pages of a website, save the job reference and continue from the same job on its next visit. The limit makes the first trial manageable before collecting a larger dataset.
Let your agent build it
You don’t need to write this code yourself. Copy the prompt and paste it into your agent. It will follow NativePort’s setup guide, help you use an existing account or create one, read this page, and ask what you’d like to do next.
View prompt
Help me use NativePort for this use case. 1. Read https://nativeport.ai/SKILL.md and follow its instructions to learn how to use NativePort. Save or update the skill if your environment supports it. 2. Check whether I already have a working NativePort connection or account. Reuse it if available. Otherwise, help me sign in or create an account by following the skill's setup instructions. Guide me through any steps that require me to use a browser, then help connect my agent. Keep credentials private. 3. Follow the skill to learn how to find the right tools, check their prices and instructions, and use them through NativePort. 4. Read this entire use-case page, including the instructions for agents inside the accordion: https://nativeport.ai/use-cases/collect-datasets-asynchronously/ Once setup is complete and you have read the page, briefly explain what you can help me do and ask me what I would like to do next. Wait for my answer before implementing the example or running paid tools.
Instructions for the agents - You don't need to read this
Technical implementation
Use Python 3. Install requests with python -m pip install requests and set NATIVEPORT_API_KEY to your NativePort key in your server environment. Keep the key out of browser code.
Set START_URL to the public website to crawl. Save this as collect.py. Run python collect.py start once, then python collect.py check on later invocations. Keep collection-run.json in the same working directory. The Actor input follows Website Content Crawler.
import json
import os
from pathlib import Path
import sys
import requests
BASE = "https://api.nativeport.ai"
HEADERS = {"Authorization": f"Bearer {os.environ['NATIVEPORT_API_KEY']}"}
state = Path("collection-run.json")
def call(method, path, **kwargs):
response = requests.request(method, BASE + path, headers=HEADERS, timeout=90, **kwargs)
response.raise_for_status()
return response.json()
if len(sys.argv) != 2 or sys.argv[1] not in {"start", "check"}:
raise SystemExit("Usage: python collect.py start|check")
if sys.argv[1] == "start":
if state.exists():
raise SystemExit("Saved run exists. Use check; do not start a duplicate job.")
run = call("POST", "/apify/v2/acts/apify~website-content-crawler/runs",
json={"startUrls": [{"url": os.environ["START_URL"]}],
"maxCrawlPages": 20})["data"]
state.write_text(json.dumps({"id": run["id"]}), encoding="utf-8")
print("Run saved:", run["id"])
else:
run_id = json.loads(state.read_text(encoding="utf-8"))["id"]
run = call("GET", f"/apify/v2/actor-runs/{run_id}")["data"]
if run["status"] in {"FAILED", "TIMED-OUT", "ABORTED"}:
raise SystemExit(f"Run ended with {run['status']}; inspect it before retrying")
if run["status"] != "SUCCEEDED":
raise SystemExit(f"Still {run['status']}; run check again later")
dataset = run["defaultDatasetId"]
output = Path("collected-pages.jsonl")
temporary = output.with_suffix(".tmp")
offset = 0
with temporary.open("w", encoding="utf-8") as handle:
while True:
items = call("GET", f"/apify/v2/datasets/{dataset}/items",
params={"limit": 100, "offset": offset})
if not isinstance(items, list):
raise RuntimeError("Expected a dataset page")
for item in items:
handle.write(json.dumps(item, ensure_ascii=False) + "\n")
offset += len(items)
if len(items) < 100:
break
temporary.replace(output)
print(f"Saved {offset} records to {output}")Run this serially: the local state file is not a multi-worker lock. If the start request times out, do not automatically retry; a job may already have started. Reconcile the run before creating another. A successful run can still contain fewer pages than requested or records with errors. Preserve source URLs and inspect the output before summarizing it. See the NativePort Apify lifecycle.
Tool costs
| Tool used in the example | Price per call |
|---|---|
| Apify — one bounded website-content collection run | Priced from Apify's reported cost for the request or run (usage-based) |
The example starts one run limited to twenty pages. Its price depends on the Actor’s execution and result charges, settled when the run finishes. Status checks and dataset reads are free. A failed run can still have consumed paid work.
Prices in USD. Usage-based tools have no fixed per-call price. View pricing.