The problem
A long recording can exceed what a single request can handle reliably. If your agent loses progress and sends the entire file again, it can repeat transcription work and make the result harder to manage.
How NativePort helps
NativePort connects your agent to Fish Audio to transcribe smaller parts of a recording. Your agent saves each completed part and combines them in order, so a later run can continue from the work already saved.
For example, it can process a one-hour interview in thirty-second parts. The agent should review words near each cut, where a sentence may continue into the next part.
Let your agent build it
You don’t need to write this code yourself. Copy the prompt and paste it into your agent. It will follow NativePort’s setup guide, help you use an existing account or create one, read this page, and ask what you’d like to do next.
View prompt
Help me use NativePort for this use case. 1. Read https://nativeport.ai/SKILL.md and follow its instructions to learn how to use NativePort. Save or update the skill if your environment supports it. 2. Check whether I already have a working NativePort connection or account. Reuse it if available. Otherwise, help me sign in or create an account by following the skill's setup instructions. Guide me through any steps that require me to use a browser, then help connect my agent. Keep credentials private. 3. Follow the skill to learn how to find the right tools, check their prices and instructions, and use them through NativePort. 4. Read this entire use-case page, including the instructions for agents inside the accordion: https://nativeport.ai/use-cases/transcribe-long-recordings/ Once setup is complete and you have read the page, briefly explain what you can help me do and ask me what I would like to do next. Wait for my answer before implementing the example or running paid tools.
Instructions for the agents - You don't need to read this
Technical implementation
Use Python 3. Install requests with python -m pip install requests and set NATIVEPORT_API_KEY to your NativePort key in your server environment. Keep the key out of browser code.
Prepare a mono, 16 kHz, 16-bit PCM WAV file and set AUDIO_FILE to it. Run in a private working directory. This bounded example uses Python’s wave module and an application-managed checkpoint per successful request. It is not a claim that the upstream accepts an hour-long file in one call.
import base64
import hashlib
import io
import json
import os
from pathlib import Path
import wave
import requests
source = Path(os.environ["AUDIO_FILE"])
headers = {"Authorization": f"Bearer {os.environ['NATIVEPORT_API_KEY']}"}
digest = hashlib.sha256(source.read_bytes()).hexdigest()
folder = Path("transcript-parts") / digest
folder.mkdir(parents=True, exist_ok=True)
assembled = []
with wave.open(str(source), "rb") as recording:
if (recording.getnchannels(), recording.getsampwidth(), recording.getframerate()) != (1, 2, 16000):
raise ValueError("Convert to mono 16-bit PCM WAV at 16 kHz before running")
if not 0 < recording.getnframes() <= 16000 * 3600:
raise ValueError("This example accepts at most sixty minutes")
index = 0
while True:
frames = recording.readframes(16000 * 30)
if not frames:
break
checkpoint = folder / f"{index:04d}.json"
if checkpoint.exists():
result = json.loads(checkpoint.read_text(encoding="utf-8"))
else:
buffer = io.BytesIO()
with wave.open(buffer, "wb") as part:
part.setnchannels(1)
part.setsampwidth(2)
part.setframerate(16000)
part.writeframes(frames)
response = requests.post("https://api.nativeport.ai/fishaudio/v1/asr",
headers=headers, json={"audio": base64.b64encode(buffer.getvalue()).decode("ascii"),
"ignore_timestamps": True}, timeout=120)
response.raise_for_status()
result = response.json()
if not isinstance(result.get("text"), str):
raise RuntimeError("Invalid transcript; progress was not saved")
temporary = checkpoint.with_suffix(".tmp")
temporary.write_text(json.dumps(result, ensure_ascii=False), encoding="utf-8")
temporary.replace(checkpoint)
assembled.append({"part": index, "start_seconds": index * 30,
"text": result["text"], "reported_duration": result.get("duration")})
index += 1
output = folder / "combined.json"
output.write_text(json.dumps(assembled, indent=2, ensure_ascii=False), encoding="utf-8")
print(f"Saved {len(assembled)} ordered parts to {output}")The start times identify chunk boundaries, not word timestamps. The example does not overlap parts or preserve speaker identities across requests. Review boundaries or use silence-aware segmentation when necessary. Do not run concurrent writers against the same checkpoint folder. A timeout can occur after a paid request was processed; reconcile uncertain calls before rerunning. See duration-based ASR billing.
Tool costs
| Tool used in the example | Price per call |
|---|---|
| Fish Audio — one ASR call per thirty-second audio part | $0.0001055 / second (usage-based) |
The example handles at most sixty minutes, split into parts of up to thirty seconds: up to 120 calls. Each call is billed on its reported duration rounded up to a second. Successfully saved parts are skipped on a rerun. Local splitting and assembly add no provider call.
Prices in USD. Usage-based tools have no fixed per-call price. View pricing.