The problem
A conversation can switch languages while it is being recorded. Forcing a single language can make it harder for your agent to recover the words accurately, and a text-only agent cannot listen to the recording by itself.
How NativePort helps
NativePort connects your agent to Fish Audio to request a transcript without a fixed language hint. Your agent can inspect the returned text and compare a short sample with the recording before processing more of it.
For example, it can try a short conversation that mixes English and Portuguese. Accuracy at language changes, names and specialist terms needs checking; automatic detection does not guarantee every switch is transcribed correctly.
Let your agent build it
You don’t need to write this code yourself. Copy the prompt and paste it into your agent. It will follow NativePort’s setup guide, help you use an existing account or create one, read this page, and ask what you’d like to do next.
View prompt
Help me use NativePort for this use case. 1. Read https://nativeport.ai/SKILL.md and follow its instructions to learn how to use NativePort. Save or update the skill if your environment supports it. 2. Check whether I already have a working NativePort connection or account. Reuse it if available. Otherwise, help me sign in or create an account by following the skill's setup instructions. Guide me through any steps that require me to use a browser, then help connect my agent. Keep credentials private. 3. Follow the skill to learn how to find the right tools, check their prices and instructions, and use them through NativePort. 4. Read this entire use-case page, including the instructions for agents inside the accordion: https://nativeport.ai/use-cases/transcribe-mixed-language-audio/ Once setup is complete and you have read the page, briefly explain what you can help me do and ask me what I would like to do next. Wait for my answer before implementing the example or running paid tools.
Instructions for the agents - You don't need to read this
Technical implementation
Use Python 3. Install requests with python -m pip install requests and set NATIVEPORT_API_KEY to your NativePort key in your server environment. Keep the key out of browser code.
Set AUDIO_FILE to a short WAV or MP3 sample. The request intentionally omits language and the optional model header, using the default ASR behavior on the NativePort Fish Audio route. The five-megabyte limit below is a conservative example guard, not a published provider maximum.
import base64
import json
import os
from pathlib import Path
import requests
audio = Path(os.environ["AUDIO_FILE"]).read_bytes()
if not audio or len(audio) > 5_000_000:
raise ValueError("Use a non-empty sample smaller than 5 MB for this first test")
response = requests.post("https://api.nativeport.ai/fishaudio/v1/asr",
headers={"Authorization": f"Bearer {os.environ['NATIVEPORT_API_KEY']}"},
json={"audio": base64.b64encode(audio).decode("ascii"),
"ignore_timestamps": True}, timeout=120)
response.raise_for_status()
result = response.json()
if not isinstance(result.get("text"), str) or not result["text"].strip():
raise RuntimeError("No transcript returned")
Path("transcript.json").write_text(json.dumps(result, indent=2, ensure_ascii=False), encoding="utf-8")
print(result["text"])
print("Reported duration:", result.get("duration"))Review the transcript against speech around each language change. Preserve uncertainty instead of silently translating or correcting unfamiliar words. This example does not assign speaker identities, translate the conversation or claim a quality benchmark for mixed-language audio.
Tool costs
| Tool used in the example | Price per call |
|---|---|
| Fish Audio — one transcription request, billed by audio duration | $0.0001055 / second (usage-based) |
The example sends one short recording. The charge uses the duration reported by Fish Audio, rounded up to a whole second, at the displayed rate. Automatic language handling does not make the call free. No translation or summary call is included.
Prices in USD. Usage-based tools have no fixed per-call price. View pricing.