NativePort
← All use cases

Create voice samples from a written description

Let your agent turn a description of a voice into short audio samples you can listen to.

The problem

Describing a voice in text does not let you hear how it will sound. Your agent needs a voice-generation tool to turn qualities such as warmth, pace and energy into audio you can compare.

How NativePort helps

NativePort connects your agent to Fish Audio to generate voice samples from a written description. Your agent can request a couple of alternatives reading the same short passage, then let you listen and choose the direction you prefer.

For example, it can create two warm, calm narration samples for a welcome message. These are audio samples to evaluate; this step does not save a reusable voice for future recordings.

Let your agent build it

You don’t need to write this code yourself. Copy the prompt and paste it into your agent. It will follow NativePort’s setup guide, help you use an existing account or create one, read this page, and ask what you’d like to do next.

View prompt
Help me use NativePort for this use case.

1. Read https://nativeport.ai/SKILL.md and follow its instructions to learn how to use NativePort. Save or update the skill if your environment supports it.

2. Check whether I already have a working NativePort connection or account. Reuse it if available. Otherwise, help me sign in or create an account by following the skill's setup instructions. Guide me through any steps that require me to use a browser, then help connect my agent. Keep credentials private.

3. Follow the skill to learn how to find the right tools, check their prices and instructions, and use them through NativePort.

4. Read this entire use-case page, including the instructions for agents inside the accordion: https://nativeport.ai/use-cases/create-voice-samples/

Once setup is complete and you have read the page, briefly explain what you can help me do and ask me what I would like to do next. Wait for my answer before implementing the example or running paid tools.
Instructions for the agents - You don't need to read this

Technical implementation

Use Python 3. Install requests with python -m pip install requests and set NATIVEPORT_API_KEY to your NativePort key in your server environment. Keep the key out of browser code.

The NativePort voice-design route uses the model: voice-design-1 header and returns JSON containing base64-encoded candidate audio. This example requests two English candidates. Save the response metadata along with the decoded files; use the returned format metadata or inspect the file signature instead of assuming all audio is MP3.

python
import base64
import json
import os
from pathlib import Path
import requests

response = requests.post("https://api.nativeport.ai/fishaudio/v1/voice-design",
    headers={"Authorization": f"Bearer {os.environ['NATIVEPORT_API_KEY']}",
             "model": "voice-design-1"},
    json={"instruction": "A warm, calm English-speaking narrator with clear diction and a relaxed pace.",
          "reference_text": "Welcome. Take a moment to settle in, and we will guide you through the next steps.",
          "language": "en", "n": 2}, timeout=120)
response.raise_for_status()
candidates = response.json().get("candidates") or []
if not candidates:
    raise RuntimeError("No voice samples returned")
metadata = []
for index, candidate in enumerate(candidates, start=1):
    audio = base64.b64decode(candidate["audio_base64"], validate=True)
    if not audio:
        raise RuntimeError("Empty candidate audio")
    if audio.startswith(b"RIFF") and audio[8:12] == b"WAVE":
        extension = "wav"
    elif audio.startswith(b"OggS"):
        extension = "ogg"
    elif audio.startswith(b"ID3") or (len(audio) > 1 and audio[0] == 255 and audio[1] & 224 == 224):
        extension = "mp3"
    else:
        extension = "bin"  # Inspect the encoding before choosing a player.
    filename = f"voice-sample-{index}.{extension}"
    Path(filename).write_bytes(audio)
    metadata.append({"file": filename, **{key: value for key, value in candidate.items()
                                          if key != "audio_base64"}})
Path("voice-samples.json").write_text(json.dumps(metadata, indent=2), encoding="utf-8")
print(json.dumps(metadata, indent=2))

Listen to the samples before evaluating tone or clarity. The candidate IDs are not persistent voice-model IDs: NativePort does not expose Fish Audio’s model-creation route here. Ask the user which sample or description they prefer before planning a separate narration workflow.

Tool costs

Tool used in the examplePrice per call
Fish Audio — one voice-design request, two candidates $0.01055 / call

The example requests two candidate samples in one successful voice-design call at the price shown above. It does not create a stored voice or include later text-to-speech calls.

Prices in USD. Usage-based tools have no fixed per-call price. View pricing.