Give your app a voice, or turn recordings into text
Create spoken audio from a script, transcribe a recording, or work with a reference voice. Explore voice and audio APIs for the feature you want to build.
Choose what you need from the audio
For narration, start with text to speech: send a script, choose a voice, and receive an audio file. Voice-cloning features can use a reference sample when you have permission to use that voice.
For recordings, choose speech to text. A transcript can make a meeting or interview searchable. Speaker labels help separate the people talking, and word timestamps let your app link the text back to the recording.
If you only need an answer about a recording, an audio-capable model may handle the question directly. Check that the model and route support audio and are available to your account before building on them.
Compare prices using your own workload. Fish Audio transcription is billed by audio duration, while its speech generation is billed by text size. ElevenLabs is billed through NativePort at a flat approximation per call. The provider pages explain the rates.
Explore the guides
Compare providers
Compare starting prices and the tasks each provider supports. Features and usage affect the total cost. For measured results, see the benchmarks and how we test.
| Provider | Best suited for | Starting price |
|---|---|---|
| Anthropic | Work with transcript text using Claude. | metered by Anthropic |
| Apify | Use available Actors for media and audio workflows. | metered by Apify |
| ElevenLabs | Generate speech or transcribe audio with speaker labels. | $0.025 per call |
| Fish Audio | Generate speech, transcribe audio, or design a voice. | metered by Fish Audio |
| Grok | Summarize or analyze transcript text. | metered by Grok (xAI) |
| OpenAI | Use supported models for audio and transcript workflows. | metered by OpenAI |
Use it with NativePort
Use one NativePort key for speech generation, transcription, and supported voice-cloning features. You can also use the same balance for a model that summarizes or works with the transcript.
Before you start
Which transcription API should I try?
The transcription guide uses ElevenLabs for speaker labels and word timestamps. Fish Audio also provides speech recognition. Test a representative recording and compare the text, features, and cost you need.
Does a longer recording cost more?
With Fish Audio’s per-second transcription rate, yes. ElevenLabs uses a flat approximation per call through NativePort. Account for any file limits and splitting in your workflow when comparing the total cost.
Can I clone a voice?
Supported providers offer voice cloning. Use a voice you have permission to use and check the provider’s consent requirements before submitting a sample.