Suno Speech generates a spoken voice and original background music together as one track, and Suno labels it beta. TTAPI exposes it at POST /suno/v1/speech, with an inspiration mode for quick briefs and a custom mode for exact scripts, voice gender, and an optional music bed. Each job is billed at 6 quota · $0.06 and returns a jobId that you read through Fetch V2.
What Suno launched with Speech (beta)
Suno introduced Speech (beta) on October 1, 2026, and describes it as the first audio model that generates voice and music together as one cohesive track. This TTAPI note summarizes the launch and shows how to call the matching Suno Speech API endpoint from your own backend.
In Suno's app, you type an idea, a poem, or your own text, then describe the voice and the musical style. Suno's examples include dramatic readings of friends' messages, epic scores for voice notes, meditations, poems, pep talks, and bedtime stories. After a month of testing with a small group, Suno opened the beta to everyone and warns that accents can drift and dramatic pauses can run long.
How the TTAPI Suno Speech API works
TTAPI exposes Speech as POST /suno/v1/speech, with the same host, TT-API-KEY header, and asynchronous job flow as the other Suno operations. A successful submit returns a jobId; the audio arrives later through Fetch V2 or your hookUrl callback.
The custom flag selects one of two modes. Inspiration mode (custom=false, the default) works from a short description of what to say and how it should sound. Custom mode (custom=true) reads the exact text you send in prompt and adds controls for delivery, voice gender, variety, and background music.
| Field | Mode | Notes |
|---|---|---|
| custom | Both | false for inspiration mode (default); true for custom mode |
| gpt_description_prompt | Inspiration | Description of the speech to generate, up to 3,000 characters |
| prompt | Custom | Required text to read aloud, up to 5,000 characters |
| tags | Custom | Tone, pacing, emotion, and scene, up to 1,000 characters |
| vocal_gender | Custom | Male or Female |
| background_music | Custom | Include a music bed under the voice; defaults to true |
| variety | Custom | off, normal (default), high, extra, or max |
| audio_format | Both | mp3, m4a, or wav for the final audioUrl |
| hookUrl | Both | Optional callback URL for job completion |
Field names and limits follow the current TTAPI Suno Speech reference. Check it before hard-coding limits in validation.
Generate speech from a short brief
Inspiration mode is the quickest way to try Speech. Describe the content, the voice, and the pacing in gpt_description_prompt, and the model writes and performs the read. It suits drafts, placeholder narration, and ideas where the exact wording is not fixed yet.
Because the model chooses the words, review the transcript as well as the sound before you publish anything. When wording matters, such as product names, approved scripts, or anything legal, switch to custom mode.
curl --request POST \
--url 'https://api.ttapi.io/suno/v1/speech' \
--header "TT-API-KEY: $TTAPI_KEY" \
--header 'Content-Type: application/json' \
--data '{
"custom": false,
"gpt_description_prompt": "Read a short evening news brief about a new city library opening, in a calm male voice at a steady pace.",
"audio_format": "mp3",
"hookUrl": "https://example.com/webhooks/suno"
}'Read an exact script with custom mode
Custom mode reads the text you send in prompt. Use tags to describe delivery, such as tone, pace, emotion, and setting, and set vocal_gender when the voice must match a character or brand. background_music defaults to true, which is the main difference from a plain text-to-speech voice; set it to false when you only need the narration.
variety controls how far the result may move from the style in tags. Keep off or normal for repeatable narration, and try high or above only when you want to explore different deliveries of the same script. The example below asks for a short bedtime story over a soft music bed.
curl --request POST \
--url 'https://api.ttapi.io/suno/v1/speech' \
--header "TT-API-KEY: $TTAPI_KEY" \
--header 'Content-Type: application/json' \
--data '{
"custom": true,
"prompt": "The lighthouse keeper climbed the last step and looked out at the quiet sea. Tonight, every ship had found its way home. He turned down the lamp, smiled, and whispered goodnight to the waves.",
"tags": "warm, unhurried bedtime story narration, gentle pauses, soft lullaby piano and light strings underneath",
"vocal_gender": "Female",
"background_music": true,
"variety": "normal",
"audio_format": "mp3",
"hookUrl": "https://example.com/webhooks/suno"
}'- Keep the script within 5,000 characters and the tags within 1,000.
- Write names, numbers, and abbreviations the way they should be spoken.
- Set background_music to false when the voice will be mixed into another soundtrack.
Fetch the result and review the read
Speech jobs run asynchronously, like Suno music jobs. Store the returned jobId with the script and settings, then read the job through the Suno Fetch V2 endpoint or wait for the hookUrl callback. Bound your polling and keep the request visible to the user while the job is pending.
Suno is clear that Speech is still beta, so listen to every result before it ships. Keep the jobId, script, and settings together so a rejected take can be regenerated with one change at a time.
curl --request GET \
--url "https://api.ttapi.io/suno/v2/fetch?jobId=$JOB_ID" \
--header "TT-API-KEY: $TTAPI_KEY"- Names and numbers are pronounced as intended.
- The accent and voice stay consistent across the read.
- Pauses and pacing match the tags.
- The music sits under the voice instead of competing with it.
Suno Speech API pricing on TTAPI
TTAPI bills each Suno Speech job at 6 quota · $0.06, the same listed rate as a Suno music generation. Check the billing record in your account after the first request before you forecast spend.
For budgeting, count every take, not just the approved one. A narration that needs three attempts costs three jobs, so a clear script and specific tags usually save more than any other setting.
| Operation | Price | Unit |
|---|---|---|
| Speech | 6 quota · $0.06 | per job |
| Music generation | 6 quota · $0.06 | per job |
| Sound effects | 1.2 quota · $0.012 | per job |
Prices from the TTAPI pricing page. Confirm current billing in your account before production use.