SpeechNewsSuno Speech API launch

Suno Speech API: voice and music in one track

Suno Speech (beta) generates spoken audio and background music as one track. Learn the TTAPI Suno Speech API modes, parameters, pricing, and cURL requests.

TTAPI5 min read
IN THIS NOTE

What you will take away

  • Understand what Suno Speech (beta) generates
  • Choose inspiration or custom mode for a request
  • Submit, fetch, and budget a Suno Speech API job
THE TAKEAWAY

Suno Speech generates a spoken voice and original background music together as one track, and Suno labels it beta. TTAPI exposes it at POST /suno/v1/speech, with an inspiration mode for quick briefs and a custom mode for exact scripts, voice gender, and an optional music bed. Each job is billed at 6 quota · $0.06 and returns a jobId that you read through Fetch V2.

What Suno launched with Speech (beta)

Suno introduced Speech (beta) on October 1, 2026, and describes it as the first audio model that generates voice and music together as one cohesive track. This TTAPI note summarizes the launch and shows how to call the matching Suno Speech API endpoint from your own backend.

In Suno's app, you type an idea, a poem, or your own text, then describe the voice and the musical style. Suno's examples include dramatic readings of friends' messages, epic scores for voice notes, meditations, poems, pep talks, and bedtime stories. After a month of testing with a small group, Suno opened the beta to everyone and warns that accents can drift and dramatic pauses can run long.

How the TTAPI Suno Speech API works

TTAPI exposes Speech as POST /suno/v1/speech, with the same host, TT-API-KEY header, and asynchronous job flow as the other Suno operations. A successful submit returns a jobId; the audio arrives later through Fetch V2 or your hookUrl callback.

The custom flag selects one of two modes. Inspiration mode (custom=false, the default) works from a short description of what to say and how it should sound. Custom mode (custom=true) reads the exact text you send in prompt and adds controls for delivery, voice gender, variety, and background music.

Suno Speech API request fields
FieldModeNotes
customBothfalse for inspiration mode (default); true for custom mode
gpt_description_promptInspirationDescription of the speech to generate, up to 3,000 characters
promptCustomRequired text to read aloud, up to 5,000 characters
tagsCustomTone, pacing, emotion, and scene, up to 1,000 characters
vocal_genderCustomMale or Female
background_musicCustomInclude a music bed under the voice; defaults to true
varietyCustomoff, normal (default), high, extra, or max
audio_formatBothmp3, m4a, or wav for the final audioUrl
hookUrlBothOptional callback URL for job completion

Field names and limits follow the current TTAPI Suno Speech reference. Check it before hard-coding limits in validation.

Generate speech from a short brief

Inspiration mode is the quickest way to try Speech. Describe the content, the voice, and the pacing in gpt_description_prompt, and the model writes and performs the read. It suits drafts, placeholder narration, and ideas where the exact wording is not fixed yet.

Because the model chooses the words, review the transcript as well as the sound before you publish anything. When wording matters, such as product names, approved scripts, or anything legal, switch to custom mode.

Inspiration-mode Speech request
cURL
curl --request POST \
  --url 'https://api.ttapi.io/suno/v1/speech' \
  --header "TT-API-KEY: $TTAPI_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
  "custom": false,
  "gpt_description_prompt": "Read a short evening news brief about a new city library opening, in a calm male voice at a steady pace.",
  "audio_format": "mp3",
  "hookUrl": "https://example.com/webhooks/suno"
}'

Read an exact script with custom mode

Custom mode reads the text you send in prompt. Use tags to describe delivery, such as tone, pace, emotion, and setting, and set vocal_gender when the voice must match a character or brand. background_music defaults to true, which is the main difference from a plain text-to-speech voice; set it to false when you only need the narration.

variety controls how far the result may move from the style in tags. Keep off or normal for repeatable narration, and try high or above only when you want to explore different deliveries of the same script. The example below asks for a short bedtime story over a soft music bed.

Custom-mode Speech request
cURL
curl --request POST \
  --url 'https://api.ttapi.io/suno/v1/speech' \
  --header "TT-API-KEY: $TTAPI_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
  "custom": true,
  "prompt": "The lighthouse keeper climbed the last step and looked out at the quiet sea. Tonight, every ship had found its way home. He turned down the lamp, smiled, and whispered goodnight to the waves.",
  "tags": "warm, unhurried bedtime story narration, gentle pauses, soft lullaby piano and light strings underneath",
  "vocal_gender": "Female",
  "background_music": true,
  "variety": "normal",
  "audio_format": "mp3",
  "hookUrl": "https://example.com/webhooks/suno"
}'
  • Keep the script within 5,000 characters and the tags within 1,000.
  • Write names, numbers, and abbreviations the way they should be spoken.
  • Set background_music to false when the voice will be mixed into another soundtrack.

Fetch the result and review the read

Speech jobs run asynchronously, like Suno music jobs. Store the returned jobId with the script and settings, then read the job through the Suno Fetch V2 endpoint or wait for the hookUrl callback. Bound your polling and keep the request visible to the user while the job is pending.

Suno is clear that Speech is still beta, so listen to every result before it ships. Keep the jobId, script, and settings together so a rejected take can be regenerated with one change at a time.

Check a Speech job
cURL
curl --request GET \
  --url "https://api.ttapi.io/suno/v2/fetch?jobId=$JOB_ID" \
  --header "TT-API-KEY: $TTAPI_KEY"
  • Names and numbers are pronounced as intended.
  • The accent and voice stay consistent across the read.
  • Pauses and pacing match the tags.
  • The music sits under the voice instead of competing with it.

Suno Speech API pricing on TTAPI

TTAPI bills each Suno Speech job at 6 quota · $0.06, the same listed rate as a Suno music generation. Check the billing record in your account after the first request before you forecast spend.

For budgeting, count every take, not just the approved one. A narration that needs three attempts costs three jobs, so a clear script and specific tags usually save more than any other setting.

TTAPI Suno prices for related audio operations
OperationPriceUnit
Speech6 quota · $0.06per job
Music generation6 quota · $0.06per job
Sound effects1.2 quota · $0.012per job

Prices from the TTAPI pricing page. Confirm current billing in your account before production use.

KEEP BUILDING

Open the Suno Music API.

Continue from the model page, then follow the current music API documentation.