eroq

API reference · Media

Text-to-speech

Synthesize speech from text. Returns MP3 audio.

POST/v1/audio/speech3 per 100 characters (Voice One) · 2 (Turbo)

The response body is the MP3 itself (audio/mpeg) — pipe it straight to a file or an audio element. Input is capped at 5,000 characters per request.

Voices: aria (female), orion (male), lyra (female), silas (male), vesper (female), elio (male), cleo (female), soren (male), rhea (female), cassius (male), sloane (female), brody (male), imogen (female), alistair (male), ines (female), mateo (male), margaux (female), bastien (male), greta (female), henrik (male), hina (female), kaito (male), soyeon (female), minho (male). Credits are charged per started block of 100 input characters — a 250-character line costs 9 credits on Voice One, 6 on Turbo.

Delivery tags (Voice One). Direct the read in square brackets — [whispering] Come closer., [excited], [angry], [laughing], [sighing], [pause], or any short description ([warm and teasing]). The voice performs them and never says them: a delivery tag holds until the next tag or line break, a sound happens once, a pause lands as a beat (never a measured gap). A tag's characters count toward the 100-character blocks like any other text. eroq-voice-turbo reads the words and skips the tags.

Request body

inputstringrequired

The text to speak. On Voice One, [bracket] tags direct the delivery — see the notes.

voicestringrequired

Roster (aria, orion, lyra, silas, vesper, elio, cleo, soren, rhea, cassius, sloane, brody, imogen, alistair, ines, mateo, margaux, bastien, greta, henrik, hina, kaito, soyeon, minho), one of YOUR cloned voice ids, a PUBLIC voice id from GET /v1/voices/public, or char:<id> — a character speaks with their assigned voice.

speednumber

Playback speed 0.7–1.3, 1 = natural.

expressivenessnumber

0 (steady) → 1 (theatrical), mapped onto the engine sampling.

folder_idstring

Library folder the voice line lands in — one of yours (GET /v1/folders). Omitted = the Library root.

flow_idstring

A flow to run on the voice line once it lands — one that starts « After a render » of voice lines (GET /v1/flows). Checked before anything is billed (403 plan_required below the plan flows take). See Flows.

modelstring

eroq-voice-one (default, most expressive) or eroq-voice-turbo (faster, 2 credits per 100 characters).

curl https://eroq.ai/v1/audio/speech \
  -H "Authorization: Bearer $EROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "eroq-voice-one",
  "input": "That noise is the reactor missing its coolant flush. [sighing] Again.",
  "voice": "aria"
}' \
  --output speech.mp3
/v1/audio/speech
→ POST /v1/audio/speech
Press run — this replays a real exchange from the docs' own data. No key, no request, no charge.

Response

audio/mpeg
(binary MP3 body — audio/mpeg)