API 레퍼런스 · Media
Text-to-speech
Synthesize speech from text. Returns MP3 audio.
The response body is the MP3 itself (audio/mpeg) — pipe it straight to a file or an audio element. Input is capped at 5,000 characters per request.
Voices: aria (female), orion (male). Credits are charged per started block of 100 input characters — a 250-character line costs 9 credits on Voice One, 6 on Turbo.
Delivery tags (Voice One). Direct the read in square brackets — [whispering] Come closer., [excited], [angry], [laughing], [sighing], [pause], or any short description ([warm and teasing]). The voice performs them and never says them: a delivery tag holds until the next tag or line break, a sound happens once, a pause lands as a beat (never a measured gap). A tag's characters count toward the 100-character blocks like any other text. eroq-voice-turbo reads the words and skips the tags.
요청 본문
The text to speak. On Voice One, [bracket] tags direct the delivery — see the notes.
Roster (aria, orion), one of YOUR cloned voice ids, a PUBLIC voice id from GET /v1/voices/public, or char:<id> — a character speaks with their assigned voice.
Playback speed 0.7–1.3, 1 = natural.
0 (steady) → 1 (theatrical), mapped onto the engine sampling.
Library folder the voice line lands in — one of yours (GET /v1/folders). Omitted = the Library root.
eroq-voice-one (default, most expressive) or eroq-voice-turbo (faster, 2 credits per 100 characters).
curl https://eroq.ai/v1/audio/speech \
-H "Authorization: Bearer $EROQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "eroq-voice-one",
"input": "That noise is the reactor missing its coolant flush. [sighing] Again.",
"voice": "aria"
}' \
--output speech.mp3응답
(binary MP3 body — audio/mpeg)