# Punctuation-driven voice direction — there is no SSML here

> Six worked before and after lines showing what each mark changes in a read, what punctuation cannot do, and where the speed and expressiveness dials fit.

Published 2026-09-11 · eroq.ai — canonical: https://eroq.ai/blog/punctuation-driven-voice-direction


There is no SSML on eroq. No `<break time="400ms"/>`, no prosody tags, no emphasis markup, no phoneme overrides. The entire direction language is the punctuation you already know, plus two dials. Most people arrive expecting a limitation and leave treating it as a discipline, because a script that performs when a person reads it is a script that performs when the engine does — and the reverse is never true of a script held together by tags.

What follows is the six marks that change a read, each one shown twice.

## Six marks, six jobs

### The ellipsis holds a pause

**Before** — Come closer, I wasn't finished with you.

**After** — Come closer… I wasn't finished with you.

The comma takes a breath. The ellipsis hangs. It also slows the words on either side of it, which is why one per paragraph is the limit — a voice that trails off in every sentence sounds unsure of all of them.

### The em dash cuts a beat

**Before** — I was going to wait, but why would I?

**After** — I was going to wait — but why would I?

A comma joins two clauses. A dash interrupts one. Use it for a change of mind mid-sentence, a self-correction, a thought that arrives late and shoves the previous one aside.

### Question marks lift the line

**Before** — You really thought I forgot about you.

**After** — You really thought I forgot? About you?

One question rises at the end. Two stacked questions rise twice, and the fragment gets the second lift all to itself. Splitting a sentence into a question plus a fragment is the cheapest way to add an attitude that no dial provides.

### Exclamation marks spike the energy

**Before** — Stop right there. Perfect, stay just like that.

**After** — Stop right there! Perfect — stay just like that.

Volume and pace go up a notch. The reason to ration them is arithmetic rather than taste — if three sentences in a row shout, none of them does, and the line reads as a tone rather than a reaction.

### Sentence length controls the pace

**Before** — Slow down and breathe, and then we can start properly.

**After** — Slow down. Breathe. Good — now we can start.

Short sentences read punchy because each period is a full stop the voice has to land on. Long comma-led sentences flow. This is the single biggest lever in the list, and it is the one writers reach for last because it means rewriting rather than adding a mark.

### Commas around a word make it land

**Before** — And that is exactly the problem.

**After** — And that, right there, is exactly the problem.

Setting a phrase apart with commas makes the read lean on it. It is the closest thing to per-word emphasis available, and unlike a tag it survives being read aloud by a human, which means you can test it without spending a credit.

All six live in the Delivery guide inside the [Voice studio](/studio/voice), where tapping one drops its example into the editor so you can hear the difference rather than take my word for it.

## Line breaks and paragraphs

A period stops a sentence; a paragraph break separates thoughts, and the read treats it as the larger of the two gaps. Break where a reader would take an actual breath — between the setup and the punchline — rather than where the page looks tidy. A wall of six sentences renders as a wall.

## What punctuation cannot do

Worth knowing before you spend an afternoon trying:

- **There is no timed pause.** You cannot ask for 400 milliseconds. If a gap has to be exact — a beat cut to picture, a gap for a sound effect — render the two halves separately and place them in your editor.
- **There is no emphasis tag.** Commas around a phrase, or rewriting so the important word lands at the end of a short sentence, are the tools.
- **Stage directions get spoken.** Everything you type is text to be read. `[angrily]` at the head of a line does not set a mood, it makes the voice say "angrily". Same for `(laughs)` and `*whispers*`.
- **There is no pronunciation override.** An acronym or an unusual name that comes out wrong gets fixed by spelling it the way it sounds — "Ess Cue Ell" or "Kess-ler" — not by a phoneme tag.

## Then the two dials

Punctuation shapes the performance; the dials set the register it happens in.

**Speed** runs 0.70× to 1.30× in steps of 0.05, with 1 as natural. **Expressiveness** runs from 0, a steady close read, to 100%, theatrical, and it governs how hard the engine leans into the punctuation you wrote. Both work identically on roster voices and on your own clones, and both are real engine parameters rather than post-processing.

Three combinations that hold up:

- **Audiobook narration** — speed 0.95, expressiveness 30%. Even, unhurried, nothing editorializing on the text.
- **Character dialogue** — speed 1.00, expressiveness 70%. The marks do the acting; the dial permits it.
- **Ad read** — speed 1.10, expressiveness 55%. Forward-leaning without becoming a caricature.

Adjust one at a time, and if a cloned voice stops sounding like its owner, bring expressiveness down before you touch speed. The sample is the anchor and both dials pull away from it.

## A paragraph, before and after

**Before**

> I can't believe you actually came back after everything that happened last time, and honestly I don't know what to say to you, so maybe you should sit down and tell me what you saw before someone notices the door is open.

**After**

> You came back. I didn't think you would… not after last time.
>
> Sit down — no, really, sit. We have maybe ten minutes before they notice the door.
>
> So. Tell me what you saw.

Same information, four seconds longer, and it reads as a person rather than a paragraph. Nothing was tagged. The whole change is five periods, an ellipsis, a dash and a paragraph break.

## Iterate on the cheap model

Rewriting punctuation is a loop, and loops should be cheap. Draft on [Voice Turbo](/models/eroq-voice-turbo) at 2 credits per 100 characters, ship on [Voice One](/models/eroq-voice-one) at 3 — same roster, same clones, same MP3 out. A 200-character line is 6 credits on Voice One and 4 on Turbo, so a dozen passes at a stubborn sentence costs less than a coffee. [Voice One vs Voice Turbo](/blog/voice-one-vs-voice-turbo) covers where the extra credit shows up.

Over the API the same two dials are `speed` and `expressiveness` on `POST /v1/audio/speech`, documented in the [speech reference](/docs/speech), and the input text is the direction exactly as it is in the studio. For laying the result under a picture, [AI voice-over for video](/blog/ai-voice-over-for-video) picks up where this leaves off.

## FAQ

### Does eroq support SSML or pause tags?

No. There is no SSML, no break tag and no prosody markup — punctuation and sentence structure are the direction language, and the speed and expressiveness controls set the register. Anything you type inside the text is spoken, including brackets.

### How do I get a pause of an exact length?

You cannot ask for one. An ellipsis produces a hanging pause and a paragraph break a larger gap, but neither is measured. When a gap has to hit a specific frame, render the two halves as separate lines and place them on a timeline in your own editor.

### Why does my line sound flat no matter what I set?

Almost always because the writing is flat. Long comma-led sentences with no full stops give the engine nothing to land on, and raising expressiveness on a flat script produces an emphatic monotone. Break it into short sentences first, then re-render.

### Do these marks work the same on a cloned voice?

Yes. Punctuation, speed and expressiveness behave identically on roster voices and on your clones, since the clone supplies the timbre and manner while the script supplies the performance. If a clone stops sounding like itself at high expressiveness, lower the dial rather than the writing.

Want to hear the six lines rather than read them? [Open the Voice studio](/studio/voice), hit the Delivery guide, and tap each one into the editor.
