# AI voice for audiobooks and podcasts — the real credit math

> Narrate a chapter on eroq with Voice One against Voice Turbo per hundred characters, voice cloning, punctuation as direction and thirteen languages.

Published 2026-09-11 · eroq.ai — canonical: https://eroq.ai/blog/ai-voice-for-audiobooks-and-podcasts


Long-form audio is a budgeting problem before it is a creative one. A chapter is not a sentence you regenerate until it sounds right; it is sixty thousand characters that you will render, listen to, fix and render again. So the first thing worth knowing about narration on eroq is the unit of billing — characters, not words, not minutes — and the second is how to spend fewer of them on drafts. The short version is on the [audiobooks and podcasts use case](/use-cases/audiobooks-podcasts).

## Two engines, the same voices

[Voice](/studio/voice) has two text-to-speech models and they share the roster and your clones:

- **[Voice One](/models/eroq-voice-one)** — the expressive flagship, 3 credits per 100 characters, MP3 out.
- **Voice Turbo** — faster, 2 credits per 100 characters, same voices, same clones.

The curated roster is small on purpose: **Aria** is warm and intimate, **Orion** is low and calm, and the roster covers 13 languages. Speed and expressiveness are sliders, and a Character you have created can speak with a roster voice or a clone of your own.

## The arithmetic for a 10,000-word chapter

English prose runs about six characters per word including spaces, so a 10,000-word chapter is roughly 60,000 characters, which is 600 blocks of 100.

- **Voice One** — 600 × 3 = **1,800 credits**
- **Voice Turbo** — 600 × 2 = **1,200 credits**

The gap is 600 credits per chapter, or about $6 at the [entry pack rate](/pricing) of $10 for 1,000 credits, and it is the reason for the only workflow rule that matters here: **proof on Turbo, ship on One**. Render the draft chapter on Voice Turbo, listen for the mispronounced surname and the sentence that reads as a question, fix the text, then render the final once on Voice One.

A twelve-chapter book proofed that way is 12 × 1,200 = 14,400 credits of drafts plus 12 × 1,800 = 21,600 credits of finals, which is **36,000 credits** — exactly the monthly allowance on the Team plan ($249 for 36,000 credits). If you only ever render finals, the same book is 21,600 credits — Studio ($149 for 20,000 credits) plus two Starter packs ($10 for 1,000 credits each), or 20,000 ÷ 1,800 = 11.1 chapters a month on Studio alone. Credits roll over forever and never expire, and failed generations refund themselves, so a book written over four months does not need four months of the same plan.

## Punctuation is the direction

There is no separate direction track. Delivery comes from punctuation, and once you know that, you write differently: a comma is a breath, an em dash is a pause with intent, an ellipsis is hesitation, and a full stop is a floor. Short sentences, one idea each, read aloud better than the elegant subordinate clauses that look good on a page.

This is a passage written for the microphone rather than the eye — quotable, and a fair test of any voice you are auditioning:

> He had rehearsed the sentence all afternoon. In the car. In the lift. At the door, with his hand flat against the wood — and then she opened it before he knocked, and every word he had practiced went somewhere else. "You're early," she said. He was not early. He was three years late.

Render that on both models before you commit to a book. Sixty seconds of listening settles an argument that a spec sheet cannot.

## Cloning your own voice

If the channel is yours, the voice should be too. [Voice cloning](/tools/ai-voice-cloning) takes your own recorded audio and gives you a voice you can use anywhere the roster works, on either model, at the same price per 100 characters. For a podcast that means a consistent intro, a corrected sentence dropped into an episode you already published, and pickups without booking the booth again.

One rule, stated plainly: clone your own voice, or a voice you have written permission to use. Real identifiable people without consent are banned outright, and that includes voices.

## Podcasts, and the parts around the audio

Narration is one of four jobs in a podcast workflow, and the other three are cheap:

- **Scripting** — [Text](/studio/text) with RP+ at 3 credits per completion, or RP mini at 1. The scene register writes immersive prose, which reads aloud better than bullet points. A tightened script in six completions is 18 credits.
- **Two voices** — a dialogue episode is two renders, one per speaker, in Aria and Orion, cut together in your editor. Split the script by speaker and you also get cheap retakes, since a re-render costs only the block you changed.
- **Transcripts and show notes** — Scribe One is speech-to-text at 5 credits per request for audio up to 8 MB, with automatic language detection. Transcribe the finished episode for chapter markers, captions and the SEO copy on the episode page.

Every render is saved in your library automatically, MP3 in hand, so the intro you made in March is still there in September.

## Thirteen languages, one book

The roster speaks 13 languages, which makes a translated edition a re-render rather than a re-casting. The arithmetic carries over directly — a translated 10,000-word chapter is roughly the same 600 blocks, so 1,800 credits on Voice One — and the punctuation you wrote for delivery survives translation better than adverbs do.

When the volume justifies it, the same models are available through the API. `audio/speech` takes the text and the voice and returns the file, and the [speech docs](/docs/speech) list the parameters. One script over a folder of chapter files is a night's work that renders a book.

## FAQ

### How many credits is one minute of narration?

Bill by characters, not minutes. Spoken English lands around 150 words a minute, so a minute is roughly 900 characters, or 9 blocks — 27 credits on Voice One and 18 on Voice Turbo. A 40-minute episode is about 36,000 characters, so 360 × 3 = 1,080 credits on Voice One.

### Can I fix one paragraph without re-rendering the chapter?

Yes, and you should. Render in chunks — a scene, a section, a speaker turn — and a correction costs only that chunk. Re-rendering 500 characters on Voice One is 5 × 3 = 15 credits instead of 1,800.

### Is Voice Turbo noticeably worse?

It is faster and cheaper at 2 credits per 100 characters, with the same roster and the same clones. Voice One is the expressive flagship, which shows most in emotional fiction and least in plain informational narration. Test your own text on both — a 1,000-character sample costs 30 credits on one and 20 on the other.

Got a chapter ready? [Open Voice](/studio/voice), paste 1,000 characters, and render the same passage on both models before you plan a book.
