# Podcast intro with AI voice — write a cold open that lands

> Script a cold open and a show ID for the ear, direct the read with punctuation, pick a voice, then lay the MP3 under a music bed in your own editor.

Published 2026-09-09 · eroq.ai — canonical: https://eroq.ai/blog/podcast-intro-with-ai-voice


The first fifteen seconds of a podcast decide whether the next forty minutes happen. Most shows spend them on a logo sting and a name, which is the audio equivalent of opening a film on the production company card. This is how to write and render the two pieces that actually work — a cold open and a show ID — in about the time it takes to make coffee.

One thing to settle up front so you plan the session correctly: eroq renders speech, not shows. There is no timeline, no music generation and no mixer. What you get back is an MP3, and the bed goes underneath it in whatever editor you already use.

## Two pieces, two different jobs

A **cold open** starts mid-thought. No greeting, no name, no context — a sentence from the middle of the episode that makes the listener need the surrounding minute. It runs eight to fifteen seconds and it is written per episode.

A **show ID** is the thing that never changes. Name of the show, what it is, who makes it, and out. Ten seconds at most, rendered once and reused for a year.

They fail in opposite directions. Cold opens fail by explaining. Show IDs fail by trying to be interesting. Write them separately, render them separately, and cut them together with the bed in your editor.

## Write for the ear, which is not how you write for the page

- **One idea per sentence, and short sentences.** A listener cannot re-read.
- **Put the subject first.** Subordinate clauses that arrive before the point get lost in audio.
- **Spell numbers the way you say them.** "Nineteen ninety-six", not "1996". "About three grand", not "$3,000".
- **Names early, once.** A name introduced in the last clause of a sentence does not stick.
- **No parentheses, no semicolons.** If a thought needs brackets, it needs its own sentence.
- **Read it aloud before you render it.** If you stumble, the engine will too.

Punctuation is the direction track. The speech engines perform commas, ellipses, dashes and sentence length — there is no separate markup — so write the pauses in rather than hoping for them.

## A cold open you can steal

Written for a true-crime-adjacent show, but the shape transfers to anything:

> She signed the last page at four in the afternoon. By six, the building was empty, the servers were wiped, and the company that had existed that morning did not exist any more.
>
> Nobody has ever been charged.
>
> This is what happened in the ninety minutes in between.

Look at what the punctuation is doing. The comma list in the first sentence accelerates. The one-line paragraph after it lands as a full stop with air around it. The last line is short because the listener should be leaning in by then, not being informed.

## A show ID you can steal

> *Nightline Files.* True stories about the hour nobody was watching.
>
> New episodes every Thursday — hosted by Dana Ruiz.

Ten seconds, no adjectives doing work the delivery should do, and an em dash where a breath belongs.

If the script is the part you get stuck on, draft it in the [Text studio](/studio/text) first — the scene register writes prose meant to be read aloud, and a completion is 3 credits on RP+ or 1 on RP mini, which is cheaper than an hour of staring.

## Pick the voice, then set two dials

The roster is two voices, cast for different registers. **Aria** is warm and intimate — she leans into the mic, which is what a cold open wants. **Orion** is low and unhurried and reads like late-night radio, which is what a show ID wants. Both read 13 languages. If the show has a host, cloning the host's own voice is the better answer for the ID, because the audience should recognize it.

Then the two controls next to the model:

- **Cold open, Aria** — speed 1.00, expressiveness 65%. The punctuation carries the drama, so the dial only has to allow it.
- **Show ID, Orion** — speed 0.95, expressiveness 35%. Slightly slow, deliberately flat. An excited station ID sounds like an advert.

Move one at a time. Expressiveness above roughly 80% starts performing punctuation you did not intend, and speed below 0.85 turns gravitas into sedation. The full range is 0.70× to 1.30× on speed and 0 to 100% on expressiveness.

## Render it, and what that costs

Open the [Voice studio](/studio/voice), paste the cold open, pick the voice, set the dials, generate, and download the MP3. Iterate on [Voice Turbo](/models/eroq-voice-turbo) at 2 credits per 100 characters while you are still fixing sentences, then do the final pass on [Voice One](/models/eroq-voice-one) at 3.

The arithmetic is small enough to ignore, which is the point. Billing counts started blocks of 100 characters, so the cold open above — about 260 characters — is 3 blocks, or 9 credits on Voice One. The show ID is about 120 characters, so 2 blocks, or 6 credits. Rendering both fifteen times over while you tune the writing comes to 225 credits, a little over two dollars at the entry pack rate, and the 50 free credits on a new account cover the first three passes outright.

Every render is saved to your library automatically, so the ID you made in March is still there in September when you need to rebuild the intro.

## Lay it under a bed, in your editor

This is the step eroq does not do, and pretending otherwise would cost you an hour. Take the MP3 into Reaper, Audition, Audacity, Premiere — whatever you already have — and build the intro there:

1. Voice on its own track, bed on another, music sourced from wherever you normally license it.
2. Duck the bed under the voice by six to ten decibels, with a short fade in and a longer fade out.
3. Let the bed start about a second before the first word and run two seconds past the last one.
4. Leave a beat of silence between the cold open and the show ID. That gap is doing more work than the sting.
5. Export at your normal loudness target, and check it on phone speakers — most of your audience is on them.

For the video side of the same episode, [AI video for podcast clips](/blog/ai-video-for-podcast-clips) covers turning the audio into something postable, and [AI voice for audiobooks and podcasts](/blog/ai-voice-for-audiobooks-and-podcasts) has the long-form credit math if you are narrating whole segments rather than an intro. The [audiobooks and podcasts use case](/use-cases/audiobooks-podcasts) is the short version of both.

## FAQ

### Can I make the music bed in eroq too?

No. There is no music generation and no timeline — the voice models return an MP3 and the mix happens in your own editor. Source the bed the way you already do and duck it under the voice there.

### Should I use a cloned voice or a roster voice for a podcast intro?

Use a clone when the audience should recognize the host, and a roster voice when the intro is meant to sound like a station announcer rather than a person. Aria suits intimate cold opens, Orion suits show IDs, and both work on either speech model.

### How long should a cold open be?

Eight to fifteen seconds, which is roughly 200 to 350 characters of script. Longer than that and it stops being a hook and starts being the episode, which is what the listener came for anyway.

### How do I make the voice pause between two lines?

With punctuation and paragraph breaks — a period stops, an ellipsis trails off, a dash cuts, and a new paragraph reads as a bigger separation. There is no pause tag, so if you need a precise gap, cut it in your editor.

Got a script? [Open the Voice studio](/studio/voice) and render the cold open first — it is the part that decides whether the rest gets heard.
