# How long should a voice sample be for a good AI voice clone

> The 30 seconds to 3 minutes range, the room, mic distance, what to read aloud so the clone has range, and the four habits that quietly ruin a sample.

Published 2026-09-02 · eroq.ai — canonical: https://eroq.ai/blog/how-long-should-a-voice-sample-be


Every question about voice cloning ends up being a question about the sample, because the sample is the clone. The engine has no idea what you sound like on a good day — it knows the ninety seconds you handed it, including the fridge in the next room. This is what to record, how much of it, and the handful of habits that quietly wreck an otherwise fine take.

## The range, and why it is a range

The clone form in the [Voice studio](/studio/voice) says it without hedging: **30 seconds to 3 minutes** of continuous natural speech. Not thirty seconds assembled from clips — thirty seconds of you actually talking.

Under thirty seconds, the engine has heard one cadence and one pitch range and will happily reproduce exactly that, forever. Past three minutes, you are mostly adding more of what it already learned, and the odds that something goes wrong in the file — a cough, a phone buzz, a sentence where you drifted off the mic — go up with every extra minute.

The other half of the guidance matters more than the numbers: **clean beats long**. One clean minute outperforms ten noisy ones. If you have a choice between a three-minute take with a hum in it and a fifty-second take without, upload the fifty seconds.

## The room decides more than the microphone

[Voice cloning](/glossary/voice-cloning) copies what it hears, and it has no concept of "that part was the room". Reverb gets learned as part of your timbre, which is why a clone recorded in a kitchen sounds like it is permanently in a kitchen.

- Kill the obvious noise sources first — air conditioning, fan, fridge, an open window onto a street.
- Record somewhere soft. Curtains, a bed, a sofa, a wardrobe full of coats. Hard parallel surfaces are what makes echo.
- Put the phone on airplane mode. A notification buzz through a desk is louder on the recording than it was in the room.

## One speaker, one distance

The form is explicit here too — solo voice only, no crosstalk, no interviewer — and the mic close enough that the room stops being audible.

Pick a distance and hold it for the entire take. Leaning in for the intimate lines and back for the loud ones is good mic technique for a live read and terrible input for a clone, because the engine hears two different voices and averages them. Angle yourself slightly off-axis so plosives miss the capsule rather than thumping it.

## Read something with range in it

This is the part most people skip, and it is the difference between a clone that can act and one that can only announce. The engine clones a manner, not just a timbre. Read one flat paragraph and you get one flat voice.

Give it a statement, a question, a list, something quiet and something sharp. Here is a script that covers all of them in about ninety seconds — read it at your normal speaking pace, and do not perform it:

> Right, let me do this properly. My name is Mara, and this is a sample for a voice clone, which is a strange sentence to say out loud.
>
> A plain statement, first. The kettle boiled ten minutes ago and nobody has made the tea.
>
> Now a question — do you honestly think that is going to work?
>
> A list, read the way I actually read a list. Keys, wallet, charger, and the one thing I always forget.
>
> Here is the quiet one. I did not want to say it like that. I am saying it anyway.
>
> And the sharp one. No. That is not what we agreed, and you know it.
>
> One last line to finish on, unhurried, the way you would read the end of a chapter to someone who is almost asleep.

Swap the name, keep the shape. If the clone is destined for one job — calm narration, an ad read, a character who bites — bias the script toward that job, but leave at least one line of contrast in so the clone knows the voice can move.

## The four habits that ruin a sample

1. **Music or ambience under the voice.** There is no separating it afterward. A backing track becomes part of the clone.
2. **Processing.** Compressors, noise gates, de-essers and the "enhance voice" switch on a recorder app all remove information the engine wanted. Raw is better input than polished.
3. **One emotion for three minutes.** Flat in, flat out.
4. **Dead air at the ends.** Trim the silence before your first word and after your last. It is not harmful, it is just wasted budget.

## Uploading it, and what happens next

The clone form takes MP3, WAV, M4A, FLAC, MP4, MOV, WEBM, WEBA, OPUS and MID, up to twenty samples and 1 GB in total — which is far more headroom than the guidance suggests you use. Name the clone something you will recognize in six months ("Mara — narration" beats "test 3").

Cloning itself is free and synchronous. The clone lands in your [voice library](/studio/voices) next to the roster voices, Aria and Orion, and works on both speech models. The paid part is the speech you render with it, at 3 credits per 100 characters on [Voice One](/models/eroq-voice-one) and 2 on Voice Turbo.

## Audition it before you trust it

Render the same line on both models before you commit a project. Something with a pause, a question and a hard stop in it does more work than a nice sentence:

> I waited. Two hours, in the rain, for that? No — don't explain. Just tell me one true thing and I'll go.

That line is 104 characters, and billing counts started blocks of 100 — so it is two blocks, 6 credits on Voice One and 4 on Turbo. If it comes back sounding like a stranger, the fix is almost always the sample rather than the settings — re-record in a softer room, closer, with more variety, and clone again. [The voice cloning guide](/blog/ai-voice-cloning-guide) covers the rest of the workflow, and [Voice One vs Voice Turbo](/blog/voice-one-vs-voice-turbo) covers which model to render the final on.

## FAQ

### Is a longer voice sample always better?

No. The guidance is 30 seconds to 3 minutes, and cleanliness beats duration inside that range. One clean minute outperforms ten noisy ones, because every extra minute is another chance to record a cough, a hum or a sentence you drifted away from the mic on.

### Can I combine several short recordings into one clone?

You can upload up to twenty samples, but continuous natural speech beats a pile of short clips. If the clips were recorded in different rooms or at different distances, the engine hears an inconsistent voice and averages it. One unbroken read is the better input.

### Does a phone work as a microphone for cloning?

Yes, if you treat it like a microphone. Hold it at a fixed distance, record in a soft room, turn off any "enhance voice" processing in the recorder app, and do not use speakerphone or a car. A phone in a wardrobe beats a good microphone in a kitchen.

### What if the clone does not sound like me?

Re-record rather than reach for the sliders. Speed and expressiveness pull the read away from the sample, so if the identity is wrong they cannot fix it. A closer, quieter, more varied sample is the actual repair.

Ready to record? [Open the Voice studio](/studio/voice), clone from the voice library, and spend the first credits on an audition line rather than a whole script.
