Voice

Diarization

Diarization is the part of transcription that works out who spoke when, labelling a transcript by speaker. It is what turns a wall of text from an interview into a readable two-column conversation.

It is a separate model from the one that turns sound into words, which is why plenty of transcription services return accurate text with no speaker labels at all.

Word-level timestamps are a related but distinct feature, and the one subtitle files actually need.

On eroq

Scribe One returns text, at 5 credits a request, with automatic language detection. It does not diarize, does not return word-level timestamps and does not produce a subtitle file — so it suits captioning your own narration rather than cutting up an interview.

Questions

Can eroq label who said what in a recording?

No. Scribe One returns the text of what was said with no speaker labels, so a multi-speaker recording comes back as one continuous transcript.

Can I get an SRT file for subtitles?

Not from eroq — there are no word-level timestamps and no subtitle output. The transcript is text you take into a captioning tool.

What is transcription good for here, then?

Turning your own narration back into text — for captions you time yourself, for a description, or for checking what a generated voice actually said.