# Diarization

> Diarization is the part of transcription that works out who spoke when, labelling a transcript by speaker. It is what turns a wall of text from an interview into a readable two-column conversation.

It is a separate model from the one that turns sound into words, which is why plenty of transcription services return accurate text with no speaker labels at all.

Word-level timestamps are a related but distinct feature, and the one subtitle files actually need.

## On eroq

Scribe One returns text, at 5 credits a request, with automatic language detection. It does not diarize, does not return word-level timestamps and does not produce a subtitle file — so it suits captioning your own narration rather than cutting up an interview.

## FAQ

### Can eroq label who said what in a recording?

No. Scribe One returns the text of what was said with no speaker labels, so a multi-speaker recording comes back as one continuous transcript.

### Can I get an SRT file for subtitles?

Not from eroq — there are no word-level timestamps and no subtitle output. The transcript is text you take into a captioning tool.

### What is transcription good for here, then?

Turning your own narration back into text — for captions you time yourself, for a description, or for checking what a generated voice actually said.

## Related

- https://eroq.ai/glossary/speech-to-text — Speech-to-text
- https://eroq.ai/docs/transcriptions — Transcriptions API
- https://eroq.ai/developers/voice — Voice API

---

This page as HTML: https://eroq.ai/glossary/diarization · Studio: https://eroq.ai/studio · Docs: https://eroq.ai/docs · Machine index: https://eroq.ai/llms.txt
