When you record a conversation or interview with multiple participants, Descript can automatically detect distinct voices and help you label each Speaker. You'll identify each participant once, and Descript applies those labels throughout your transcript.
This article covers
- How speaker detection works
- Speaker detection modes (automatic vs. manual)
- Detect and identify speakers
How speaker detection works
Speaker detection works in two steps:
- Detect: When you import and transcribe a file, Descript's AI automatically analyzes the audio and identifies distinct voices. Automatic Speaker detection is available for files up to 5 hours long.
- Identify: Once detection finishes, you'll get a notification prompting you to assign a name to each Speaker. Descript plays a sample of each voice so you can tell them apart and create labels.
After you identify your Speakers, Descript automatically applies labels throughout your script.
Speaker detection modes
Speaker detection works differently depending on your selected mode (automatic or manual). This App setting can be different between web and desktop apps, so check your settings in both Descript for Web and the desktop app.
| Automatic detection (default) | Manual detection |
|---|---|
| Runs automatically without asking for the number of speakers | Prompts you to select the number of speakers first |
| Up to 10 hours | Up to 3 hours |
| Always ask before detecting speakers is OFF | Always ask before detecting speakers is ON |
If your file exceeds these limits, the Detect speakers option will be disabled. You can check this by right-clicking the file — a tooltip will explain the time restriction.
Detect and identify speakers
- Import your file into a project. Descript begins detecting speakers in the background automatically as part of the automatic transcription process.
- When detection is complete, you’ll get a notification. Click Identify speakers.
- The Identify speakers modal will open. Listen to the audio sample for each detected voice to confirm who's speaking.
- Add each participant's name, or select an existing name from the list.
- When you’ve listened to all the samples, click Close.
- Add the file to your script. The speaker labels will appear automatically in your transcript.
To fix a name, reposition a label, or merge two Speakers afterward, see Speaker labels.
Manually trigger speaker detection
If automatic detection is disabled, speaker detection won't run when you import files. To trigger it manually, click the ellipsis next to your file in the Project panel and select Detect speakers. You'll be prompted to enter the number of speakers before detection begins.
If the Detect speakers option is unavailable, your file may exceed the 3-hour time limit for manual detection.