Automatic speaker detection

When you record a conversation or interview with multiple participants, Descript can automatically detect distinct voices and help you label each Speaker. You'll identify each participant once, and Descript applies those labels throughout your transcript.

Speaker labels applied in a Descript transcript after automatic speaker detection

This article covers

How speaker detection works

Speaker detection works in two steps:

  1. Detect: When you import and transcribe a file, Descript's AI automatically analyzes the audio and identifies distinct voices. Automatic Speaker detection is available for files up to 5 hours long.
  2. Identify: Once detection finishes, you'll get a notification prompting you to assign a name to each Speaker. Descript plays a sample of each voice so you can tell them apart and create labels.

After you identify your Speakers, Descript automatically applies labels throughout your script.

Speaker detection modes

Speaker detection works differently depending on your selected mode (automatic or manual). This App setting can be different between web and desktop apps, so check your settings in both Descript for Web and the desktop app.

Automatic detection (default) Manual detection
Runs automatically without asking for the number of speakers Prompts you to select the number of speakers first
Up to 10 hours Up to 3 hours
Always ask before detecting speakers is OFF Always ask before detecting speakers is ON

If your file exceeds these limits, the Detect speakers option will be disabled. You can check this by right-clicking the file — a tooltip will explain the time restriction.

Detect and identify speakers

  1. Import your file into a project. Descript begins detecting speakers in the background automatically as part of the automatic transcription process.
  2. When detection is complete, you’ll get a notification. Click Identify speakers.
    Identify speakers prompt in Descript displayed after automatic speaker detection completes
  3. The Identify speakers modal will open. Listen to the audio sample for each detected voice to confirm who's speaking.
  4. Add each participant's name, or select an existing name from the list.
    Add speaker name dialog in Descript
  5. When you’ve listened to all the samples, click Close.
  6. Add the file to your script. The speaker labels will appear automatically in your transcript.

To fix a name, reposition a label, or merge two Speakers afterward, see Speaker labels.

Manually trigger speaker detection

If automatic detection is disabled, speaker detection won't run when you import files. To trigger it manually, click the ellipsis next to your file in the Project panel and select Detect speakers. You'll be prompted to enter the number of speakers before detection begins.

Project file actions menu to manually initiate speaker detection

If the Detect speakers option is unavailable, your file may exceed the 3-hour time limit for manual detection.