> ## Documentation Index
> Fetch the complete documentation index at: https://help.descript.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate Text-to-speech Audio

> Generate AI speech from a script using a stock voice or your own voice clone.

| Supported languages |                     |                       |         |
| ------------------- | ------------------- | --------------------- | ------- |
| English (US)        | Finnish             | Portuguese (Portugal) | Slovak  |
| Croatian            | French (FR)         | Romanian              | Turkish |
| Czech               | German              | Malay                 | Danish  |
| Hungarian           | Polish              | Spanish (US)          | Dutch   |
| Italian             | Portuguese (Brazil) | Swedish               |         |

## First, choose a voice generation model

Descript supports two AI speech models from ElevenLabs. Switch between them anytime in [App Settings > AI models](https://web.descript.com/view/settings/account?active=ai-models).

* **Multilingual v2 (default):** Reliable, fast generation across all supported languages. The default model for new projects.
* **ElevenLabs v3:** More natural and expressive AI speech, with support for **tone tags** to direct delivery (e.g. whisper, laugh, sigh). Slightly slower generation than v2. Not available on legacy plans.

All [**custom voice clones**](/ai-speech/custom-speaker) are affected by your selected model. Custom voices tend to sound more consistent and respond reliably to tone tags.

[**Stock voices**](/ai-speech/stock-voices) are more expressive on the v3 model. They respond to tone tags more reliably and deliver more natural, animated speech. Because of these new capabilities, some voices may sound slightly different or vary in tone across generations.

The following stock voices are tuned for v3: Edward, Elizabeth, Grace, Joshua, Kyle, Libby, Michael, Owen, Ryan, Sarah, Simon, Ursula, Vernon.

## Then, generate text-to-speech audio

The exact workflow is slightly different for single speaker vs multi-speaker TTS generation.

### **Single speaker**

1. Click **Add speaker** at the top of your composition and select a Speaker.
2. Enter [Write mode](/ai-speech/write-mode) and type your script.
3. When finished, click **Done writing**. Descript automatically generates AI speech for the entire script using your assigned Speaker.<br /><img src="https://mintcdn.com/descript-5bf56f3f/NQQ8Oln3iV-r_AmP/images/external/descript-45807563119885-b9a8b39b8e.png?fit=max&auto=format&n=NQQ8Oln3iV-r_AmP&q=85&s=391d50742033257fbcac304a19cf7d12" alt="Done writing button in Write mode" width="3024" height="1656" data-path="images/external/descript-45807563119885-b9a8b39b8e.png" />

### **Multiple speakers**

1. Enter [Write mode](/ai-speech/write-mode) before selecting a Speaker.
2. Write or paste your full script into the script panel. When you're done, click **Done writing** to exit Write mode.
3. Highlight a paragraph (or any portion of text), press the **@** key, and assign a Speaker. Descript will generate TTS audio in that voice for the selection. Repeat for each section that needs a different Speaker.

### Direct AI voice performance with tone tags

With the ElevenLabs v3 model, you can add **tone tags** to direct how the AI speaks — for example, telling it to whisper, laugh, sigh, or speak seriously. Tone tags appear in your script as gray text inside parentheses and are interpreted by the model when audio is generated.

Tone input isn't available on non-TTS content, on AI speech that's been [converted to audio](/ai-speech/covert-to-audio), or when using the v2 model.

#### Add a tone tag

1. Highlight a selection in your script.
2. Click the **Tone** button (wave icon) in the selection toolbar.<br /><img src="https://mintcdn.com/descript-5bf56f3f/NQQ8Oln3iV-r_AmP/images/external/descript-45807563120909-7cba5befb7.png?fit=max&auto=format&n=NQQ8Oln3iV-r_AmP&q=85&s=2ca7fdbe0c77337bdc6c2a44b66256e6" alt="toneTagsHoverMenu.png" width="1347" height="753" data-path="images/external/descript-45807563120909-7cba5befb7.png" />
3. Pick a preset like *serious*, *sigh*, *laugh*, or *long pause* — or choose **Custom** to write your own.

The tag appears in your script as gray text inside parentheses, and the AI uses it to shape its delivery on your next generation. Custom tone tags can't be longer than 150 characters.

<Note>
  **Other ways to add a tone tag**

  The Tone button is the easiest way, but you can also:

  * **Type it manually.** While in Write Mode, place your cursor where you want the tag, type an opening parenthesis `(`, enter your tag, and close it with `)`.
  * **Use** [**the action bar**](/descript-tour/action-bar). Open with `command/ctrl + K` and search for **Inline note**.
</Note>

#### Where tone tags appear

| Surface                     | Visible?                                   |
| --------------------------- | ------------------------------------------ |
| Your script (in the editor) | Yes — gray text in parentheses             |
| Generated audio             | Interpreted by the model, not spoken aloud |
| Exported transcripts        | Yes — included as text                     |
| Captions                    | No                                         |
| Share page transcript       | No                                         |

<Warning>
  **Square brackets no longer supported for tone**

  In earlier versions, you could type prompts directly into your script using square brackets — for example, `[whispers]` or `[sigh]`. That approach is no longer supported on v3.

  If your script contains raw `[...]`, Descript will block AI speech generation and prompt you to fix it. To migrate an existing script, replace any square-bracket prompts with parentheses, or use the **Tone** button to insert tags through the new flow.
</Warning>

## After you generate

AI-generated speech behaves differently from recorded audio. To make timeline edits, apply fades or crossfades, or get precise control over playback, you'll need to [convert it to a standard audio layer](/ai-speech/covert-to-audio) first.

<img src="https://mintcdn.com/descript-5bf56f3f/NQQ8Oln3iV-r_AmP/images/external/descript-45807563121421-d3e287aeac.png?fit=max&auto=format&n=NQQ8Oln3iV-r_AmP&q=85&s=e3a3ecb1fc5a5795ac2c22af5deeaf82" alt="Convert to audio option in the clip context menu" width="822" height="450" data-path="images/external/descript-45807563121421-d3e287aeac.png" />

## Known limitations

* **Generation speed.** Voice generation with ElevenLabs v3 is slightly slower than Multilingual v2, especially for long paragraphs or when tone tags are included.
* **Tone and voice continuity.** You may notice inconsistencies in how the output from this model sounds across paragraphs. This can include shifts in tone, volume, accent, or even speaker identity. These variations affect both custom and stock voices and are more likely to occur in longer or more complex scripts.
* **Custom tone tags are experimental.** Short, direct instructions like `(annoyed)` or `(slowly)` tend to work, but longer descriptive prompts may be spoken aloud instead of treated as direction. If a custom tag isn't working as expected, try shortening it.
* **Tone tag scope.** A tone tag affects delivery in the paragraph where it's added. The model decides how long the effect lasts — Descript doesn't control which exact words a tag applies to. Use paragraph breaks to reset the delivery.

## FAQs and troubleshooting

<AccordionGroup>
  <Accordion title="Audio isn't generating">
    Confirm the Speaker has a voice assigned. If not, assign one in the Speaker card.
  </Accordion>

  <Accordion title="Mispronounced words">
    AI Speakers may occasionally mispronounce words. See our [pronunciation guide](/ai-speech/pronunciation).
  </Accordion>

  <Accordion title="Black frames appear after generating TTS">
    TTS and Regenerate don't work over sequences. If used on a sequence, video may be removed.

    **Workaround:**

    1. [Convert the AI voice clip into an audio layer](/ai-speech/covert-to-audio).
    2. Cut the AI audio clip from the script (`Cmd + X` / `Ctrl + X`).
    3. Restore the original script track by expanding the clip in the timeline or using Undo.
    4. Paste the AI audio as a layer above the original script track.
    5. Split and mute the original script section using [the Blade tool](/timeline/timeline-tools).
  </Accordion>

  <Accordion title="Unexpected background noise">
    Artifacts usually come from the original training audio. Try to avoid:

    * Static or sudden loud sounds
    * Background noise (traffic, appliances, music)
    * Excessive mouth noise or breathing
  </Accordion>
</AccordionGroup>
