Text to Speech

Create a speech item

  1. Open the add menu on an audio track.

  2. Choose Text to Speech.

  3. Select the new item and enter its text in the inspector.

  4. Choose a model and set the controls it shows.

  5. Generate the speech.

Generate speech for captions

Select caption clips and choose Generate Speech from the timeline context menu. The same action on a caption track reads every caption that has text.

Pick a model that supports caption timing. Each caption is spoken at its current duration. The audio lands on a track where it does not overlap existing clips.

Models and licenses

The server supports Qwen3 TTS (built-in voices, voice cloning, and voice design) and IndexTTS 2 and 2.5.

IndexTTS is under the bilibili Model Use License Agreement. Read that agreement before downloading the model.