Transcription

Transcription turns selected audio into timed captions on a new caption track.

Create captions

  1. Select one or more audio clips, or select an audio track.

  2. Open the timeline context menu and choose Transcribe.

  3. Choose one of the speech-to-text models offered by the compute server.

  4. Select Transcribe and wait for the operation to finish.

The current server can offer Parakeet, Qwen3 ASR, Whisper, and Distil-Whisper. If none appear, check the selected server in Preferences ‣ External.

Follow edit points

Keep Follow cuts enabled when caption boundaries should follow nearby audio or video edits. Snap source chooses which cuts to follow, while Snap tolerance controls how close a generated boundary must be before it moves to a cut.

The defaults are suitable for most projects. Disable Follow cuts when you want the transcription model’s continuous timing without edit-based chunking.

If no speech is detected, Shrimply leaves the project unchanged.