Audio and Captions

Captions

Caption tracks store timed text independently from visual and audio tracks. The inspector controls text, writing direction, layout, and appearance. Shrimply can import and export WebVTT captions.

Transcription

Transcription converts selected audio into timed text. The available model catalog comes from the connected compute server, so start the server before using transcription. See Compute Server.

Text-to-speech and voice conversion

Audio tracks can contain generated speech. The compute server determines which text-to-speech and voice-conversion models appear in Shrimply.

Lip sync expressions

Expressions can call mouth() to obtain Rhubarb mouth cues for the project audio mix or selected audio tracks. See Lip sync for cue values, selection syntax, and cache behavior.

Audio cleanup

Timeline actions can remove silence or export the audio for selected content. Use audio modifiers for nondestructive cleanup and sound design within the project.