Compute Server

The compute server is optional. It runs transcription, text-to-speech, video segmentation, voice conversion, camera tracking, and video generation on this machine.

Run locally

The server needs Python 3.14. From the repository root:

$ make dev-server

From the server directory:

$ uv run --locked src/main.py

Models download the first time you use them. Leave the server running until the job finishes.

Connect Shrimply

The server listens at http://127.0.0.1:8787. In Shrimply, open Preferences ‣ External, select the local server under Inference Servers, and pick a device.

Shrimply shows the features that server offers. Anything it does not offer stays out of the editor.

Share access

SHRIMPLY_SERVER_SHARE=1 opens a temporary public gradio.live URL. That URL can run every model. Share it with people you trust, and stop the process when you want it gone.

$ SHRIMPLY_SERVER_SHARE=1 uv run --locked src/main.py

Containers

Docker Compose enables the GPU and keeps downloaded models between runs.

$ cd server
$ docker compose up

Features

Troubleshooting

If a model is missing or the connection fails, check the selected server in Preferences ‣ External. The first request waits while its model downloads.