Compute Server¶
The compute server is optional. It runs transcription, text-to-speech, video segmentation, voice conversion, camera tracking, and video generation on this machine.
Run locally¶
The server needs Python 3.14. From the repository root:
$ make dev-server
From the server directory:
$ uv run --locked src/main.py
Models download the first time you use them. Leave the server running until the job finishes.
Connect Shrimply¶
The server listens at http://127.0.0.1:8787. In Shrimply, open
, select the local server under
Inference Servers, and pick a device.
Shrimply shows the features that server offers. Anything it does not offer stays out of the editor.
Containers¶
Docker Compose enables the GPU and keeps downloaded models between runs.
$ cd server
$ docker compose up
Features¶
Troubleshooting¶
If a model is missing or the connection fails, check the selected server in . The first request waits while its model downloads.