VoiceStudio

August 10, 2026 · View on GitHub

For headless servers, dedicated GPUs, or "I want one command" deployments. The docker image bundles the backend; the UI is served over HTTP and you open it in a normal browser.

Official images: ghcr.io/debpalash/omnivoice-studio and palashdeb/omnivoice-studio on Docker Hub — same images, same tags.

Image ↔ version mapping

TagWhat you get
:latestRolling preview — latest commit on main, at or ahead of the last release. This is the preview channel; pin :stable for production.
:stableMost recent versioned release (updated on every v* git tag)
:0.4.1Exact release version
:0.4Latest patch within the 0.4 minor
:mainAlias of the same rolling main build as :latest
:sha-xxxxxxxSpecific commit (produced by manual workflow dispatch)
:rocmAMD GPU (ROCm) build of the rolling preview — the ROCm analogue of :latest
:stable-rocm, :0.4.1-rocm, :0.4-rocm, :sha-xxxxxxx-rocmROCm builds of the corresponding CUDA tags above

Versioning rule: preview builds always come from main and never version-sort below :stable — upgrades flow naturally.

Note on the update-channel toggle: The update-channel UI (Settings → About → Update channel) is part of the Tauri desktop app's built-in auto-updater. It does not apply to the Docker image — the Docker image is the headless web-server build. To update your Docker deployment, pull the new image tag and recreate the container (docker compose pull && docker compose up -d).

Pull and run (CPU)

docker pull ghcr.io/debpalash/omnivoice-studio:latest

docker run -d --name omnivoice \
  -p 127.0.0.1:3900:3900 \
  -v omnivoice-data:/app/omnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:latest

Docker Hub mirror: the same images are published to palashdeb/omnivoice-studio on Docker Hub with identical tags — swap the image for palashdeb/omnivoice-studio:latest if you prefer Docker Hub. Tag semantics (:latest = rolling main preview, :stable/:X.Y.Z = releases) are the same on both registries.

Open http://localhost:3900. The first run downloads ~2.4 GB of model weights — follow docker logs -f omnivoice to watch.

Pull and run (NVIDIA GPU)

docker run -d --name omnivoice --gpus all \
  -p 127.0.0.1:3900:3900 \
  -v omnivoice-data:/app/omnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:latest

GPU mode requires the NVIDIA Container Toolkit on the host.

Pull and run (AMD GPU / ROCm)

AMD GPUs use the dedicated :rocm image variant — the default (CUDA) image runs CPU-only on AMD hardware. The ROCm userspace ships inside the image; the host only needs the amdgpu kernel driver. Pass the GPU through as plain device nodes (no container toolkit needed):

docker run -d --name omnivoice \
  --device /dev/kfd --device /dev/dri \
  -p 127.0.0.1:3900:3900 \
  -v omnivoice-data:/app/omnivoice_data \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/debpalash/omnivoice-studio:rocm

The same flags work with Podman (podman run --device /dev/kfd --device /dev/dri …); in a Quadlet unit that's two AddDevice= lines:

# ~/.config/containers/systemd/omnivoice.container
[Container]
Image=ghcr.io/debpalash/omnivoice-studio:rocm
AddDevice=/dev/kfd
AddDevice=/dev/dri
PublishPort=127.0.0.1:3900:3900
Volume=omnivoice-data:/app/omnivoice_data

Release pins exist too: :stable-rocm, :0.4.1-rocm, :0.4-rocm mirror the CUDA tags exactly.

Consumer cards and APUs (RX 6000/7000, Strix Point/Halo): the backend auto-sets HSA_OVERRIDE_GFX_VERSION when — and only when — your card's GFX ID is missing from the shipped ROCm build's architecture list, so try without any override first. Overriding a natively-supported GPU (gfx1151 on ROCm 7.x, for example) only forces it onto foreign kernels. If the GPU still isn't used, force it explicitly with -e HSA_OVERRIDE_GFX_VERSION=11.0.0 (user-set on the container — it is deliberately not baked into the image, because the right value depends on your card); a value you set is always respected as-is.

Rootless / non-root hosts: if /dev/kfd is group-owned, the container user needs those groups too — add --group-add for your host's render and video GIDs (getent group render video).

Verify the container sees the GPU:

docker exec <container> python3 -c \
  "import torch; ok = torch.cuda.is_available(); print(ok, torch.cuda.get_device_name(0) if ok else 'unavailable')"

Use omnivoice for the docker run examples above. Docker Compose names the ROCm container omnivoice-studio-rocm (CPU: omnivoice-studio, NVIDIA: omnivoice-studio-gpu); docker compose ps shows the exact active name.

(ROCm-built PyTorch reports through torch.cuda.*True plus your card's name means torch can see the GPU.) That check alone isn't proof the app is using it: Settings → System shows the device VoiceStudio actually resolved. If it reads cpu while the command above prints True, the backend log line starting Falling back to CPU: names the architecture mismatch it hit.

The image installs and launches VoiceStudio through that same python3 interpreter. To verify this invariant on an older or custom image, compare docker exec <container> python3 -c "import sys, torch; print(sys.executable, torch.version.hip)" with docker exec <container> sh -c 'tr "\\0" " " </proc/1/cmdline'; PID 1 must begin with python3 -m uvicorn.

If the command prints False, Settings → System now says why, and the three answers need different fixes:

What it saysWhat to do
/dev/kfd is not presentThe container was started without --device /dev/kfd --device /dev/dri, or the host's amdgpu driver isn't loaded.
this process cannot open itA group problem. Run ls -l /dev/kfd /dev/dri/render* on the host, and pass those GIDs with --group-add. The numbers differ between machines — a --group-add 39 copied from someone else's command grants nothing.
no GPU was enumeratedThe device nodes are fine and the runtime still found nothing — usually a card newer than the image's ROCm. Check rocminfo on the host, and see the HSA_OVERRIDE_GFX_VERSION note above.
# CPU
docker compose -f deploy/docker-compose.yml --profile cpu up -d

# NVIDIA GPU
docker compose -f deploy/docker-compose.yml --profile gpu up -d

# AMD GPU (ROCm)
docker compose -f deploy/docker-compose.yml --profile rocm up -d

The docker-compose.yml shipped in deploy/ defaults to 127.0.0.1:3900 on the host. The backend inside the container binds to 0.0.0.0 so the host port mapping can forward — the host-side 127.0.0.1 binding is what enforces loopback-only.

LAN access

To expose VoiceStudio on your LAN (e.g. you're running it on a homelab box and opening the UI from a laptop), change the host port mapping:

# deploy/docker-compose.yml
services:
  omnivoice:
    ports:
      - "0.0.0.0:3900:3900"   # ← was 127.0.0.1:3900:3900

The VoiceStudio frontend defaults to the same origin the page was served from, so opening the UI from http://<lan-ip>:3900 Just Works for both the page load and the API/media requests it makes afterwards.

If you front the app with a reverse proxy and the API and UI land on different origins, pin the API base explicitly. Use OMNIVOICE_PUBLIC_API_BASE — a runtime env var the backend injects into the page, so it works with the prebuilt image via docker run -e (the older VITE_OMNIVOICE_API is inlined at build time and cannot be set on a prebuilt image):

docker run -e OMNIVOICE_PUBLIC_API_BASE=https://api.your-host.example \
  -p 0.0.0.0:3900:3900 \
  ghcr.io/debpalash/omnivoice-studio:latest

OMNIVOICE_PUBLIC_API_BASE must be a plain http(s)://… URL; anything else is ignored and the app falls back to same-origin. If you build from source you may instead bake VITE_OMNIVOICE_API at build time, but the runtime var above is simpler and image-agnostic.

Security: VoiceStudio ships no authentication. Anything on your LAN with the URL can use the app. Put it behind a reverse proxy with basic_auth (Caddy / nginx + htpasswd) or a private network overlay (Tailscale, ZeroTier) before exposing publicly.

Volume mounts

Two paths are worth persisting across container restarts:

MountPurposeWhy
omnivoice_data:/app/omnivoice_dataProject DB, user voices, settingsSurvives upgrade; encrypted HF token lives here
~/.cache/huggingface:/root/.cache/huggingfaceHF model cacheRe-using your host's cache saves ~2.4 GB of re-downloads

Troubleshooting

  • Container reports 0.2.7 but image is tagged 0.3.x: This was a workflow bug (fixes #249, #251) — the :latest tag was not being updated on release tag pushes. Pull the image again after the fix is merged: docker pull ghcr.io/debpalash/omnivoice-studio:latest. The running version is now shown in Settings → About → Version (read live from the backend), so the web UI no longer displays a dash in Docker.
  • Checking which version is running: docker exec <container> python3 -c "import importlib.metadata; print(importlib.metadata.version('omnivoice'))", or hit the /health endpoint — it returns {"status": "ok", "device": ..., "version": "0.3.x"}. Use the container name listed by docker compose ps (or omnivoice for the docker run examples).
  • "Loopback origin required" errors (and a blank version): the desktop build restricts the /system/* and /api/settings/* routes to a loopback origin, but Docker's NAT makes every request look non-loopback, so the gate used to 403 the whole admin UI (issue #261). The image now ships with OMNIVOICE_SERVER_MODE=1, which relaxes that gate for the headless deployment — exposure is instead governed by your -p port mapping (keep the 127.0.0.1: prefix to stay local) plus the optional share PIN. If you front the container with your own auth proxy on loopback, set OMNIVOICE_SERVER_MODE=0 to re-enable the strict gate.
  • Media-preview 404 in LAN mode: see the LAN access section above — the window.location.host fix shipped in v0.3.
  • GPU not detected (NVIDIA): verify docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu22.04 nvidia-smi succeeds first.
  • GPU not detected (AMD): make sure you pulled the :rocm tag (the default image is CUDA-only) and passed --device /dev/kfd --device /dev/dri. Check the container sees the card with docker exec omnivoice rocminfo | grep -i gfx. On consumer cards, run without any HSA_OVERRIDE_GFX_VERSION first — the backend sets it itself when your card needs it, and overriding a natively-supported GPU only forces it onto foreign kernels. See Pull and run (AMD GPU / ROCm) above for when to set one by hand.
  • More entries: docs/install/troubleshooting.md.