cabef1bfe0
* fix(docker): add ROCm GPU support via compose overlay Fixes #618. The Docker image installs CPU-only PyTorch from PyPI by default, so even when users correctly pass /dev/kfd and /dev/dri device nodes into the container, torch.cuda.is_available() returns False and the GPU is reported as "None (CPU only)". Changes: - Dockerfile: add PYTORCH_VARIANT build arg (default: cpu). When set to "rocm", the ROCm-enabled PyTorch wheels are installed from the pytorch.org/whl/rocm6.3 index before requirements.txt runs, so pip sees the ROCm build as already satisfying the torch>=2.2.0 constraint and does not overwrite it with the CPU wheel. The render and video groups are created with parameterised GIDs (RENDER_GID / VIDEO_GID, defaulting to Ubuntu 22.04 values) and the voicebox user is added to both groups so it can open /dev/kfd and /dev/dri. - docker-compose.rocm.yml: new compose overlay that wires everything together — PYTORCH_VARIANT=rocm build arg, /dev/kfd + /dev/dri device passthrough, group_add for render/video, HSA_OVERRIDE_GFX_VERSION (defaults to 11.0.0 for RDNA3/Strix Halo with a comment listing values for RDNA2/RDNA1/Vega), and PYTORCH_HIP_ALLOC_CONF for the memory allocator. Usage: docker compose -f docker-compose.yml -f docker-compose.rocm.yml up --build - docker-compose.yml: add a comment pointing to the ROCm overlay. The CPU default path is unchanged — no extra build time, no size increase. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(docker): address review comments on ROCm overlay Two issues raised in PR review: 1. CodeRabbit: `docker compose up --build-arg` is not supported by the `up` subcommand. Replaced the GID override instructions with the correct env-var export pattern. Added RENDER_GID and VIDEO_GID to `build.args` using ${VAR:-default} interpolation so a single export covers both the Dockerfile group creation and the runtime group_add. Changed group_add entries from hardcoded strings to the same interpolated vars so host GIDs stay in sync end-to-end. 2. @Xarianne: ROCm 6.3 does not support RDNA 4 (RX 9000 series) cards. Added a ROCM_VERSION build arg (default 6.3) to both the Dockerfile and docker-compose.rocm.yml so users can set ROCM_VERSION=7.2 for RDNA 4 support without editing any files. Added RDNA 4 / 12.0.0 to the HSA_OVERRIDE_GFX_VERSION comment table. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
48 lines
1.2 KiB
YAML
48 lines
1.2 KiB
YAML
# Voicebox — CPU build (default)
|
|
# For AMD ROCm GPU acceleration use the overlay:
|
|
# docker compose -f docker-compose.yml -f docker-compose.rocm.yml up --build
|
|
|
|
services:
|
|
voicebox:
|
|
build: .
|
|
container_name: voicebox
|
|
restart: unless-stopped
|
|
|
|
ports:
|
|
# Host-side moved to 17600 so the dev/installed Voicebox can keep 17493.
|
|
# Container still listens on its native port internally.
|
|
- "127.0.0.1:17600:17493"
|
|
|
|
volumes:
|
|
# Bind-mount for generated audio (customize the host path as needed)
|
|
# Host side: ./output/
|
|
# Container side: /app/data/generations/
|
|
- ./output:/app/data/generations
|
|
|
|
# Named volume for profiles, DB, cache (persists across container restarts)
|
|
- voicebox-data:/app/data
|
|
|
|
# HuggingFace model cache (so models aren't re-downloaded on rebuild)
|
|
- huggingface-cache:/home/voicebox/.cache/huggingface
|
|
|
|
environment:
|
|
- LOG_LEVEL=info
|
|
- NUMBA_CACHE_DIR=/tmp/numba_cache
|
|
|
|
networks:
|
|
- voicebox-net
|
|
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
cpus: '4'
|
|
memory: 8G
|
|
|
|
networks:
|
|
voicebox-net:
|
|
driver: bridge
|
|
|
|
volumes:
|
|
voicebox-data:
|
|
huggingface-cache:
|