Skip to content

Commit 7b402c8

Browse files
fix(container): start on CPU when device=auto and align run commands
Resolve SENSEVOICE_DEVICE=auto for CPU Docker hosts Keeps the local-build path documented while public registry pulls remain private. Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
1 parent e8bdeb0 commit 7b402c8

9 files changed

Lines changed: 111 additions & 8 deletions

File tree

‎.github/workflows/sensevoice-container.yml‎

Lines changed: 15 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,23 +4,37 @@ on:
44
pull_request:
55
paths:
66
- Dockerfile
7+
- docker-compose.yaml
78
- requirements.txt
89
- api.py
910
- model.py
11+
- utils/device_env.py
12+
- README.md
13+
- README_zh.md
14+
- README_ja.md
15+
- CONTRIBUTING.md
1016
- .github/workflows/sensevoice-container.yml
1117
- tests/test_container_contract.py
18+
- tests/test_device_env.py
1219
push:
1320
branches:
1421
- main
1522
tags:
1623
- "v*"
1724
paths:
1825
- Dockerfile
26+
- docker-compose.yaml
1927
- requirements.txt
2028
- api.py
2129
- model.py
30+
- utils/device_env.py
31+
- README.md
32+
- README_zh.md
33+
- README_ja.md
34+
- CONTRIBUTING.md
2235
- .github/workflows/sensevoice-container.yml
2336
- tests/test_container_contract.py
37+
- tests/test_device_env.py
2438
workflow_dispatch:
2539

2640
permissions:
@@ -37,7 +51,7 @@ jobs:
3751
steps:
3852
- uses: actions/checkout@v4
3953
- name: Validate container contract
40-
run: python -m unittest tests.test_container_contract
54+
run: python -m unittest tests.test_container_contract tests.test_device_env
4155

4256
build:
4357
needs: contract

‎CONTRIBUTING.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -145,12 +145,14 @@ You can also run SenseVoice using Docker:
145145
docker build -t sensevoice .
146146

147147
# Run with GPU
148-
docker run --gpus all -p 50000:50000 sensevoice
148+
docker run --rm --gpus all -p 50000:50000 -v sensevoice-models:/models sensevoice
149149

150150
# Run on CPU
151-
docker run -e SENSEVOICE_DEVICE=cpu -p 50000:50000 sensevoice
151+
docker run --rm -e SENSEVOICE_DEVICE=cpu -p 50000:50000 -v sensevoice-models:/models sensevoice
152152
```
153153

154+
Do not advertise `docker pull` for `ghcr.io/qwenaudio/sensevoice` or the old Aliyun registry while those paths return HTTP 401 for anonymous clients. Port is 50000, not 60001.
155+
154156
## Questions?
155157

156158
- Open an issue using the [Questions template](https://github.com/QwenAudio/SenseVoice/issues/new?template=ask_questions.md).

‎README.md‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -339,18 +339,22 @@ docker build -t sensevoice .
339339
340340
### Run (GPU – default)
341341
```bash
342-
docker run --gpus all -p 50000:50000 sensevoice
342+
docker run --rm --gpus all -p 50000:50000 -v sensevoice-models:/models sensevoice
343343
```
344344
### Run (CPU-only)
345345
```bash
346346
docker run --rm -e SENSEVOICE_DEVICE=cpu -p 50000:50000 -v sensevoice-models:/models sensevoice
347347
```
348+
349+
The container listens on port 50000. After it is healthy, open `http://127.0.0.1:50000/docs`.
350+
348351
### Docker Compose
349-
Docker Compose provides an easier way to run SenseVoice with persistent model caching, networking etc.
352+
Docker Compose builds the same image and keeps the model cache on the `sensevoice-models` volume. The default compose file does not request GPUs, so it starts on CPU hosts. For GPU, use the `docker run --gpus all` command above.
350353

351354
### Start Stack
352355
```bash
353356
docker compose up --build
357+
# CPU is the default. Equivalent: SENSEVOICE_DEVICE=cpu docker compose up --build
354358
```
355359
### Data prepare
356360

‎api.py‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
11
# Set the device with environment, default is cuda:0
22
# export SENSEVOICE_DEVICE=cuda:1
33

4-
import os, re
4+
import re
55
from fastapi import FastAPI, File, Form, UploadFile
66
from fastapi.responses import HTMLResponse
77
from typing_extensions import Annotated
@@ -11,6 +11,7 @@
1111
from model import SenseVoiceSmall
1212
from funasr.utils.postprocess_utils import rich_transcription_postprocess
1313
from io import BytesIO
14+
from utils.device_env import resolve_sensevoice_device
1415

1516
TARGET_FS = 16000
1617

@@ -26,7 +27,9 @@ class Language(str, Enum):
2627

2728

2829
model_dir = "iic/SenseVoiceSmall"
29-
m, kwargs = SenseVoiceSmall.from_pretrained(model=model_dir, device=os.getenv("SENSEVOICE_DEVICE", "cuda:0"))
30+
m, kwargs = SenseVoiceSmall.from_pretrained(
31+
model=model_dir, device=resolve_sensevoice_device()
32+
)
3033
m.eval()
3134

3235
regex = r"<\|.*\|>"

‎docker-compose.yaml‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -18,7 +18,9 @@ services:
1818
options:
1919
max-size: "5m"
2020
max-file: "3"
21-
gpus: all
21+
# No GPU reservation. A default one makes `docker compose up` fail on
22+
# hosts without the NVIDIA runtime, which is the CPU path from issue 339.
23+
# GPU users should use `docker run --gpus all`.
2224

2325
volumes:
2426
sensevoice-models:

‎tests/__init__.py‎

Whitespace-only changes.

‎tests/test_container_contract.py‎

Lines changed: 32 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -43,11 +43,43 @@ def test_readme_uses_real_service_port_and_no_retired_registry(self):
4343
readme = (ROOT / "README.md").read_text(encoding="utf-8")
4444

4545
self.assertIn("-p 50000:50000", readme)
46+
self.assertNotIn("60001", readme)
4647
self.assertNotIn(
4748
"registry.cn-hangzhou.aliyuncs.com/funasr/sensevoice",
4849
readme,
4950
)
5051

52+
def test_verified_docker_run_is_identical_across_docs(self):
53+
gpu = (
54+
"docker run --rm --gpus all -p 50000:50000 "
55+
"-v sensevoice-models:/models sensevoice"
56+
)
57+
cpu = (
58+
"docker run --rm -e SENSEVOICE_DEVICE=cpu -p 50000:50000 "
59+
"-v sensevoice-models:/models sensevoice"
60+
)
61+
for name in ("README.md", "README_zh.md", "README_ja.md", "CONTRIBUTING.md"):
62+
text = (ROOT / name).read_text(encoding="utf-8")
63+
self.assertIn(gpu, text, msg=name)
64+
if name != "README_ja.md":
65+
self.assertIn(cpu, text, msg=name)
66+
67+
def test_default_compose_starts_without_a_gpu_reservation(self):
68+
compose = (ROOT / "docker-compose.yaml").read_text(encoding="utf-8")
69+
70+
self.assertIn("50000:50000", compose)
71+
self.assertIn("sensevoice-models:/models", compose)
72+
self.assertNotIn("gpus:", compose)
73+
74+
def test_api_uses_resolved_device_not_raw_auto(self):
75+
api = (ROOT / "api.py").read_text(encoding="utf-8")
76+
77+
self.assertIn("resolve_sensevoice_device()", api)
78+
self.assertNotIn(
79+
'SenseVoiceSmall.from_pretrained(model=model_dir, device=os.getenv("SENSEVOICE_DEVICE", "cuda:0"))',
80+
api,
81+
)
82+
5183
def test_readmes_do_not_offer_private_ghcr_image_as_anonymous_pull(self):
5284
for name in ("README.md", "README_zh.md", "README_ja.md"):
5385
readme = (ROOT / name).read_text(encoding="utf-8").lower()

‎tests/test_device_env.py‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,27 @@
1+
import unittest
2+
3+
from utils.device_env import resolve_sensevoice_device
4+
5+
6+
class ResolveSensevoiceDeviceTest(unittest.TestCase):
7+
def test_explicit_cpu(self):
8+
self.assertEqual(resolve_sensevoice_device("cpu", cuda_available=True), "cpu")
9+
10+
def test_explicit_cuda(self):
11+
self.assertEqual(
12+
resolve_sensevoice_device("cuda:1", cuda_available=False), "cuda:1"
13+
)
14+
15+
def test_auto_prefers_cuda_when_present(self):
16+
self.assertEqual(
17+
resolve_sensevoice_device("auto", cuda_available=True), "cuda:0"
18+
)
19+
20+
def test_auto_falls_back_to_cpu(self):
21+
self.assertEqual(
22+
resolve_sensevoice_device("auto", cuda_available=False), "cpu"
23+
)
24+
25+
26+
if __name__ == "__main__":
27+
unittest.main()

‎utils/device_env.py‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
import os
2+
3+
4+
def resolve_sensevoice_device(requested=None, cuda_available=None):
5+
"""Map SENSEVOICE_DEVICE to a FunASR device string.
6+
7+
The image defaults to ``auto``. FunASR/torch want ``cuda:0`` or ``cpu``.
8+
Passing ``auto`` through would fail on hosts without a GPU, which is the
9+
path the original Docker issue was hitting.
10+
"""
11+
if requested is None:
12+
requested = os.getenv("SENSEVOICE_DEVICE", "cuda:0")
13+
if requested != "auto":
14+
return requested
15+
if cuda_available is None:
16+
import torch
17+
18+
cuda_available = torch.cuda.is_available()
19+
return "cuda:0" if cuda_available else "cpu"

0 commit comments

Comments
 (0)