Skip to content

Field Report: MuseTalk V1.5 working on RTX 5060 Ti (Blackwell sm_120) with Python 3.12 + mediapipe patch #409

Description

@chefboyrdave21

Summary

Got MuseTalk V1.5 running end-to-end on an RTX 5060 Ti (Blackwell, sm_120, 16GB VRAM) with Python 3.12 on Ubuntu 24.04. Sharing findings for the community since Blackwell GPU + Python 3.12 is a common setup that currently doesn't work out of the box.

Two Issues & Solutions

1. PyTorch + Blackwell (sm_120)

MuseTalk recommends PyTorch 2.0.1+cu118, but Blackwell GPUs need cu128+ for native sm_120 kernel support.

PyTorch CUDA sm_120?
2.6.0+cu124 12.4 ❌ no kernel image errors
2.10.0+cu126 12.6 ❌ Same errors
2.10.0+cu128 12.8 ✅ Works!
pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128

2. mmpose/mmcv on Python 3.12

mmcv has no pre-built wheels for Python 3.12 on any CUDA index. Building from source fails due to pkg_resources removal in Python 3.12.

Workaround: Replace mmpose with mediapipe (Tasks API) + face_alignment in musetalk/utils/preprocessing.py:

  • pip install mediapipe face-alignment
  • Use mediapipe's 478-point face mesh instead of mmpose's wholebody model
  • Map mediapipe landmarks → MuseTalk's nose bridge indices (28-30):
    • face_lm[28] = pts_478[6] (nose bridge top)
    • face_lm[29] = pts_478[197] (nose bridge mid)
    • face_lm[30] = pts_478[195] (nose bridge lower)
  • Use fa.get_landmarks_from_image() instead of deprecated get_detections_for_batch()

Performance

  • 7sec audio → 30sec inference → MP4 output
  • ~3.6GB VRAM for MuseTalk models
  • Reference image face detection cached after first call

Suggestion

Consider adding mediapipe as an alternative preprocessing backend for users on Python 3.12+ where mmpose isn't installable. Happy to contribute a PR if there's interest.

Setup: RTX 5060 Ti 16GB | Ubuntu 24.04 | Python 3.12.3 | PyTorch 2.10.0+cu128 | MuseTalk V1.5

Activity

  1. HeimdallCore commented on Aug 5, 2026

    @HeimdallCore

    We ran into the same Blackwell (RTX 5060 Ti, sm_120) situation, but hit a different, harder blocker than the Python-3.12 issue described here: mmcv (we tried both 2.0.1 and the latest 2.2.0) fails to compile against current PyTorch with an ambiguous-overload C++ error around Float8_e5m2fnuz operators — a genuine upstream incompatibility, not fixable with a flag.

    Instead of patching around mmcv or converting mediapipe's 478-point mesh to the 68-point dlib format this repo expects (worth flagging: a couple of the index mappings floating around, e.g. for points 28-30, don't line up when we cross-checked them against independent mediapipe→dlib conversion tables), we replaced mmpose/dwpose landmark detection entirely with the standalone face-alignment PyPI package (MIT, FAN-based) — no MuseTalk dependency, no CUDA/C++ build step, and it returns native 68-point dlib-format landmarks directly, so no conversion table is needed at all.

    Our patch is 2 files, musetalk/utils/preprocessing.py:

    • swap the mmpose.apis.init_model init for face_alignment.FaceAlignment(LandmarksType.TWO_D, ...)
    • swap both inference_topdown()/merge_data_samples()/keypoints[0][23:91] extraction blocks for landmark_model.get_landmarks_from_image(...)

    We verified this end-to-end (real MuseTalk V1.5 inference, 361 frames, correct lip-sync in the output). Happy to open this as its own PR if that's useful — didn't want to duplicate effort if someone's already down this path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions