Skip to content

To achieve higher fps #377

Description

@ljz1998zzzz

Hi there, I noticed there are lots of issues about failing to achieve the same fps as stated in this project ( 30 fps on v100 ).

Here is one tip:
users should ensure that [round(audio time length * instructed fps) / batch size = N (N∈Z)]

following this, the inference performance can be increased significantly

Activity

  1. ljz1998zzzz commented on Sep 10, 2025

    @ljz1998zzzz
    Author

    Here are some detailed experiment results:
    audio length sun.wav : 22 s 488 ms
    when I set 16 fps and batch size 8, inference fps =359/22.88=15.7, notice there is a long delay at the beginning of the last batch. this should be the problem. ps: the selected area stands for the time between two consecutive entries into the iteration in process_frames function.

    Image

    then I set 21 fps and batch size 8, inference fps =472/17.39=27.1, notice there is no long delay at the beginning of the last batch.

    Image

    Hope this is helpful.

  2. ljz1998zzzz commented on Sep 12, 2025

    @ljz1998zzzz
    Author

    Further exploring this issue, there is also a long delay at the beginning of the first batch inference. So I add a warm up module. Now we get 50 fps without saving images and 33 fps with saving images on single A100.

    Image Image

    yet when I do this, the audio feature extraction part took longer time (3.4s than 1.2s), which needs further research.

    here is warm up module welcome to use:

    `@torch.no_grad()
     def warm_up():
        warmup_audio_path = 'your/path/to/audio_demo/sun.wav'
        latent_batch = torch.rand(args.batch_size,8,32,32,device=device,dtype=unet.model.dtype)
         audio_feature_batch = torch.rand(args.batch_size,50,384,device=device,dtype=unet.model.dtype)
         ############################################## warmup audio feature ##############################################
        start_time = time.time()
         # Extract audio features
         whisper_input_features, librosa_length = audio_processor.get_audio_feature(warmup_audio_path, weight_dtype=weight_dtype)
         whisper_chunks = audio_processor.get_whisper_chunk(
             whisper_input_features,
             device,
             weight_dtype,
             whisper,
             librosa_length,
             fps=args.fps,
             audio_padding_length_left=args.audio_padding_length_left,
             audio_padding_length_right=args.audio_padding_length_right,
         )
         audio_time = time.time()
         print(f"audio feature warmup: costs {(audio_time - start_time) * 1000}ms")
    
         ############################################## warmup latents ##############################################
         pred_latents = unet.model(latent_batch,
                                   timesteps,
                                   encoder_hidden_states=audio_feature_batch).sample
         pred_latents = pred_latents.to(device=device, dtype=vae.vae.dtype)
         recon = vae.decode_latents(pred_latents)
         latent_time = time.time()
         print(f"latents warmup: costs {(latent_time - audio_time) * 1000}ms")`
    
  3. EsraaK-Fouad commented on Oct 29, 2025

    @EsraaK-Fouad

    thanks a lot for sharing this detailed explanation and the warm-up module, it’s really helpful!

    I tried the warm-up technique you mentioned, but I didn’t notice much difference compared to the original inference process. Could you please clarify what exactly the warm-up improves ?

    Also, you mentioned achieving 50 FPS (without saving) and 33 FPS (with saving) on a single A100 . could you please elaborate a bit more on how you reached 40 FPS+ in your setup?
    Are there any specific configurations (e.g., batch size, audio length, or environment optimizations) that make the biggest impact?

    Thanks again for your great work and for taking the time to share these insights 🙏

  4. ljz1998zzzz commented on Oct 29, 2025

    @ljz1998zzzz
    Author

    Warm-up is actually a pretty common technique applied in DNN. As we know, there is always longer time during the first round inference (because of compute unit loading and blabla~). So warm-up simply put this round ahead, and make the second or third round as the actual 1st round in our application.

    If you take apart each round and see the computation delay during them, you will find the 1st the longest. And possibly, the last is also really long. That's because the last batch-size changes during inference due to mismatch between audio duration and batch size. Further, there is some compute unit reload issues when batch-size changes, which cause additional delay during inference.

    So, leave warm-up alone, [round(audio time length * instructed fps) / batch size = N (N∈Z)] would be the key to increase the overall inference fps. It guarantees the constant batch-size during inference and avoids the reload time for some units/matrices.

    Hope this helpful.

  5. EsraaK-Fouad commented on Oct 29, 2025

    @EsraaK-Fouad

    Thanks for the clarification! I applied the warm-up before inference as shown in the attached image, but I didn’t notice any improvement in the inference time when testing with a 7-second audio clip.
    Image
    Could you please check if I might have missed something?

  6. ljz1998zzzz commented on Oct 29, 2025

    @ljz1998zzzz
    Author

    Usually we put warm-up module in front of the whole application to ensure the real-time performance of our app. In this case, you may call warm-up module before GPU computation is called if you wanna actually see the improvement.

    Here is my testing code for reference.

    # Load model weights
    398     vae, unet, pe = load_all_model(
    399         unet_model_path=args.unet_model_path,
    400         vae_type=args.vae_type,
    401         unet_config=args.unet_config,
    402         device=device
    403     )
    404     timesteps = torch.tensor([0], device=device)
    405
    406     pe = pe.half().to(device)
    407     vae.vae = vae.vae.half().to(device)
    408     unet.model = unet.model.half().to(device)
    409
    410     # Initialize audio processor and Whisper model
    411     audio_processor = AudioProcessor(feature_extractor_path=args.whisper_dir)
    412     weight_dtype = unet.model.dtype
    413     whisper = WhisperModel.from_pretrained(args.whisper_dir)
    414     whisper = whisper.to(device=device, dtype=weight_dtype).eval()
    415     whisper.requires_grad_(False)
    416
    417     # Initialize face parser with configurable parameters based on version
    418     if args.version == "v15":
    419         fp = FaceParsing(
    420             left_cheek_width=args.left_cheek_width,
    421             right_cheek_width=args.right_cheek_width
    422         )
    423     else:  # v1
    424         fp = FaceParsing()
    425
    426     warm_up()
    427     inference_config = OmegaConf.load(args.inference_config)
    428     print(inference_config)
    429
    430     for avatar_id in inference_config:
    431         data_preparation = inference_config[avatar_id]["preparation"]
    432         video_path = inference_config[avatar_id]["video_path"]
    433         if args.version == "v15":
    434             bbox_shift = 0
    435         else:
    436             bbox_shift = inference_config[avatar_id]["bbox_shift"]
    437         avatar = Avatar(
    438             avatar_id=avatar_id,
    439             video_path=video_path,
    440             bbox_shift=bbox_shift,
    441             batch_size=args.batch_size,
    442             preparation=data_preparation)
    443
    444         audio_clips = inference_config[avatar_id]["audio_clips"]
    445         for audio_num, audio_path in audio_clips.items():
    446             print("Inferring using:", audio_path)
    447             avatar.inference(audio_path,
    448                            audio_num,
    449                            args.fps,
    450                            args.skip_save_images)
    
    
  7. EsraaK-Fouad commented on Oct 29, 2025

    @EsraaK-Fouad

    Thanks for the detailed explanation and for sharing your testing code!
    I moved the warm_up() call to execute right after loading all model weights and components (UNet, VAE, Whisper, etc.) onto the GPU similar to your example. I’ll run another test to check its impact on real-time performance.
    Currently, the inference time is 19 ms with a batch size of 8 and FPS set to 21.

    Image
  8. jackrx259 commented on Jul 8, 2026

    @jackrx259

    这个我在4090上做到了190fps,用蒸馏的vae decoder就行,另外还有些前后处理也优化下就好了

  9. 2996637314 commented on Jul 11, 2026

    @2996637314

    这个我在4090上做到了190fps,用蒸馏的vae decoder就行,另外还有些前后处理也优化下就好了

    看起来你牺牲了一些视觉效果 我试过蒸馏decoder显然这会让效果大幅下降 如果您有更好的方案或者一些具体的实现步骤 您可以给我留言 这对我有很大的帮助 谢谢

  10. EsraaK-Fouad commented on Jul 12, 2026

    @EsraaK-Fouad

    @jackrx259
    Thanks for sharing! If you've found a way to maintain the quality while achieving around 190 FPS, I'd love to hear more about your approach. If you could share some implementation details or optimization tips, I'd really appreciate it. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions