Repository navigation
To achieve higher fps #377
Description
Activity
Here are some detailed experiment results:
audio length sun.wav : 22 s 488 ms
when I set 16 fps and batch size 8, inference fps =359/22.88=15.7, notice there is a long delay at the beginning of the last batch. this should be the problem. ps: the selected area stands for the time between two consecutive entries into the iteration in process_frames function.
then I set 21 fps and batch size 8, inference fps =472/17.39=27.1, notice there is no long delay at the beginning of the last batch.
Hope this is helpful.
Further exploring this issue, there is also a long delay at the beginning of the first batch inference. So I add a warm up module. Now we get 50 fps without saving images and 33 fps with saving images on single A100.
yet when I do this, the audio feature extraction part took longer time (3.4s than 1.2s), which needs further research.
here is warm up module welcome to use:
`@torch.no_grad() def warm_up(): warmup_audio_path = 'your/path/to/audio_demo/sun.wav' latent_batch = torch.rand(args.batch_size,8,32,32,device=device,dtype=unet.model.dtype) audio_feature_batch = torch.rand(args.batch_size,50,384,device=device,dtype=unet.model.dtype) ############################################## warmup audio feature ############################################## start_time = time.time() # Extract audio features whisper_input_features, librosa_length = audio_processor.get_audio_feature(warmup_audio_path, weight_dtype=weight_dtype) whisper_chunks = audio_processor.get_whisper_chunk( whisper_input_features, device, weight_dtype, whisper, librosa_length, fps=args.fps, audio_padding_length_left=args.audio_padding_length_left, audio_padding_length_right=args.audio_padding_length_right, ) audio_time = time.time() print(f"audio feature warmup: costs {(audio_time - start_time) * 1000}ms") ############################################## warmup latents ############################################## pred_latents = unet.model(latent_batch, timesteps, encoder_hidden_states=audio_feature_batch).sample pred_latents = pred_latents.to(device=device, dtype=vae.vae.dtype) recon = vae.decode_latents(pred_latents) latent_time = time.time() print(f"latents warmup: costs {(latent_time - audio_time) * 1000}ms")`thanks a lot for sharing this detailed explanation and the warm-up module, it’s really helpful!
I tried the warm-up technique you mentioned, but I didn’t notice much difference compared to the original inference process. Could you please clarify what exactly the warm-up improves ?
Also, you mentioned achieving 50 FPS (without saving) and 33 FPS (with saving) on a single A100 . could you please elaborate a bit more on how you reached 40 FPS+ in your setup?
Are there any specific configurations (e.g., batch size, audio length, or environment optimizations) that make the biggest impact?Thanks again for your great work and for taking the time to share these insights 🙏
Warm-up is actually a pretty common technique applied in DNN. As we know, there is always longer time during the first round inference (because of compute unit loading and blabla~). So warm-up simply put this round ahead, and make the second or third round as the actual 1st round in our application.
If you take apart each round and see the computation delay during them, you will find the 1st the longest. And possibly, the last is also really long. That's because the last batch-size changes during inference due to mismatch between audio duration and batch size. Further, there is some compute unit reload issues when batch-size changes, which cause additional delay during inference.
So, leave warm-up alone, [round(audio time length * instructed fps) / batch size = N (N∈Z)] would be the key to increase the overall inference fps. It guarantees the constant batch-size during inference and avoids the reload time for some units/matrices.
Hope this helpful.
Usually we put warm-up module in front of the whole application to ensure the real-time performance of our app. In this case, you may call warm-up module before GPU computation is called if you wanna actually see the improvement.
Here is my testing code for reference.
# Load model weights 398 vae, unet, pe = load_all_model( 399 unet_model_path=args.unet_model_path, 400 vae_type=args.vae_type, 401 unet_config=args.unet_config, 402 device=device 403 ) 404 timesteps = torch.tensor([0], device=device) 405 406 pe = pe.half().to(device) 407 vae.vae = vae.vae.half().to(device) 408 unet.model = unet.model.half().to(device) 409 410 # Initialize audio processor and Whisper model 411 audio_processor = AudioProcessor(feature_extractor_path=args.whisper_dir) 412 weight_dtype = unet.model.dtype 413 whisper = WhisperModel.from_pretrained(args.whisper_dir) 414 whisper = whisper.to(device=device, dtype=weight_dtype).eval() 415 whisper.requires_grad_(False) 416 417 # Initialize face parser with configurable parameters based on version 418 if args.version == "v15": 419 fp = FaceParsing( 420 left_cheek_width=args.left_cheek_width, 421 right_cheek_width=args.right_cheek_width 422 ) 423 else: # v1 424 fp = FaceParsing() 425 426 warm_up() 427 inference_config = OmegaConf.load(args.inference_config) 428 print(inference_config) 429 430 for avatar_id in inference_config: 431 data_preparation = inference_config[avatar_id]["preparation"] 432 video_path = inference_config[avatar_id]["video_path"] 433 if args.version == "v15": 434 bbox_shift = 0 435 else: 436 bbox_shift = inference_config[avatar_id]["bbox_shift"] 437 avatar = Avatar( 438 avatar_id=avatar_id, 439 video_path=video_path, 440 bbox_shift=bbox_shift, 441 batch_size=args.batch_size, 442 preparation=data_preparation) 443 444 audio_clips = inference_config[avatar_id]["audio_clips"] 445 for audio_num, audio_path in audio_clips.items(): 446 print("Inferring using:", audio_path) 447 avatar.inference(audio_path, 448 audio_num, 449 args.fps, 450 args.skip_save_images)Reacted by Esraa Khaled AhmedThanks for the detailed explanation and for sharing your testing code!
I moved the warm_up() call to execute right after loading all model weights and components (UNet, VAE, Whisper, etc.) onto the GPU similar to your example. I’ll run another test to check its impact on real-time performance.
Currently, the inference time is 19 ms with a batch size of 8 and FPS set to 21.
这个我在4090上做到了190fps,用蒸馏的vae decoder就行,另外还有些前后处理也优化下就好了
这个我在4090上做到了190fps,用蒸馏的vae decoder就行,另外还有些前后处理也优化下就好了
看起来你牺牲了一些视觉效果 我试过蒸馏decoder显然这会让效果大幅下降 如果您有更好的方案或者一些具体的实现步骤 您可以给我留言 这对我有很大的帮助 谢谢
@jackrx259
Thanks for sharing! If you've found a way to maintain the quality while achieving around 190 FPS, I'd love to hear more about your approach. If you could share some implementation details or optimization tips, I'd really appreciate it. Thanks!

Hi there, I noticed there are lots of issues about failing to achieve the same fps as stated in this project ( 30 fps on v100 ).
Here is one tip:
users should ensure that [round(audio time length * instructed fps) / batch size = N (N∈Z)]
following this, the inference performance can be increased significantly