Repository navigation
[Feature]: Migration from Model Runner v1 to Model Runner v2 #41286
Description
Activity
@yewentao256 Is there anything I can help with? I’d be happy to help the migration with MoE model. Are there any things I should pay attention to, or any suggestions you have?
@yewentao256 Is there anything I can help with? I’d be happy to help the migration with MoE model. Are there any things I should pay attention to, or any suggestions you have?
Thanks for the interests! We will make very quick iterations through this and issues bump up with CI, so it may require fast responses which might be hard for casual conrtibutors. If you do want to contribute, please take a look at CIs of migration PRs and see if there are something you can help with.
One PR for this issue #42963 :) And I'm happy to spend time helping fix v2-related issues. Feel free to assign me if there're more CI failure about it. Thanks!
Reacted by Wentao YeThanks @gcanlin ! Do you have a L1 machine so that you may be able to reproduce this OOM issue? https://buildkite.com/vllm/ci/builds/66440#019e2c86-4e73-4692-92a2-c1b0970e79df
Reacted by Canlin GuoThanks @gcanlin ! Do you have a L1 machine so that you may be able to reproduce this OOM issue? https://buildkite.com/vllm/ci/builds/66440#019e2c86-4e73-4692-92a2-c1b0970e79df
Sorry. May I ask what does
L1 machinemean? I have A100 GPUs, not sure whether it's enough to run the test.@gcanlin Ohh sorry, a typo with that, I meant L4
<html> <body> <!--StartFragment--> [2026-05-15T17:28:32Z] +-----------------------------------------------------------------------------------------+ -- | [2026-05-15T17:28:32Z] \| NVIDIA-SMI 570.133.20 Driver Version: 570.133.20 CUDA Version: 13.0 \| | [2026-05-15T17:28:32Z] \|-----------------------------------------+------------------------+----------------------+ | [2026-05-15T17:28:32Z] \| GPU Name Persistence-M \| Bus-Id Disp.A \| Volatile Uncorr. ECC \| | [2026-05-15T17:28:32Z] \| Fan Temp Perf Pwr:Usage/Cap \| Memory-Usage \| GPU-Util Compute M. \| | [2026-05-15T17:28:32Z] \| \| \| MIG M. \| | [2026-05-15T17:28:32Z] \|=========================================+========================+======================\| | [2026-05-15T17:28:32Z] \| 0 NVIDIA L4 Off \| 00000000:35:00.0 Off \| 0 \| | [2026-05-15T17:28:32Z] \| N/A 37C P0 30W / 72W \| 0MiB / 23034MiB \| 4% Default \| | [2026-05-15T17:28:32Z] \| \| \| N/A \| <!--EndFragment--> </body> </html>
@gcanlin Ohh sorry, a typo with that, I meant L4
Oh, sorry. I'm afraid that I can't access to L4 recently :(
Reacted by Wentao YeA performance data point on V2 vs V1, offered here since this issue tracks the road to making V2 the default. Not a bug report — V2 runs correctly and MTP acceptance is identical; it's just measurably slower than V1 for one class of workload on an unusual platform, which seemed worth recording while V2 matures toward the "switch by default" gate.
Platform (deliberately atypical): NVIDIA GB10 / DGX Spark,
sm_121a, aarch64, 128 GB unified LPDDR5X (~273 GB/s). Single-GPU, no TP/PP. This is a strongly bandwidth-bound regime — very different from the HBM datacenter parts V2 is usually measured on — and the workload is spec-decode heavy, so it stresses the parts of the runner (drafting/verify loop, per-step host overhead) rather differently.Model / config (identical across both runs, only the runner env var changes):
Qwen3.5-122B-A10B, INT4 (AutoRound) weights — MoE, ~10B active- Speculative decoding: MTP, num_speculative_tokens=2
- Attention backend: FLASH_ATTN
compile_sizes=[3], greedy decode, fixed prompt, single stream (bs=1)- vLLM
0.23.1rc1.dev1302+ge765bbc97(nightly) - This arch (
Qwen3_5MoeForConditionalGeneration) is not inDEFAULT_V2_MODEL_RUNNER_ARCHITECTURESand is MoE, so it defaults to V1; V2 was opted in explicitly withVLLM_USE_V2_MODEL_RUNNER=1.
Result (single-stream decode, steady state):
Runner Decode throughput MTP acceptance len V1 ( VLLM_USE_V2_MODEL_RUNNER=0)59.6 tok/s 1.77–1.81 V2 ( VLLM_USE_V2_MODEL_RUNNER=1)54.5 tok/s 1.77–1.81 Δ −8.5% unchanged Methodology (3 lines):
- Same process, same flags, same fixed prompt; only
VLLM_USE_V2_MODEL_RUNNERtoggled between runs. - Warm-up run discarded; reported numbers are the mean of runs 2+ (steady state, post CUDA-graph capture / compile).
- MTP acceptance length is identical across runners (1.77–1.81), so the gap is not from worse draft acceptance — it looks like per-step runtime/host overhead on a bandwidth-bound part, not a spec-decode quality change.
Because acceptance is unchanged, the −8.5% appears to be per-step overhead that a fast HBM part would hide but a ~273 GB/s unified-memory part exposes. Happy to share more detail or narrow it down.
Offer: we have this GB10 / DGX Spark available and are glad to run a torch-profiler trace on both runners (V1 vs V2, same prompt) if that would help localize where V2 spends the extra time on a bandwidth-bound target — just let us know what config/annotations you'd want.
Minor related note: with
VLLM_USE_V2_MODEL_RUNNER=1the engine warns that V2 does not yet support thethinking_token_budgetrequest parameter. This already appears to be tracked (Reasoning Budget in #47172, PR #46727), so just noting it for completeness rather than as a separate ask.Also cross-linking #40755 as a related but distinct V2 perf observation (that one is the
_gumbel_sample_kernelon H800; ours is greedy so sampling isn't the path, but both are "V2 slower on specific hardware" data points).Reacted by Wentao Ye and David
🚀 The feature, motivation and pitch
We are going to migrate from model runner v1 to model runner v2 gradually, here is the roadmap:
Tasks:
logprob_token_idssupport #40559 @yewentao256num_gpu_runner_capture_triggersandnum_cudagraph_captured#41285 @yewentao256pre_forwardorder #42676 @yewentao256Sizes of tensors must matcherror #42778 @yewentao256Triton Error [CUDA]: device-side assert triggered#43139 @yewentao256AttributeError: 'CohereASRDecoder' object has no attribute 'embed_input_ids'#44568 @yewentao256openai.InternalServerError: Error code: 500 - 'list index out of range'#45467 @yewentao256WhisperModelState#46096 @njhillRemaining TODOs:
Alternatives
No response
Additional context
No response
Before submitting a new issue...