Skip to content

[Feature]: Migration from Model Runner v1 to Model Runner v2 #41286

Description

@yewentao256

🚀 The feature, motivation and pitch

We are going to migrate from model runner v1 to model runner v2 gradually, here is the roadmap:

  • Start with dense model, namely "Qwen/Qwen3-0.6B" and "facebook/opt-125m" as they covered most of the CI tests
  • Then, with moe model, like "deepseek-ai/DeepSeek-V2-lite"
  • Finally, test with popular model, like "deepseek-ai/DeepSeek-V4-Pro"

Tasks:

Remaining TODOs:

Alternatives

No response

Additional context

No response

Before submitting a new issue...

  • Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

Activity

  1. self-assigned this
    on Apr 29, 2026
  2. SouthWest7 commented on May 14, 2026

    @SouthWest7
    Contributor

    @yewentao256 Is there anything I can help with? I’d be happy to help the migration with MoE model. Are there any things I should pay attention to, or any suggestions you have?

  3. yewentao256 commented on May 14, 2026

    @yewentao256
    MemberAuthor

    @yewentao256 Is there anything I can help with? I’d be happy to help the migration with MoE model. Are there any things I should pay attention to, or any suggestions you have?

    Thanks for the interests! We will make very quick iterations through this and issues bump up with CI, so it may require fast responses which might be hard for casual conrtibutors. If you do want to contribute, please take a look at CIs of migration PRs and see if there are something you can help with.

  4. gcanlin commented on May 18, 2026

    @gcanlin
    Contributor

    One PR for this issue #42963 :) And I'm happy to spend time helping fix v2-related issues. Feel free to assign me if there're more CI failure about it. Thanks!

  5. yewentao256 commented on May 18, 2026

    @yewentao256
    MemberAuthor

    Thanks @gcanlin ! Do you have a L1 machine so that you may be able to reproduce this OOM issue? https://buildkite.com/vllm/ci/builds/66440#019e2c86-4e73-4692-92a2-c1b0970e79df

  6. gcanlin commented on May 18, 2026

    @gcanlin
    Contributor

    Thanks @gcanlin ! Do you have a L1 machine so that you may be able to reproduce this OOM issue? https://buildkite.com/vllm/ci/builds/66440#019e2c86-4e73-4692-92a2-c1b0970e79df

    Sorry. May I ask what does L1 machine mean? I have A100 GPUs, not sure whether it's enough to run the test.

  7. yewentao256 commented on May 18, 2026

    @yewentao256
    MemberAuthor

    @gcanlin Ohh sorry, a typo with that, I meant L4

    <html>
    <body>
    <!--StartFragment-->
    [2026-05-15T17:28:32Z] +-----------------------------------------------------------------------------------------+
    --
      | [2026-05-15T17:28:32Z] \| NVIDIA-SMI 570.133.20             Driver Version: 570.133.20     CUDA Version: 13.0     \|
      | [2026-05-15T17:28:32Z] \|-----------------------------------------+------------------------+----------------------+
      | [2026-05-15T17:28:32Z] \| GPU  Name                 Persistence-M \| Bus-Id          Disp.A \| Volatile Uncorr. ECC \|
      | [2026-05-15T17:28:32Z] \| Fan  Temp   Perf          Pwr:Usage/Cap \|           Memory-Usage \| GPU-Util  Compute M. \|
      | [2026-05-15T17:28:32Z] \|                                         \|                        \|               MIG M. \|
      | [2026-05-15T17:28:32Z] \|=========================================+========================+======================\|
      | [2026-05-15T17:28:32Z] \|   0  NVIDIA L4                      Off \|   00000000:35:00.0 Off \|                    0 \|
      | [2026-05-15T17:28:32Z] \| N/A   37C    P0             30W /   72W \|       0MiB /  23034MiB \|      4%      Default \|
      | [2026-05-15T17:28:32Z] \|                                         \|                        \|                  N/A \|
    
    <!--EndFragment-->
    </body>
    </html>
  8. gcanlin commented on May 18, 2026

    @gcanlin
    Contributor

    @gcanlin Ohh sorry, a typo with that, I meant L4

    Oh, sorry. I'm afraid that I can't access to L4 recently :(

  9. tobby168 commented on Jul 23, 2026

    @tobby168

    A performance data point on V2 vs V1, offered here since this issue tracks the road to making V2 the default. Not a bug report — V2 runs correctly and MTP acceptance is identical; it's just measurably slower than V1 for one class of workload on an unusual platform, which seemed worth recording while V2 matures toward the "switch by default" gate.

    Platform (deliberately atypical): NVIDIA GB10 / DGX Spark, sm_121a, aarch64, 128 GB unified LPDDR5X (~273 GB/s). Single-GPU, no TP/PP. This is a strongly bandwidth-bound regime — very different from the HBM datacenter parts V2 is usually measured on — and the workload is spec-decode heavy, so it stresses the parts of the runner (drafting/verify loop, per-step host overhead) rather differently.

    Model / config (identical across both runs, only the runner env var changes):

    • Qwen3.5-122B-A10B, INT4 (AutoRound) weights — MoE, ~10B active
    • Speculative decoding: MTP, num_speculative_tokens=2
    • Attention backend: FLASH_ATTN
    • compile_sizes=[3], greedy decode, fixed prompt, single stream (bs=1)
    • vLLM 0.23.1rc1.dev1302+ge765bbc97 (nightly)
    • This arch (Qwen3_5MoeForConditionalGeneration) is not in DEFAULT_V2_MODEL_RUNNER_ARCHITECTURES and is MoE, so it defaults to V1; V2 was opted in explicitly with VLLM_USE_V2_MODEL_RUNNER=1.

    Result (single-stream decode, steady state):

    Runner Decode throughput MTP acceptance len
    V1 (VLLM_USE_V2_MODEL_RUNNER=0) 59.6 tok/s 1.77–1.81
    V2 (VLLM_USE_V2_MODEL_RUNNER=1) 54.5 tok/s 1.77–1.81
    Δ −8.5% unchanged

    Methodology (3 lines):

    1. Same process, same flags, same fixed prompt; only VLLM_USE_V2_MODEL_RUNNER toggled between runs.
    2. Warm-up run discarded; reported numbers are the mean of runs 2+ (steady state, post CUDA-graph capture / compile).
    3. MTP acceptance length is identical across runners (1.77–1.81), so the gap is not from worse draft acceptance — it looks like per-step runtime/host overhead on a bandwidth-bound part, not a spec-decode quality change.

    Because acceptance is unchanged, the −8.5% appears to be per-step overhead that a fast HBM part would hide but a ~273 GB/s unified-memory part exposes. Happy to share more detail or narrow it down.

    Offer: we have this GB10 / DGX Spark available and are glad to run a torch-profiler trace on both runners (V1 vs V2, same prompt) if that would help localize where V2 spends the extra time on a bandwidth-bound target — just let us know what config/annotations you'd want.

    Minor related note: with VLLM_USE_V2_MODEL_RUNNER=1 the engine warns that V2 does not yet support the thinking_token_budget request parameter. This already appears to be tracked (Reasoning Budget in #47172, PR #46727), so just noting it for completeness rather than as a separate ask.

    Also cross-linking #40755 as a related but distinct V2 perf observation (that one is the _gumbel_sample_kernel on H800; ours is greedy so sampling isn't the path, but both are "V2 slower on specific hardware" data points).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions