Skip to content

[deps] Migrate custom vLLM-router fork on top of 0.1.15 release of vllm-router #2070

Description

@SumanthRH

Summary

We currently use a custom vllm router fork: vllm-project/router@main...SumanthRH:router:sticky_least_loaded with SkyRL

The fork has two fixes on top of 0.1.14:

  1. Allowing extra args on top of chat completions: [fix] Preserve unknown fields in ChatCompletionRequest vllm-project/router#162
  2. Custom load-aware sticky routing policy sticky_least_loaded

We should rebase the fork on top of 0.1.15 after the recent release and upgrade SkyRL to use the new wheels.

We should also try to upstream the routing policy to vllm-router.

Activity

  1. self-assigned this
    on Aug 19, 2026
  2. bvolpato commented on Sep 3, 2026

    @bvolpato
    Contributor

    Opened vllm-project/router#235 as a draft for the missing sticky_least_loaded policy, adapted from the public fork implementation. It assigns new sessions by active-session count, preserves affinity across turns, and adds explicit release plus idle expiry. The draft documents process-local state and model-scoped session IDs when models share a policy.

    This does not complete the migration by itself. An official router release still needs the typed generation paths in router#199, consolidated request accounting and response-lifetime handling from router#200/router#216 (including the stalled-stream and decode-body cases), and this session policy. Removing downstream HTTP cancellation also depends on SkyRL#2121.

    The new draft passes local CPU/mock-worker tests. It does not establish tensor-parallel backend abort latency or GPU behavior.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions