You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[deps] Migrate custom vLLM-router fork on top of 0.1.15 release of vllm-router #2070
Opened vllm-project/router#235 as a draft for the missing sticky_least_loaded policy, adapted from the public fork implementation. It assigns new sessions by active-session count, preserves affinity across turns, and adds explicit release plus idle expiry. The draft documents process-local state and model-scoped session IDs when models share a policy.
This does not complete the migration by itself. An official router release still needs the typed generation paths in router#199, consolidated request accounting and response-lifetime handling from router#200/router#216 (including the stalled-stream and decode-body cases), and this session policy. Removing downstream HTTP cancellation also depends on SkyRL#2121.
The new draft passes local CPU/mock-worker tests. It does not establish tensor-parallel backend abort latency or GPU behavior.
Summary
We currently use a custom vllm router fork: vllm-project/router@main...SumanthRH:router:sticky_least_loaded with SkyRL
The fork has two fixes on top of 0.1.14:
ChatCompletionRequestvllm-project/router#162We should rebase the fork on top of 0.1.15 after the recent release and upgrade SkyRL to use the new wheels.
We should also try to upstream the routing policy to
vllm-router.