-
Notifications
You must be signed in to change notification settings - Fork 54
Pull requests: mudler/vllm.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
tenstorrent: decode IQ3_XXS matmul weights on the int8-dot kernel (QUANT-GGUF-IQ-TENSTORRENT wave 1)
#3193
opened Sep 14, 2026 by
lu-zero
Collaborator
Loading…
feat(KERNEL-QUANT-CIQ-GEMM-ROCM-RDNA3): enable gfx1100 WMMA prefill
#3187
opened Sep 14, 2026 by
VikashLoomba
Contributor
Loading…
docs: correct Qwen3.8-Flash-Next generation claims
#3186
opened Sep 14, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs(QUANT-EXL3): align support and usage with code
#3176
opened Sep 13, 2026 by
localai-org-maint-bot
Collaborator
•
Draft
11 of 12 tasks
feat(MODEL-MM-GLM53-FLASH): W9c-3 device compose forward in glm5_next_device.cpp
#3175
opened Sep 13, 2026 by
localai-org-maint-bot
Collaborator
Loading…
feat(MODEL-MM-deepseek-v4-deepseek-v4-for-causal-lm): run the DeepSeek-V4-Flash-Vision tower on CUDA against llama.cpp b10766
#3163
opened Sep 12, 2026 by
localai-org-maint-bot
Collaborator
Loading…
record(ENG-RECORD-CONFLICT-SURFACES): file two records-tooling defects
#3152
opened Sep 12, 2026 by
localai-org-maint-bot
Collaborator
Loading…
fix(QUANT-GGUF-CPU-THREADPOOL): floor the oversubscription threshold at 100 us so 4-vCPU CI does not flake
#3140
opened Sep 11, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs(ENG-MM-INPUT-PIPELINE): correct HTTP support claims
#3118
opened Sep 10, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs: clarify Qwen3.8 benchmark limits
#3104
opened Sep 9, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs(KV-WARMUP-PROFILE): specify startup memory profiling
#3050
opened Sep 8, 2026 by
VikashLoomba
Contributor
•
Draft
perf(KERNEL-QUANT-CIQ-GEMM-ROCM): cooperative-tile WMMA kernels — Shared rejected, BigTile accepted
#3036
opened Sep 7, 2026 by
joral
Contributor
Loading…
feat(KERNEL-QUANT-CIQ-GEMM-ROCM-IQUANT): port IQ4_XS/IQ3_XXS to ROCm
#3029
opened Sep 6, 2026 by
joral
Contributor
Loading…
fix(MODEL-TEXT-GLM4-MOE-LITE-GATE-2839): the near-tie predicate cannot fire, so apply the bar the oracle capture licenses -- and it FAILS 69/128
#2906
opened Sep 4, 2026 by
localai-org-maint-bot
Collaborator
Loading…
perf(GFX1100-TG200): T14 row-split greedy argmax arm
#2876
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T9 cooperative gated norm arm
#2875
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T6b cooperative attn preamble arm
#2868
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T6a cooperative GDN scan arm
#2866
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
feat(GFX1100-TG200): T25 keep ssm_out as Q5_K with runtime input permutation
#2807
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN projections
#2804
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T16 YTILE=4 default for wvSplitK decode-skinny GEMV
#2787
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T2b flips ROCm support_static_graph_mode
#2777
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
feat(MODEL-MM-GLM53-FLASH-KPOOL-CUDA): give GLM-5.3-Flash's k-pool indexer a device it can run on
#2432
opened Aug 31, 2026 by
localai-org-maint-bot
Collaborator
Loading…
feat(BACKEND-VULKAN-TQ1_0): TQ1_0 ternary keep-quant matmul, MoE, and rope shaders for Vulkan
#2248
opened Aug 29, 2026 by
phantomic12
Loading…
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.