-
-
Notifications
You must be signed in to change notification settings - Fork 21.2k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[XPU] Add env overrides for DiffKV attention 2D launch and tile size
intel-gpu
Related to Intel GPU
#53804
opened Aug 25, 2026 by
slokesha
Loading…
4 tasks done
Fix mamba align retention checkpoints
kv-cache-manager
mrv2
Model Runner V2 specific
#53803
opened Aug 25, 2026 by
ptorsten
Loading…
[Bugfix] Align hybrid prefix-cache hit boundaries
bug
Something isn't working
kv-cache-manager
scheduler
#53802
opened Aug 25, 2026 by
ptorsten
Loading…
[Bugfix] Close the socket test_loopback_bind creates when bind fails
bug
Something isn't working
#53801
opened Aug 25, 2026 by
eyad6789
Loading…
fix MI455 customed paged attn
rocm
Related to AMD ROCm
#53800
opened Aug 25, 2026 by
Concurrensee
Contributor
•
Draft
[Bugfix][Scheduler] Handle missing req_id in update_from_output gracefully
bug
Something isn't working
scheduler
#53799
opened Aug 25, 2026 by
dhakshin32
Loading…
[Bugfix] Seed align-mode mamba state_idx in mamba block units
bug
Something isn't working
mrv2
Model Runner V2 specific
#53798
opened Aug 25, 2026 by
ptorsten
Loading…
Add support for loading dflash2 model in speculators format
dflash
#53797
opened Aug 25, 2026 by
fynnsu
Contributor
Loading…
3 of 4 tasks
[Core][Frontend] Add per-request prefix cache telemetry
documentation
Improvements or additions to documentation
frontend
kv-cache-manager
scheduler
#53795
opened Aug 25, 2026 by
cook1e-0707
Loading…
6 tasks done
Fuse ReLU2 with static FP8 activation quantization
quantization
#53793
opened Aug 25, 2026 by
samnordmann
Contributor
•
Draft
[ROCm] Resolve the indexer fp8 cache dtype once at import
rocm
Related to AMD ROCm
#53792
opened Aug 25, 2026 by
amd-sriram
Loading…
[RFC] Fuse Nemotron shared ReLU2 and static FP8 quantization
quantization
#53791
opened Aug 25, 2026 by
samnordmann
Contributor
•
Draft
[Bugfix] NemotronHMTP: add hf_to_vllm_mapper so quant exclusions reach the MTP draft
bug
Something isn't working
quantization
#53790
opened Aug 25, 2026 by
juhi10071998
Contributor
Loading…
[ROCm][Perf] Bound the decode paged-MQA-logits sanitize to the top-k read region
rocm
Related to AMD ROCm
#53789
opened Aug 25, 2026 by
amd-sriram
Loading…
[Do Not Merge][Attention] Enable masked MHA for GLM-5 head dimensions
ci/build
glm
performance
Performance-related issues
#53785
opened Aug 25, 2026 by
MatthewBonanni
Member
Loading…
[Distributed] Support pre-shared ncclUniqueId rendezvous for weight transfer
#53784
opened Aug 25, 2026 by
dharak-cohere
•
Draft
[4/N] Expose HiSparse cache metrics
kv-connector
mrv2
Model Runner V2 specific
scheduler
#53782
opened Aug 25, 2026 by
MatthewBonanni
Member
•
4/6
Loading…
[3/N] HiSparse: host-resident sparse-MLA decode hot-buffering
ci/build
cpu
Related to CPU backends
deepseek
Related to DeepSeek models
documentation
Improvements or additions to documentation
DSv4
glm
kv-cache-manager
kv-connector
minimax
mrv2
Model Runner V2 specific
nvidia
ready
ONLY add when PR is ready to merge/full CI is needed
scheduler
v1
#53781
opened Aug 25, 2026 by
MatthewBonanni
Member
•
3/6
Loading…
[2/N][KV Connector][NIXL] Support per-region transfer geometry
kv-connector
ready
ONLY add when PR is ready to merge/full CI is needed
#53780
opened Aug 25, 2026 by
MatthewBonanni
Member
•
2/6
Loading…
[1/N][KV Connector] Identify externally transferable KV cache groups
kv-cache-manager
kv-connector
ready
ONLY add when PR is ready to merge/full CI is needed
#53779
opened Aug 25, 2026 by
MatthewBonanni
Member
•
1/6
Loading…
[Build] Restore generic wheel for no_device targets
ci/build
#53778
opened Aug 25, 2026 by
Silv3S
Loading…
3 of 4 tasks
[Bugfix][XPU] Fix async output deadlock from blocking cross-stream event
bug
Something isn't working
intel-gpu
Related to Intel GPU
#53776
opened Aug 25, 2026 by
yeonsily
Loading…
3 of 4 tasks
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.