-
Notifications
You must be signed in to change notification settings - Fork 2.7k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[None][test] Add Qwen 3.8 MAX and Flash-next performance coverage
#19012
opened Sep 10, 2026 by
yufeiwu-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6676352][doc] Fix two copy-paste-broken doc snippets
#19009
opened Sep 10, 2026 by
BowenFu
Contributor
Loading…
1 task done
[None][refactor] Centralize FMHA availability and support capability checks
#19008
opened Sep 10, 2026 by
yuxianq
Collaborator
Loading…
1 task done
[https://nvbugs/6731971][doc] Fix broken relative paths flagged by test_relative_path_validity
#19007
opened Sep 10, 2026 by
nv-guomingz
Collaborator
Loading…
2 tasks done
[None][fix] Jointly manage target and draft KV caches
#19006
opened Sep 10, 2026 by
liji-nv
Collaborator
Loading…
1 task done
[https://nvbugs/6727262][fix] Bound default Qwen hybrid state memory
#19005
opened Sep 10, 2026 by
VALLIS-NERIA
Collaborator
Loading…
1 task done
[None][feat] Enable KVCacheManagerV2 by default for Llama and Llama4
#19004
opened Sep 10, 2026 by
yizhang-nv
Member
Loading…
1 task done
[None][feat] Kimi K3: unlock the CUTEDSL MoE backend for NVFP4 SiTU
#19003
opened Sep 10, 2026 by
xguannv
Contributor
Loading…
1 task done
[visual_gen] Optimize Wan2.2 Dataflow (Zhen Xie from VibeHPC)
VisualGen
#19002
opened Sep 10, 2026 by
zhen-xie
Loading…
[TRTLLM-15097][test] Prune MiniMax-M2 and MiniMax-M2.5 tests
ci: full pre-merge approved
#19001
opened Sep 10, 2026 by
xinhe-nv
Collaborator
Loading…
1 task done
[None][feat] Track accepted native changes in perf optimize
#19000
opened Sep 10, 2026 by
GuanhuaWang2001
Collaborator
Loading…
[None][perf] Triton top-k combine for Marlin NVFP4 MoE
#18999
opened Sep 10, 2026 by
rmeghwal-nv
Contributor
Loading…
[TRTLLM-15621][perf] Skip generation CUDA graph capture on a disagg context worker
api-compatible
Accepted LLM API contract change that is backwards-compatible
#18997
opened Sep 10, 2026 by
zhaoyangwang-nvidia
Collaborator
Loading…
1 task done
[None][fix] Update vendored PrimTS with FlashInfer #4829 follow-ups
ci: full pre-merge approved
#18996
opened Sep 10, 2026 by
yuxianq
Collaborator
Loading…
1 task done
[None][perf] Helix post-process: stop writing partial_o twice on both alltoall backends
#18995
opened Sep 10, 2026 by
xguannv
Contributor
Loading…
1 task done
[TRTLLM-16179][test] Port detokenization stop-word tests to Qwen3-0.6B
#18994
opened Sep 10, 2026 by
xinhe-nv
Collaborator
Loading…
1 task done
[#18971][fix] Fail the request when a worker has no multimodal encoder
#18993
opened Sep 10, 2026 by
dibyo10
Loading…
1 task done
[None][test] Cover the V1 hybrid manager in the vocab_size forwarding tests
#18991
opened Sep 10, 2026 by
dibyo10
Loading…
1 task done
[None][feat] perf-sanity: upload per-request disagg lifecycle spans to OpenSearch
#18990
opened Sep 10, 2026 by
chenfeiz0326
Collaborator
Loading…
3 tasks done
[None][fix] Drop every media payload on a generation_only disagg request, not just multi_modal_data
#18989
opened Sep 10, 2026 by
eopXD
Collaborator
Loading…
1 task done
[https://nvbugs/6732123][fix] Correct V2 KV cache quota estimation
#18988
opened Sep 10, 2026 by
yizhang-nv
Member
Loading…
3 tasks done
[None][perf] Introduce incremental HCA compressor summaries
#18987
opened Sep 10, 2026 by
mingyangHao
Collaborator
Loading…
1 task done
[https://nvbugs/6727262][fix] Pin a feasible
max_batch_size default for the family via the existing…
#18985
opened Sep 10, 2026 by
trtllm-agent
Collaborator
Loading…
2 tasks done
[None][fix] Size seq-slot pool to cover disagg-gen KV admission to mitigate hangs
#18983
opened Sep 9, 2026 by
brb-nv
Collaborator
Loading…
1 task done
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-10.