Skip to content

Commit d3e0e3f

Browse files
functionstackxclaude
authored andcommitted
[Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 (SemiAnalysisAI#2875)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
1 parent b105a0b commit d3e0e3f

2 files changed

Lines changed: 9 additions & 1 deletion

File tree

‎configs/nvidia-master.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7529,7 +7529,7 @@ minimaxm3-fp8-h100-vllm-agentic-mtp:
75297529
- { tp: 8, spec-decoding: mtp, kv-offloading: dram, kv-offload-backend: { name: mooncake, version: "0.3.11.post1" }, conc-list: [6, 8] }
75307530

75317531
minimaxm3-fp8-h200-vllm-agentic-mtp:
7532-
image: vllm/vllm-openai:v0.27.1
7532+
image: vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
75337533
model: MiniMaxAI/MiniMax-M3-MXFP8
75347534
model-prefix: minimaxm3
75357535
runner: cluster:h200-dgxc

‎perf-changelog.yaml‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6930,3 +6930,11 @@
69306930
description:
69316931
- "Update SGLang image to lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 (2026-09-07 cu13 dev nightly, digest sha256:19b8fa1223cc339c1eae7a5b703f1a8c2543b5b119155bf3d7efaef18f77f007, tag commit sgl-project/sglang@30705c00; Docker Hub last pushed 2026-09-07T01:43:42Z) for both H200 Qwen3.5 FP8 SGLang AgentX MTP recipes: qwen3.5-fp8-h200-sglang-agentic-mtp from lmsysorg/sglang:v0.5.16-cu130 (v0.5.16 release) and qwen3.5-fp8-h200-sglang-agentic-hicache-mtp from lmsysorg/sglang:nightly-dev-cu13-20260815-a5ba081f (2026-08-15 nightly). Both route to benchmarks/single_node/agentic/qwen3.5_fp8_h200_mtp.sh, which is unchanged: SGLANG_ENABLE_SPEC_V2 EAGLE MTP at 3 steps, golden acceptance length 3.39, flashinfer attention with allreduce fusion, fp8 quantization and fp8_e4m3 KV, HiCache kernel IO / page_first layout. Concurrency grids unchanged. Same tag the B200 Qwen3.5 FP8/FP4 SGLang AgentX recipes moved to in #2861/#2862."
69326932
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2868
6933+
6934+
- config-keys:
6935+
- minimaxm3-fp8-h200-vllm-agentic-mtp
6936+
scenario-type:
6937+
- agentic-coding
6938+
description:
6939+
- "Update vLLM image from vllm/vllm-openai:v0.27.1 (v0.27.1 release) to vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 (2026-09-07 upstream nightly, digest sha256:254eebf919e8b7b0d530d97fccc606c36f380ff6724d64bc190951bec1aee838, tag commit vllm-project/vllm@d9105ea8; Docker Hub last pushed 2026-09-07T06:16:01Z), the same tag the B200 MiniMax-M3 AgentX recipe moved to in #2860 and the ROCm counterpart of the MI325X/MI300X bumps in #2872/#2873. benchmarks/single_node/agentic/minimaxm3_fp8_h200_mtp.sh is unchanged: TRITON_ATTN attention, fp8 KV, EAGLE3 with the Inferact MiniMax-M3 EAGLE3-GQA draft pinned to FLASH_ATTN and the committed golden synthetic acceptance length 2.78, Mooncake 0.3.11.post1 DRAM offload on the host-tier arm; the resident TP8 c1/c2/c4/c6/c8/c10 and Mooncake DRAM offload c12/c14 grid is unchanged."
6940+
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2875

0 commit comments

Comments
 (0)