Repository navigation
Add various optimizations and Mega MoE benchmarks - #316
Conversation
To clarify here, is this the number of tokens per rank as it's stated, or the number of tokens across the node? I think I'm just not being smart, but to clarify, across the entire node you would have 512 times 8 equals 4096 tokens in the second row? |
|
@kiankyars Yes, the batch size listed is the number of tokens per rank as stated. So for the row showing 512 tokens per rank under EP8, the total across the node would be 512 × 8 = 4,096. Your understanding is correct. |
|
thanks! |
…ions and Mega MoE benchmarks) into nv_dev Brings in deepseek-ai PR deepseek-ai#316 'Add various optimizations and Mega MoE benchmarks' on top of the nv_dev branch (currently at c491439, the merge of PR deepseek-ai#314). Conflicts resolved against the NV-side commits from PR deepseek-ai#314 (f6d98f2, 12cb7c0, a97b74d): * deep_gemm/include/deep_gemm/scheduler/paged_mqa_logits.cuh: keep NV's host-passed num_next_n_atoms runtime parameter, plus 891d57b's kIsVarlen template + binary-search SM work distribution + varlen path. * deep_gemm/include/deep_gemm/impls/sm90_fp8_paged_mqa_logits.cuh: keep NV's kNumKVMulticast template + cluster multicast logic; add 891d57b's kIsVarlen template parameter with static_assert(!kIsVarlen) (SM90 doesn't support varlen). * csrc/jit_kernels/impls/smxx_fp8_fp4_paged_mqa_logits.hpp: unified Args struct (num_next_n_atoms + is_varlen + indices); SM90 instantiation passes is_varlen=false; SM100 TMA-box rule uses 891d57b's (is_varlen or next_n >= 2) ? 2 : 1. * csrc/apis/attention.hpp: host computes num_next_n_atoms to match the post-deepseek-ai#316 SM100 kernel rule (kNextNAtom = (is_varlen or next_n >= 2) ? 2 : 1, kNumNextNAtoms = ceil_div(next_n, kNextNAtom)); SM90 multicast hard-codes num_next_n_atoms = 1. Without this fix, next_n=5 over-schedules by 5/3 and triggers IMA on SM100. * tests/test_attention.py: keep NV's block_kv in {32, 64} and next_n=4 for SM90 multicast; add 891d57b's outer is_varlen loop (varlen is SM100-only). Validated on B300 SXM6 (sm103) and H200 NVL (sm90): test_attention.py passes the full enumeration on both architectures with no IMA; numerical diff < 1e-3 (FP8) / 0.02 (FP4).
Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks)
|
@LyricZhao Hi, sorry to bother you. I 've a question about MegaMoE. The activation input is seems already placed in symmetric memory. Should we force fp4 quantization outputs to target symmetric memory? Do the model scripts need modification? |
…s) into nv_dev Brings deepseek-ai PR #347 (88965b0, also pulls 714dd1a 'Update test_mega_moe.py') onto nv_dev (currently at ac1f285 = #316 merge #328 + PR #342 IMA guard). Merge base is 891d57b. Conflicts resolved (5 files): * scheduler/paged_mqa_logits.cuh: take #347's refresh_num_kv_and_advance + reversed metadata allocation; drop nv_dev's PR #342 exist_q_atom_idx guard and get_atom_advance (both subsumed — reversed alloc keeps current_q_atom_idx in-bounds, refresh_num_kv_and_advance reproduces the varlen 1-or-2 advance). NV's host-passed num_next_n_atoms survives in the auto-merged metadata kernel (verified: single param, no inline recompute, coherent with reversed alloc). * csrc/apis/gemm.hpp: drop redundant C/D dtype asserts (#347 cleanup; the pre-existing d-dtype check + d==c check already cover them). * sm100_fp8_fp4_gemm_1d1d.cuh: take #347's unconditional BF16/FP32 C/D assert (supersedes nv_dev's accumulation-only assert). * tests/generators.py: take #347's broadened BF16-accumulation enumeration (out_dtype==bf16 and (dtype==bf16 or arch==10)) — strict superset of nv_dev's fp8+SM100 case. * tests/test_attention.py: merge both — keep NV's SM90 next_n=4 multicast + num_clusters metadata sizing; add #347's pool-size limits, batch_size=4096, and block_table/context_lens sanity asserts. Auto-merged metadata kernel + smxx_fp8_fp4_paged_mqa_logits.hpp + attention.hpp verified coherent across all three layers (kernel signature / launch / host call all carry num_next_n_atoms + is_varlen + indices). Validated on B-card (SM100) and H200 (SM90) incl next_n=4 multicast. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
Hi, I wanted to follow up on the test results — would you mind letting me know what physical device version and software build were used? Thank you! |
|
Hi, I have a question regarding this operator. |
The original PR is for blackwell platform with NVFP4 support, hence FP8 x FP4 mlp. The codes are guarded by "#if (defined(CUDA_ARCH) and (CUDA_ARCH >= 1000)) or defined(CLION_IDE)", and extensively uses TMEM features. However efforts for Hopper is still on the way :
For the moments , without TMEM, the performance is not significant, hence MegaMoE can be replaced with Fp4 EP V2 + FP8 DeepGeem + PDL in hopper platform. Note there are experiments show that fp4 (4 bit S12E1M) dispatch can have little performance influences. You need to modify EP v2 to support FP4 dispatch and then you can have full acceleration in Hopper platform. |
We benchmarked Mega MoE on DeepSeek-V4-Flash and DeepSeek-V4-Pro under 8-way expert parallelism (EP8), testing at various batch sizes (i.e., the number of tokens per rank) to cover different serving scenarios. All values are averaged across 8 ranks.
DeepSeek-V4-Flash
DeepSeek-V4-Flash has 256 experts with top-k=6 (each token is routed to 6 experts), a hidden dimension of 4096, and an intermediate hidden dimension of 2048.
DeepSeek-V4-Pro
DeepSeek-V4-Pro has 384 experts with top-k=6, a hidden dimension of 7168, and an intermediate hidden dimension of 3072.