Skip to content

Add various optimizations and Mega MoE benchmarks - #316

Merged
LyricZhao merged 4 commits into
mainfrom
mega-update
Apr 24, 2026
Merged

LyricZhao merged 4 commits into
mainfrom
mega-update

Conversation

@zheanxu

@zheanxu zheanxu commented Apr 24, 2026 •

Copy link
Copy Markdown
Collaborator

We benchmarked Mega MoE on DeepSeek-V4-Flash and DeepSeek-V4-Pro under 8-way expert parallelism (EP8), testing at various batch sizes (i.e., the number of tokens per rank) to cover different serving scenarios. All values are averaged across 8 ranks.

DeepSeek-V4-Flash

DeepSeek-V4-Flash has 256 experts with top-k=6 (each token is routed to 6 experts), a hidden dimension of 4096, and an intermediate hidden dimension of 2048.

Batch Size Time (us) Compute (TFLOPS) Global Memory (GB/s) Interconnect (GB/s) Speedup (vs legacy)
1 56.5 5 1311 1 1.96x
512 146.5 1056 3192 266 1.73x
8192 1283.1 1928 998 499 1.56x
32768 4855.5 2038 794 529 1.62x

DeepSeek-V4-Pro

DeepSeek-V4-Pro has 384 experts with top-k=6, a hidden dimension of 7168, and an intermediate hidden dimension of 3072.

Batch Size Time (us) Compute (TFLOPS) Global Memory (GB/s) Interconnect (GB/s) Speedup (vs legacy)
1 108.1 7 1758 1 1.61x
512 369.6 1098 4619 182 1.54x
8192 2818.5 2304 1094 393 1.50x
32768 10655.2 2438 692 417 1.54x

@zheanxu
zheanxu requested a review from LyricZhao April 24, 2026 07:40
@LyricZhao
LyricZhao merged commit 891d57b into main Apr 24, 2026
@kiankyars

kiankyars commented Apr 26, 2026 •

Copy link
Copy Markdown

We benchmarked Mega MoE on DeepSeek-V4-Flash and DeepSeek-V4-Pro under 8-way expert parallelism (EP8), testing at various batch sizes (i.e., the number of tokens per rank) to cover different serving scenarios. All values are averaged across 8 ranks.

To clarify here, is this the number of tokens per rank as it's stated, or the number of tokens across the node?

I think I'm just not being smart, but to clarify, across the entire node you would have 512 times 8 equals 4096 tokens in the second row?

@zheanxu

zheanxu commented Apr 27, 2026

Copy link
Copy Markdown
Collaborator Author

@kiankyars Yes, the batch size listed is the number of tokens per rank as stated. So for the row showing 512 tokens per rank under EP8, the total across the node would be 512 × 8 = 4,096. Your understanding is correct.

@kiankyars

Copy link
Copy Markdown

thanks!

Barry-Delaney added a commit to Barry-Delaney/DeepGEMM that referenced this pull request May 7, 2026
…ions and Mega MoE benchmarks) into nv_dev

Brings in deepseek-ai PR deepseek-ai#316 'Add various optimizations and Mega MoE benchmarks'
on top of the nv_dev branch (currently at c491439, the merge of PR deepseek-ai#314).

Conflicts resolved against the NV-side commits from PR deepseek-ai#314 (f6d98f2, 12cb7c0,
a97b74d):

  * deep_gemm/include/deep_gemm/scheduler/paged_mqa_logits.cuh: keep NV's
    host-passed num_next_n_atoms runtime parameter, plus 891d57b's kIsVarlen
    template + binary-search SM work distribution + varlen path.
  * deep_gemm/include/deep_gemm/impls/sm90_fp8_paged_mqa_logits.cuh: keep NV's
    kNumKVMulticast template + cluster multicast logic; add 891d57b's kIsVarlen
    template parameter with static_assert(!kIsVarlen) (SM90 doesn't support
    varlen).
  * csrc/jit_kernels/impls/smxx_fp8_fp4_paged_mqa_logits.hpp: unified Args
    struct (num_next_n_atoms + is_varlen + indices); SM90 instantiation passes
    is_varlen=false; SM100 TMA-box rule uses 891d57b's
    (is_varlen or next_n >= 2) ? 2 : 1.
  * csrc/apis/attention.hpp: host computes num_next_n_atoms to match the
    post-deepseek-ai#316 SM100 kernel rule
    (kNextNAtom = (is_varlen or next_n >= 2) ? 2 : 1,
     kNumNextNAtoms = ceil_div(next_n, kNextNAtom));
    SM90 multicast hard-codes num_next_n_atoms = 1. Without this fix, next_n=5
    over-schedules by 5/3 and triggers IMA on SM100.
  * tests/test_attention.py: keep NV's block_kv in {32, 64} and next_n=4 for
    SM90 multicast; add 891d57b's outer is_varlen loop (varlen is SM100-only).

Validated on B300 SXM6 (sm103) and H200 NVL (sm90):
test_attention.py passes the full enumeration on both architectures with no
IMA; numerical diff < 1e-3 (FP8) / 0.02 (FP4).
RayWang96 added a commit that referenced this pull request May 8, 2026
Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks)
@GwilliamHu

GwilliamHu commented May 26, 2026 •

Copy link
Copy Markdown

@LyricZhao Hi, sorry to bother you. I 've a question about MegaMoE. The activation input is seems already placed in symmetric memory. Should we force fp4 quantization outputs to target symmetric memory? Do the model scripts need modification?

RayWang96 pushed a commit that referenced this pull request Jun 2, 2026
…s) into nv_dev

Brings deepseek-ai PR #347 (88965b0, also pulls 714dd1a 'Update test_mega_moe.py')
onto nv_dev (currently at ac1f285 = #316 merge #328 + PR #342 IMA guard).
Merge base is 891d57b.

Conflicts resolved (5 files):
  * scheduler/paged_mqa_logits.cuh: take #347's refresh_num_kv_and_advance +
    reversed metadata allocation; drop nv_dev's PR #342 exist_q_atom_idx guard
    and get_atom_advance (both subsumed — reversed alloc keeps current_q_atom_idx
    in-bounds, refresh_num_kv_and_advance reproduces the varlen 1-or-2 advance).
    NV's host-passed num_next_n_atoms survives in the auto-merged metadata kernel
    (verified: single param, no inline recompute, coherent with reversed alloc).
  * csrc/apis/gemm.hpp: drop redundant C/D dtype asserts (#347 cleanup; the
    pre-existing d-dtype check + d==c check already cover them).
  * sm100_fp8_fp4_gemm_1d1d.cuh: take #347's unconditional BF16/FP32 C/D assert
    (supersedes nv_dev's accumulation-only assert).
  * tests/generators.py: take #347's broadened BF16-accumulation enumeration
    (out_dtype==bf16 and (dtype==bf16 or arch==10)) — strict superset of nv_dev's
    fp8+SM100 case.
  * tests/test_attention.py: merge both — keep NV's SM90 next_n=4 multicast +
    num_clusters metadata sizing; add #347's pool-size limits, batch_size=4096,
    and block_table/context_lens sanity asserts.

Auto-merged metadata kernel + smxx_fp8_fp4_paged_mqa_logits.hpp + attention.hpp
verified coherent across all three layers (kernel signature / launch / host call
all carry num_next_n_atoms + is_varlen + indices).

Validated on B-card (SM100) and H200 (SM90) incl next_n=4 multicast.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@hope1262946533

Copy link
Copy Markdown

Hi, I wanted to follow up on the test results — would you mind letting me know what physical device version and software build were used? Thank you!

@justice-dance

Copy link
Copy Markdown

Hi, I have a question regarding this operator.
Is this operator currently enabled for both the prefill and decode phases?
Does it yield favorable performance gains in low-latency, high-throughput and high-concurrency scenarios?
Or will we switch to alternative operators under certain specific workloads?

@yiakwy-xpu-ml-framework-team

yiakwy-xpu-ml-framework-team commented Aug 12, 2026 •

Copy link
Copy Markdown

Hi, I wanted to follow up on the test results — would you mind letting me know what physical device version and software build were used? Thank you!

The original PR is for blackwell platform with NVFP4 support, hence FP8 x FP4 mlp. The codes are guarded by "#if (defined(CUDA_ARCH) and (CUDA_ARCH >= 1000)) or defined(CLION_IDE)", and extensively uses TMEM features.

However efforts for Hopper is still on the way :

For the moments , without TMEM, the performance is not significant, hence MegaMoE can be replaced with Fp4 EP V2 + FP8 DeepGeem + PDL in hopper platform.

Note there are experiments show that fp4 (4 bit S12E1M) dispatch can have little performance influences. You need to modify EP v2 to support FP4 dispatch and then you can have full acceleration in Hopper platform.

@hope1262946533

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants