Skip to content

Public release 26/09/30 - #462

Merged
LyricZhao merged 1 commit into
mainfrom
0930-release
Sep 30, 2026
Merged

LyricZhao merged 1 commit into
mainfrom
0930-release

Conversation

@LyricZhao

@LyricZhao LyricZhao commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

This update introduces new features, performance improvements, and important API changes:

New features

  • Locality-aware Mega MoE: optional weight localization and domain-aware scheduling, with helpers for localized allocation (via MLOPart) and SM locality discovery. Saving 100W power and boosting 5-10% GPU frequency for heavy workload.
  • New deep_gemm.epilogue classes: Identity, Alpha, FP8Quantization, and opt-in BF16StochasticRounding for direct BF16 output on SM100.
  • BF16 batched einsums now support accumulation and FP32 output for bhr,hdr->bhd and bhd,hdr->bhr.
  • Sparse Indexer now supports head counts divisible by 4, from 4 through 32, for both contiguous and paged MXFP4/MXFP8 inputs.
  • Support for padded K strides in packed UE8M0 scaling factors.

Performance improvements

  • Improved Lightning Indexer kernels for V4.1, with better scheduling, KV pipelines, and head reduction for prefill and paged decoding.
  • Mega MoE epilogue prefetch and shared memory optimizations enable deeper pipelines for FP8xFP4 prefill.

Fixes

  • Fixed Sparse Indexer memory bounds and strengthened paged-cache and metadata validation.
  • Fixed Mega MoE combine reduction order for bitwise agreement with the shared-expert reference, and improved locality probing and fallback behavior.
  • Fixed runtime shutdown ordering for cuBLASLt resources and a missing FP8 header in Mega mHC.
  • Removed the Mega mHC test baseline's dependency on external tile-kernels code.

Breaking changes

  • The non-paged fp8_fp4_mqa_logits API now requires a positive max_seqlen_k and returns compressed logits. The dense non-paged and paged MQA APIs no longer accept clean_logits or logits_dtype; callers must mask entries beyond each row's valid KV span.
  • Dense MQA on SM100 now requires MXFP8/MXFP4 inputs with packed UE8M0 scales and BF16 weights, and returns BF16 logits. SM90 retains non-MX FP8 inputs with FP32 weights and logits. Weight rows must be 16-byte aligned.
  • Paged MQA on SM100 now requires the variable-length interface: pass indices and use next_n=1.
  • Removed the legacy fp8_mqa_logits, fp8_paged_mqa_logits, and fp8_gemm_nt_skip_head_mid APIs. Use the fp8_fp4_*mqa_logits APIs for MQA workloads.
  • Mega Gate now supports only scoring_func="sqrtsoftplus", which is also the default. get_bf16_mega_gate_config has been removed.

Related release

Contributors

  • Locality-aware Mega MoE and SM probing: @Yongqi-Zhuo @LyricZhao
  • Lightning Indexer, Sparse Indexer, and Mega MoE correctness: @zheanxu
  • GEMM epilogue classes, stochastic rounding, and BF16 einsum extensions: @xyf007
  • Mega MoE epilogue prefetch and shared-memory optimizations: @THUTwenTy
  • Packed scaling-factor strides: @xay5421

Add locality-aware Mega MoE execution, GEMM epilogue classes with BF16
stochastic rounding, wider sparse MQA head support, and packed SF strides.
Improve Indexer and Mega MoE pipelines, update attention and Mega Gate
interfaces, and include runtime and correctness fixes.

Export the current source through the OSS filter while retaining public
release history and the Mega MoE task-info release-ordering fix.

Co-authored-by: Anyi Xu <41318296+xay5421@users.noreply.github.com>
Co-authored-by: Chenggang Zhao <chenggangz@deepseek.com>
Co-authored-by: huahuaquq <289927700+huahuaquq@users.noreply.github.com>
Co-authored-by: TwenTy_ <47315699+THUTwenTy@users.noreply.github.com>
Co-authored-by: xyf007 <41227294+xyf007@users.noreply.github.com>
Co-authored-by: Yongqi Zhuo <Yongqi-Zhuo@users.noreply.github.com>
Co-authored-by: yukuai26 <1010238651@qq.com>
Co-authored-by: Zhean Xu <94977922+zheanxu@users.noreply.github.com>
@LyricZhao
LyricZhao merged commit 057ca59 into main Sep 30, 2026
2 of 3 checks passed
@GSM2017PMK-OSV

Copy link
Copy Markdown

Ое

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants