Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

104 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLM Grandmaster Notes

📚The path to LLM mastery is paved with broken embeddings and resurrected gradients.

  • base
    • transformer
    • vit transformer
    • lm head
    • kv cache
    • GPU Architecture
      • SM80
      • SM90
      • SM100
      • SM120
      • Memory Hierarchy (HBM, L2 Cache, Shared Memory/SMEM, Register File)
      • Warp Specialization (Producer-Consumer Model used in FA3)
      • Distributed Shared Memory (DSM) / Thread Block Clusters
  • activation & mlp
    • SwiGLU / GeGLU / ReGLU
    • LLaMA MLP (Gated Linear Units)
  • attention
    • self attention
    • online attention
    • flash attention
    • flash attention 2
    • flash attention 3
    • flash decoding
    • flash decoding++
    • scaled dot-product attention (SDPA)
    • multi-head self-attention (MHSA)
    • multi-head attention (MHA)
    • grouped-query attention (GQA)
    • multi-query attention (MQA)
    • multi-head latent attention (MLA)
    • multi-token attention (MTA)
    • sage attention 1
    • sage attention 2
    • sage attention 2++
    • sage attention 3
    • paged attention
    • ring attention
    • ring flash attention
    • linear attention
    • lightning attention
    • native sparse attention (NSA)
    • grouped latent attention (GLA)
    • grouped-tied attention (GTA)
  • softmax
    • softmax
    • safe softmax
    • online softmax
  • kv cache optimization
    • sparse
    • quantization
    • allocator
    • window
    • share
  • norm
    • Batch Norm
    • Layer Norm
    • RMS Norm
  • position embedding
    • RoPE
    • AliBi
    • 2D RoPE
    • 3D RoPE
    • NTK-Award RoPE
    • Yarn
  • quantization
    • smooth quant
    • AWQ
    • KIVI
    • GPTQ
    • FP8 Training & Inference (E4M3 / E5M2 formats, AMAX scaling)
    • FP4 / FP6 / NF4 (QLoRA)
    • NVFP4 / MXFP4 / MXFP8 (Microscaling Formats, OCP specification)
    • KV Cache Quantization (INT4 / INT8 / FP8 KV)
    • AQLM / QuIP / QuIP# (Advanced vector quantization)
  • speculative decoding
    • Medusa
    • Lookahead decoding
    • NGram
    • OSD
    • Eagle 1,2,3
    • multi-token prediction (MTP)
    • Dflash
  • design
    • chunked prefill
    • continous batching
    • sliding window
    • CUDA Graph (Minimizing CPU launch overhead for short decoding steps)
    • Radix Attention / Prompt Cache (SGLang style prefix caching)
    • FlashDecoding with KV Splitting (Split-K attention for long contexts)
    • Chunked Prefill & Decode Co-run (handling prefill/decode interference)
  • reinforcement learning
    • PPO
    • GRPO
    • DAPO
    • GPG
    • DPO
    • KTO
    • IPO
    • SimPO
    • Rejection Sampling Fine-Tuning (RFT)
    • Online DPO / RLAIF (RL from AI Feedback)
  • gemm
    • deep gemm
    • cutlass
      • cooperative and ping-pong gemm scheduler
    • cublas
  • open source
    • flash mla
  • ptx instructions
    • mbarrier
    • cp.async
    • ldmatrix
    • mma
    • wgmma
    • cp.async.bulk / TMA (Tensor Memory Accelerator for SM90/SM100)
    • tcgen05.alloc / tcgen05.mma (Blackwell 5th Gen Tensor Core MMA)
    • TMEM (Tensor Memory) access primitives (Blackwell new memory space)
  • distributed parallel
    • Tensor Parallelism (TP) (Megatron-style)
    • Pipeline Parallelism (PP) (1F1B, Interleaved 1F1B, Zero-Bubble)
    • Sequence Parallelism (SP) (Megatron SP, DeepSpeed Ulysses)
    • Context Parallelism (CP) (Ring-based long context)
    • Expert Parallelism (EP) (for MoE)
    • ZeRO (Zero Redundancy Optimizer) / FSDP (ZeRO-1, ZeRO-2, ZeRO-3)
    • Activation Checkpointing (Gradient Checkpointing)
    • Communication Primitives (AllReduce, AllGather, ReduceScatter, P2P)
    • Hardware Topology (NVLink, NVSwitch, PCIe, InfiniBand, RoCE)
  • mixture of experts (moe)
    • Top-k Routing (Sparse MoE)
    • Shared Experts (DeepSeekMoE style)
    • Device-Limited Routing / Segmented Routing
    • Soft MoE / Fully-Differentiable MoE
    • Expert Capacity & Auxiliary Loss (Load Balancing)
    • Dropless MoE
    • Expert Offloading (for constrained inference)

About

🎓The path to LLM mastery is paved with broken embeddings and resurrected gradients. (CHINESE NOTES)

Topics

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages