Repository navigation
SkyRL Q3 2026 Roadmap #1846
Description
Activity
- pinned this issue
on Jun 30, 2026 For weight syncing LoRA adapters, it seems a good goal would be to ensure we run
merge_lora=Falseso we only sync the adapter and not the full model.
#1852Reacted by Eric Tang and Ben RhodesIn this blog post, your roadmap section says "Multimodal support is still early and we have several major features planned, including sequence packing, Megatron backend support, long context training with context parallelism, and step-wise training"
Just wondering if any of those sit on the Q3 roadmap, or are more like Q4?
This is a great project btw, and very excited for the next quarter!
Reacted by Sumanth R Hegde and Catherine LeeMultimodal support is still early and we have several major features planned, including sequence packing, Megatron backend support, long context training with context parallelism, and step-wise training
Hi @benrhodes26 , currently multimodal is not a focus in Q3, but we are very happy to validate and guide community contributions for this. (Recent example: community contributed SFT support for VLMs: #1752)
Hello, I've addressed delta weight sync in #1902
Is batched multi-lora forward/backward in scope for the SkyRL tinker engine? Something along the lines of this work from Osmosis to enable batching of multiple LoRA jobs within the same forward pass (it seems now the Tinker backend will serialize the per-model batches and run each with _forward_backward_single_model_batch)?
hey @kailash109, good question, this is currently not planned - in general with large enough
max_tokens_per_microbatch, the trainer should be already compute bound, and so batching multiple tenant requests together won't speed up overall throughput - we expect the gains from multi-lora to largely be from the inference sidehowever, if you're interested, one open item on multi-lora fwd_bwd is that the current tinker engine only returns a batch once all fwd_bwd requests for all
model_ids are completed - this could be improved to return early for a givenmodel_idso that different trainers block minimally on one another@erictang000 Thanks for the response, that makes sense
And yes here is the PR for the early return #1920
Adding some thoughts on this roadmap from Trajectory!
Planned items:
• [P0] Fix B300 flashinfer issues
• [P1] Show metrics for the Tinker Engine: concurrent requests, number of tenants, training throughput.
- Suggestion: add on-demand profiling so we can use an environment flag to start the profiler, and send trace data to a remote store or local disk for users to examine. (#1846)
- Suggestion: add a memory OOM snapshot callback https://zdevito.github.io/2022/08/16/memory-snapshots.html
• [P0] Add streaming data loading for large data sets. Curious if you've considered the MosaicML streaming library for this. (#1846)
• [P0] Enable full FP8 training. Curious how mxfp8 and fp8 compare — one useful metric to track: % of activation values that overflow or truncate to zero after fp8 quantization/dequantization. (#1846)
• [P1] Add low-precision expert quantization. Happy to help drive this one as it's a direct step toward FP4 support. Would bump priority if FP4 is a goal. (https://github.com/NovaSky-AI/SkyRL/issues/1846%7C#1846, https://fireworks.ai/blog/scaling-optimizing-frontier-model-training)New feature work:
• [P1] Add rclone or fsspec support to store LoRA checkpoints across different cloud storage systems.
• [P1] Use the Kimi K3 model instead of Kimi K2.6. Roadmap currently only lists K2.6.
• [P1] Support the Nemotron 3.5 Nano model including its NVFP4 quantized version.Happy to help out on these items as well. Looking forward to an exciting Q3!
Reacted by Eric Tang and Charlie Ruan
Note to Community
This roadmap is a living document, and we welcome community input and feedback. We will update this roadmap to link to specific Issues and PRs for each sub-task.
Overview
SkyRL's primary focuses in Q3 2026 will continue to be pushing performance for async RL on large scale MoE models, and improving support for SkyRL's native Tinker API compatible training engine.
Tinker Engine
old tracker: #1380
Large Scale MoE Training (Megatron)
Megatron
old tracker: #1392
Quantization
See full quantization (FP8, MXFP8, NVFP4, INT4-QAT) tracker here: #1943
Router Replay
Large Scale Inference (vLLM)
Weight Syncing
Model Support
LoRA
Agent Integration
SFT
SFTTrainer#1839 @avigyabbAlgorithm
Miscellaneous