Skip to content

SkyRL Q3 2026 Roadmap #1846

Description

@erictang000

Note to Community

This roadmap is a living document, and we welcome community input and feedback. We will update this roadmap to link to specific Issues and PRs for each sub-task.

Overview

SkyRL's primary focuses in Q3 2026 will continue to be pushing performance for async RL on large scale MoE models, and improving support for SkyRL's native Tinker API compatible training engine.

Tinker Engine

  • [P0] Improve sqlite DB path support for higher concurrency
  • [P0] Avoid inference engine initialization for the SFT code path [fix] Avoid inference engine initialization for the SFT code path #1878
  • [P1] Expose metrics for observability on the Tinker Engine (i.e. concurrent requests, number of tenants, etc...)
  • [P1] Provide support for external postgres deployments as an alternative for sqlite for high concurrency/prod deployments
  • [P1] Support storing only K-most recent checkpoints to save disk space.
  • [P1] Router replay (R3) support for the tinker engine
  • [P2] Support additional Tinker LoraConfig params: [tinker] Support additional Tinker LoraConfig parameters #1632
  • [P2] Support ref model in SkyRLTrainBackend

old tracker: #1380

Large Scale MoE Training (Megatron)

Megatron

old tracker: #1392

Quantization

See full quantization (FP8, MXFP8, NVFP4, INT4-QAT) tracker here: #1943

Router Replay

  • [P0] Add router replay test to CI: [megatron] add router replay tests to CI #1770
  • [P0] Validate multi-node inference for R3 with new vLLM implementation + evaluate data transfer with HTTP
  • [P1] Optimize data transfer between training and inference for R3

Large Scale Inference (vLLM)

Weight Syncing

Model Support

LoRA

Agent Integration

  • [P0] Add TITO proxy (handling rollout logprobs, router replay expert indices) for easier multi-turn agent harness integrations @xinze-zheng @CharlieFRuan

SFT

Algorithm

Miscellaneous

Activity

  1. pinned this issue on Jun 30, 2026
  2. casper-hansen commented on Jun 30, 2026

    @casper-hansen
    Contributor

    For weight syncing LoRA adapters, it seems a good goal would be to ensure we run merge_lora=False so we only sync the adapter and not the full model.
    #1852

  3. benrhodes26 commented on Jul 6, 2026

    @benrhodes26

    In this blog post, your roadmap section says "Multimodal support is still early and we have several major features planned, including sequence packing, Megatron backend support, long context training with context parallelism, and step-wise training"

    Just wondering if any of those sit on the Q3 roadmap, or are more like Q4?

    This is a great project btw, and very excited for the next quarter!

  4. SumanthRH commented on Jul 11, 2026

    @SumanthRH
    Member

    Multimodal support is still early and we have several major features planned, including sequence packing, Megatron backend support, long context training with context parallelism, and step-wise training

    Hi @benrhodes26 , currently multimodal is not a focus in Q3, but we are very happy to validate and guide community contributions for this. (Recent example: community contributed SFT support for VLMs: #1752)

  5. kailash109 commented on Jul 15, 2026

    @kailash109
    Contributor

    Hello, I've addressed delta weight sync in #1902

  6. kailash109 commented on Jul 17, 2026

    @kailash109
    Contributor

    Is batched multi-lora forward/backward in scope for the SkyRL tinker engine? Something along the lines of this work from Osmosis to enable batching of multiple LoRA jobs within the same forward pass (it seems now the Tinker backend will serialize the per-model batches and run each with _forward_backward_single_model_batch)?

  7. erictang000 commented on Jul 17, 2026

    @erictang000
    CollaboratorAuthor

    hey @kailash109, good question, this is currently not planned - in general with large enough max_tokens_per_microbatch, the trainer should be already compute bound, and so batching multiple tenant requests together won't speed up overall throughput - we expect the gains from multi-lora to largely be from the inference side

    however, if you're interested, one open item on multi-lora fwd_bwd is that the current tinker engine only returns a batch once all fwd_bwd requests for all model_ids are completed - this could be improved to return early for a given model_id so that different trainers block minimally on one another

  8. kailash109 commented on Jul 17, 2026

    @kailash109
    Contributor

    @erictang000 Thanks for the response, that makes sense

    And yes here is the PR for the early return #1920

  9. j316chuck commented on Jul 22, 2026

    @j316chuck
    Contributor

    Adding some thoughts on this roadmap from Trajectory!

    Planned items:
    • [P0] Fix B300 flashinfer issues
    • [P1] Show metrics for the Tinker Engine: concurrent requests, number of tenants, training throughput.
    - Suggestion: add on-demand profiling so we can use an environment flag to start the profiler, and send trace data to a remote store or local disk for users to examine. (#1846)
    - Suggestion: add a memory OOM snapshot callback https://zdevito.github.io/2022/08/16/memory-snapshots.html
    • [P0] Add streaming data loading for large data sets. Curious if you've considered the MosaicML streaming library for this. (#1846)
    • [P0] Enable full FP8 training. Curious how mxfp8 and fp8 compare — one useful metric to track: % of activation values that overflow or truncate to zero after fp8 quantization/dequantization. (#1846)
    • [P1] Add low-precision expert quantization. Happy to help drive this one as it's a direct step toward FP4 support. Would bump priority if FP4 is a goal. (https://github.com/NovaSky-AI/SkyRL/issues/1846%7C#1846,  https://fireworks.ai/blog/scaling-optimizing-frontier-model-training)

    New feature work:
    • [P1] Add rclone or fsspec support to store LoRA checkpoints across different cloud storage systems.
    • [P1] Use the Kimi K3 model instead of Kimi K2.6. Roadmap currently only lists K2.6.
    • [P1] Support the Nemotron 3.5 Nano model including its NVFP4 quantized version.

    Happy to help out on these items as well. Looking forward to an exciting Q3!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions