English | 中文
This repository contains the code for our paper, "FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates," along with instructions for downloading the released training data.
FlexiSLM is the first spoken language model that supports dynamic and controllable frame rates on both speech input and output. A single trained model can be steered between 12.5 Hz and 4.0 Hz without retraining, while its dynamic frame-rate mechanism adapts to the varying complexity of speech. FlexiSLM matches state-of-the-art 7B models even in reduced 6.25Hz frame rates. It also supports controllable frame rate generation.
- August 21, 2026: FlexiSLM is accepted to EMNLP 2026 Main Conference!
- August 20, 2026: Checkpoint release. We released the FlexiSLM-7B Stage 2 checkpoint and FlexiSLM-0.5B Stage 2 checkpoint reproduced with this codebase.
- August 6, 2026: Data release. We released FlexiSLM-Data-4M-s2s, FlexiSLM-Data-2M-s2s-compact, and FlexiSLM-Data-5M-t2t.
- August 2, 2026: Code release. We released the FlexiSLM-7B training and inference code.
git clone --recurse-submodules https://github.com/AmphionTeam/FlexiSLM.git
cd FlexiSLM
pip install -r requirements.txt- FlexiSLM-Data Details
- Inference Guide
- Training Guide
- Evaluation with Kimi-Audio-Evalkit
- Evaluation Results
- Citation
- Acknowledgements
- Project File Structure
We open-source the data produced by the following pipeline:
- Prompt collection and response generation. Text prompts are collected from public QA, instruction-following, and dialogue datasets. Responses are generated with Qwen3-Omni-30B-A3B. The resulting text pairs are released as
.
- Speech synthesis. Responses are synthesized with Qwen3-TTS, while prompts are synthesized with Fish-Audio using randomly sampled speaker prompts. The resulting 4.2M samples and approximately 26K hours of audio are released as
. Download size is about 2.8TB.
- Quality filtering and compression. Stricter filtering is applied and all audio is converted to MP3. The compact release contains 2.43M samples and approximately 14.8K hours of audio in about 385 GB:
We use this compact dataset for training.
We believe this is one of the largest open-source datasets for spoken language model training, and we hope this will especially benefit new researchers in this area. For data preview and statistics, please refer to the links above.
Use the checkpoint flag to select an inference checkpoint. The default is stage2_7B (faster inference). Use stage3_7B for the Stage 3 full-finetuned 7B weights.
| Flag | Hugging Face repo | Will be downloaded to |
|---|---|---|
stage2_7B (default) |
FlexiSLM/FlexiSLM-7B-Stage2 | models/FlexiSLM-7B-Stage2 |
stage3_7B |
FlexiSLM/FlexiSLM-7B-Stage3 | models/FlexiSLM-7B-Stage3 |
stage2_0.5B |
FlexiSLM/FlexiSLM-0_5B-Stage2 | models/FlexiSLM-0_5B-Stage2 |
Set auto_download=True to download the selected checkpoint (stage2_7B by default, or stage3_7B / stage2_0.5B), plus the Qwen2.5-Omni audio encoder, SenseVoice, FlexiCodec, flow-matching decoder, and vocoder files into models/ on first run. Later runs reuse the local copies.
from pathlib import Path
import soundfile as sf
import torch
from src.inference_flexislm import (
FlexiSLMInferenceConfig,
FlexiSLMInference,
)
config = FlexiSLMInferenceConfig(
auto_download=True,
checkpoint="stage2_7B", # or "stage3_7B" / "stage2_0.5B"
use_flow_matching_decoder=True,
flow_matching_prompt_audio_path=str(
Path("examples/input.wav").resolve()
),
enable_flexible_framerate=True,
default_input_framerate=8.0,
default_output_framerate=8.0,
torch_dtype="bfloat16",
attn_implementation="flash_attention_2",
)
engine = FlexiSLMInference(config, device="cuda:0")
def save_audio(result, output_path):
waveform = result.get("audio")
if waveform is None:
raise RuntimeError("The model did not return decoded audio")
if torch.is_tensor(waveform):
waveform = waveform.detach().float().cpu().numpy()
sf.write(Path(output_path), waveform.squeeze(), 24_000)
# Text-to-speech
result = engine.generate_tts(
sentence="This is a test sentence.",
output_framerate=8.0,
)
save_audio(result, "tts.wav")
# Automatic speech recognition
result = engine.generate_from_audio(
audio_path="examples/input.wav",
text_query="Please transcribe the audio.",
input_framerate=8.0,
output_framerate=8.0,
output_text_only=True,
)
print(result["text"])
# Audio question answering
result = engine.generate_from_audio(
audio_path="examples/question.wav",
text_query="",
input_framerate=8.0,
output_framerate=8.0,
output_text_only=True,
)
print(result["text"])
# Speech-to-speech generation
result = engine.generate_from_audio(
audio_path="examples/input.wav",
text_query="",
input_framerate=8.0,
output_framerate=8.0,
output_text_only=False,
)
save_audio(result, "s2s.wav")Download the checkpoint you want to run. Auxiliary encoder and codec files are shared by both sizes.
MODEL_ROOT="$PWD/models"
# if you want to run stage2_7B (default; faster inference)
hf download FlexiSLM/FlexiSLM-7B-Stage2 --local-dir "$MODEL_ROOT/FlexiSLM-7B-Stage2"
# if you want to run stage3_7B
hf download FlexiSLM/FlexiSLM-7B-Stage3 --local-dir "$MODEL_ROOT/FlexiSLM-7B-Stage3"
# if you want to run stage2_0.5B
hf download FlexiSLM/FlexiSLM-0_5B-Stage2 --local-dir "$MODEL_ROOT/FlexiSLM-0_5B-Stage2"
# required files
hf download FlexiSLM/Qwen2_5-Omni-Audio_Encoder --local-dir "$MODEL_ROOT/Qwen2_5-Omni-Audio_Encoder"
hf download FunAudioLLM/SenseVoiceSmall --local-dir "$MODEL_ROOT/SenseVoiceSmall"
hf download jiaqili3/flexicodec \
12hz_v1_half_config.yaml \
nartts_flexicodec_only.safetensors \
nartts.safetensors \
--local-dir "$MODEL_ROOT/FlexiCodec"
hf download amphion/dualcodec-tts vocos_emilia.safetensors \
--local-dir "$MODEL_ROOT/FlexiCodec"Then reuse the Python API example from Section 1, replacing only the config = FlexiSLMInferenceConfig(...) block. Set checkpoint to match the weights you downloaded, or pass model_path directly:
model_root = Path.cwd() / "models"
config = FlexiSLMInferenceConfig(
checkpoint="stage2_7B", # or "stage3_7B" / "stage2_0.5B"
model_path=str(model_root / "FlexiSLM-7B-Stage2"), # or FlexiSLM-7B-Stage3 / FlexiSLM-0_5B-Stage2
qwen25o_encoder_path=str(model_root / "Qwen2_5-Omni-Audio_Encoder"),
qwen25o_encoder_config_path=str(
model_root / "Qwen2_5-Omni-Audio_Encoder/config.json"
),
flexicodec_ckpt_path=str(
model_root / "FlexiCodec/nartts_flexicodec_only.safetensors"
),
flexicodec_config_path=str(model_root / "FlexiCodec/12hz_v1_half_config.yaml"),
sensevoice_path=str(model_root / "SenseVoiceSmall"),
use_flow_matching_decoder=True,
flow_matching_ckpt_path=str(model_root / "FlexiCodec/nartts.safetensors"),
flow_matching_vocoder_path=str(
model_root / "FlexiCodec/vocos_emilia.safetensors"
),
flow_matching_prompt_audio_path=str(
Path("examples/input.wav").resolve()
),
enable_flexible_framerate=True,
default_input_framerate=8.0,
default_output_framerate=8.0,
torch_dtype="bfloat16",
attn_implementation="flash_attention_2",
)A minimal notebook is available at examples/inference.ipynb.
Batch inference reads requests from JSONL and uses a YAML file for model, input, output, and multi-GPU runtime settings. Committed examples are examples/requests.jsonl and examples/infer_7b.yaml.
examples/requests.jsonl:
{"index": 0, "task": "tts", "input": {"text": "FlexiSLM supports controllable speech generation."}, "metadata": {"sample_id": "tts-demo"}}
{"index": 1, "task": "asr", "input": {"audio_path": "examples/input.wav"}, "metadata": {"sample_id": "asr-demo"}}
{"index": 2, "task": "audio_qa", "input": {"audio_path": "examples/question.wav", "model_prompt": ""}, "metadata": {"sample_id": "qa-demo"}}
{"index": 3, "task": "s2s", "input": {"audio_path": "examples/input.wav"}, "metadata": {"sample_id": "s2s-demo"}}examples/infer_7b.yaml (abbreviated; see the file for the full config):
engine:
config:
checkpoint: stage2_7B # or stage3_7B / stage2_0.5B
model_path: models/FlexiSLM-7B-Stage2 # or models/FlexiSLM-7B-Stage3 / models/FlexiSLM-0_5B-Stage2
qwen25o_encoder_path: models/Qwen2_5-Omni-Audio_Encoder
# ... encoder / FlexiCodec / SenseVoice / flow-matching paths ...
use_flow_matching_decoder: true
enable_flexible_framerate: true
default_input_framerate: 8.0
default_output_framerate: 8.0
output_sample_rate: 24000
torch_dtype: bfloat16
attn_implementation: flash_attention_2
# input/output paths are resolved relative to the repository root
input:
path: examples/requests.jsonl
output:
trace_path: outputs/inference/traces.jsonl
audio_dir: outputs/inference/audio
error_path: outputs/inference/errors.jsonl
inference:
checkpoint: models/FlexiSLM-7B-Stage2 # or models/FlexiSLM-0_5B-Stage2
target_framerate_hz: 8.0
output_sample_rate: 24000
runtime:
devices: [cuda:0]
workers_per_device: 1
fail_fast: falseRun after downloading a released checkpoint and shared encoder/codec files (see Section 2):
python -m src.infer examples/infer_7b.yamlinput/output paths are resolved relative to the repository root. engine.config model paths and JSONL audio_path values are relative to the working directory (run from the repo root). engine.config.checkpoint selects stage2_7B (default), stage3_7B, or stage2_0.5B. inference.checkpoint is the local weights path recorded in traces. To fetch weights automatically instead of setting model_path, use engine.config.auto_download: true with the desired engine.config.checkpoint. Optional inference.transcribe_model_path (for example models/whisper-large-v3) ASR-transcribes generated s2s audio; download Whisper first if you enable it.
The runner writes one unified JSONL trace and stores generated speech under output.audio_dir.
FlexiSLM training has three stages:
- Talker and input-module pre-training. Freeze the Qwen backbone and train the Talker, audio embeddings, and input frame-merging module. For better performance, out released checkpoint had the talker module separately trained with TTS-only and merged its talker weights into stage 1 weights. But we confirmed that without talker weights in stage 1, the performance in Stage 2 is still good.
- Multi-task LoRA fine-tuning. Train the Talker and input modules while adapting the Thinker with LoRA.
- Full fine-tuning. Merge the Stage 2 LoRA weights into the Thinker, enable the Talker-to-Thinker connection, and train all model components.
Our released checkpoints are trained with the same settings using 8 A100 GPUs. The configs here are also adapted for 8 A100 GPUs.
Our training uses on-the-fly data extraction, so it is very convenient to use.
MODEL_ROOT="$PWD/models"
TRAIN_DATA_ROOT="$PWD/data/training"
BENCHMARK_DATA_ROOT="$PWD/data/benchmarks"
# Previously downloaded in inference guide
hf download FlexiSLM/Qwen2_5-Omni-Audio_Encoder --local-dir "$MODEL_ROOT/Qwen2_5-Omni-Audio_Encoder"
hf download FunAudioLLM/SenseVoiceSmall --local-dir "$MODEL_ROOT/SenseVoiceSmall"
hf download jiaqili3/flexicodec 12hz_v1_half_config.yaml nartts_flexicodec_only.safetensors --local-dir "$MODEL_ROOT/FlexiCodec"
# Required for training
hf download Qwen/Qwen2.5-7B-Instruct --local-dir "$MODEL_ROOT/Qwen2.5-7B-Instruct"
hf download Qwen/Qwen2.5-0.5B-Instruct --local-dir "$MODEL_ROOT/Qwen2.5-0.5B-Instruct"
hf download openai/whisper-large-v3 --local-dir "$MODEL_ROOT/whisper-large-v3"
# S2S Data for Stage 2 and 3 (includes data/ and data_part2_webq_trivia/)
hf download FlexiSLM/FlexiSLM-Data-2M-s2s-compact \
--repo-type dataset \
--local-dir "$TRAIN_DATA_ROOT/FlexiSLM-Data-2M-s2s-compact"
# ASR+TTS Data for Stage 1, 2, and 3
hf download FlexiSLM/asrtts_packed_webdataset \
--repo-type dataset \
--local-dir "$TRAIN_DATA_ROOT/asrtts_packed_webdataset"
Training recipes set report_to: swanlab. Before launching, export your SwanLab API key in the shell:
export SWANLAB_API_KEY="your_swanlab_api_key"Training arguments are stored in YAML files under config/; launchers live under scripts/:
| Stage | Configuration | Data config | Launcher | Initialization |
|---|---|---|---|---|
| Stage 1 (7B) | config/train_stage1_7B.yaml |
config/datasets/train_stage1.yaml |
scripts/train_stage1_7B.sh |
Qwen2.5-7B Instruct model |
| Stage 2 (7B) | config/train_stage2_7B.yaml |
config/datasets/train_stage2_3.yaml |
scripts/train_stage2_7B.sh |
released Stage 1 (Hub) |
| Stage 3 (7B) | config/train_stage3_7B.yaml |
config/datasets/train_stage2_3.yaml |
scripts/train_stage3_7B.sh |
released Stage 2 with LoRA merged (Hub) |
| Stage 1 (0.5B) | config/train_stage1_0_5B.yaml |
config/datasets/train_stage1.yaml |
scripts/train_stage1_0_5B.sh |
Qwen2.5-0.5B Instruct model |
| Stage 2 (0.5B) | config/train_stage2_0_5B.yaml |
config/datasets/train_stage2_3.yaml |
scripts/train_stage2_0_5B.sh |
released Stage 1 (Hub) |
| Stage 3 (0.5B) | config/train_stage3_0_5B.yaml |
config/datasets/train_stage2_3.yaml |
scripts/train_stage3_0_5B.sh |
merged 0.5B Stage 2 checkpoint |
Stage 2 sets resume_from_checkpoint to the released Stage 1 Hub repo (downloaded into models/ if missing). Stage 3 resumes from Stage 2: the launcher first merges LoRA from the released Stage 2 Hub repo into models/FlexiSLM-7B-Stage2-merged, then runs full-parameter training without LoRA (config/train_stage3_7B.yaml, matching the released Stage 3 recipe). Launch each stage after updating its YAML:
bash scripts/train_stage1_7B.sh
bash scripts/train_stage2_7B.sh
bash scripts/train_stage3_7B.sh
# for 0.5B
bash scripts/train_stage1_0_5B.sh
bash scripts/train_stage2_0_5B.sh
bash scripts/train_stage3_0_5B.shYAML values can be overridden on the command line:
bash scripts/train_stage2_7B.sh \
--resume_from_checkpoint FlexiSLM/FlexiSLM-7B-Stage1 \
--output_dir outputs/train_stage2_7B \
--learning_rate 2e-5
bash scripts/train_stage3_7B.sh \
--output_dir outputs/train_stage3_7B \
--max_steps 30000FlexiSLM uses the bundled Kimi-Audio-Evalkit submodule to evaluate VoiceBench, OpenAudioBench, and LibriSpeech. Inference and scoring are separate. Run all commands below from the FlexiSLM repository root.
Use a separate Python environment for Evalkit scoring. The Evalkit dependency set (see Kimi-Audio-Evalkit/requirements.txt) pins packages such as sacrebleu==1.5.1 and an older PyTorch stack that conflict with FlexiSLM training/inference. Keep your training/inference env (for example pip install -r requirements.txt) unchanged, and install Evalkit requirements only in a dedicated env such as kimi-audio-evalkit or eval.
A fresh --recurse-submodules clone already contains the Evalkit. For an existing clone, initialize it with git submodule update --init --recursive. Then create/activate the Evalkit env, install its requirements, and download the benchmark data:
# Example: dedicated conda env (do NOT install this into the FlexiSLM train/infer env)
conda create -n kimi-audio-evalkit python=3.10 -y
conda activate kimi-audio-evalkit
pip install -r Kimi-Audio-Evalkit/requirements.txt
BENCHMARK_DATA_ROOT="$PWD/data/benchmarks"
python Kimi-Audio-Evalkit/data/download_benchmark.py \
--datasets VoiceBench,OpenAudioBench,LibriSpeech \
--output-dir "$BENCHMARK_DATA_ROOT"Build requests and run FlexiSLM inference in your training/inference environment (not the Evalkit env). VoiceBench and OpenAudioBench use s2s (spoken answers, 24 kHz WAVs). MMSU and OpenBookQA are omitted. LibriSpeech stays ASR (text only):
# activate the FlexiSLM train/infer env first
export TRANSCRIBE_MODEL_PATH="${TRANSCRIBE_MODEL_PATH:-models/whisper-large-v3}"
python local/build_vb_oab_requests.py \
--benchmark voicebench \
--data-root data/benchmarks \
--out-dir outputs/evaluation/requests/voicebench \
--task s2s \
--transcribe-model-path "$TRANSCRIBE_MODEL_PATH" \
--subsets sd-qa advbench ifeval alpacaeval_full commoneval
python local/build_vb_oab_requests.py \
--benchmark openaudiobench \
--data-root data/benchmarks \
--out-dir outputs/evaluation/requests/openaudiobench \
--task s2s \
--transcribe-model-path "$TRANSCRIBE_MODEL_PATH"
python local/build_librispeech_requests.py \
--data-root data/benchmarks \
--out-dir outputs/evaluation/requests/librispeechRun inference with a rate-specific CONFIG. The launcher uses every GPU in CUDA_VISIBLE_DEVICES when set, otherwise all visible GPUs (nvidia-smi), and assigns one job per GPU per wave. Generated s2s WAVs are stored next to each trace under audio/.
# 12.5 Hz in/out → outputs/evaluation/traces/12_5hz/
CONFIG=config/infer_benchmarks_12_5hz.yaml bash scripts/infer_benchmarks.sh
# 6.25 Hz in/out → outputs/evaluation/traces/6_25hz/
CONFIG=config/infer_benchmarks_6_25hz.yaml bash scripts/infer_benchmarks.shSwitch to the Evalkit environment, export the DeepSeek API key, and run the eval YAML that matches the inference rate. The YAML lists every job; no --job flags are required.
conda activate kimi-audio-evalkit
export DEEPSEEK_API_KEY="your_deepseek_api_key"
export PYTHONPATH="$PWD:$PWD/Kimi-Audio-Evalkit${PYTHONPATH:+:$PYTHONPATH}"
# After CONFIG=config/infer_benchmarks_12_5hz.yaml
python -m src.eval config/eval_benchmarks_12_5hz.yaml
# After CONFIG=config/infer_benchmarks_6_25hz.yaml
python -m src.eval config/eval_benchmarks_6_25hz.yamlResults are written under outputs/evaluation/results/{12_5,6_25}hz/. LibriSpeech WER does not need DEEPSEEK_API_KEY; VoiceBench / OpenAudioBench LLM-judge jobs do.
The table below uses DeepSeek-V4-Flash (4 Flash / Deepseek-V4-Flash-0731) as the judge. Released checkpoints were evaluated with matching input/output frame rates via the scripts above. For FlexiSLM s2s traces, s2t is the model text channel (output.text) and s2s is Whisper ASR of the spoken response.
| Benchmark | Metric | Qwen2.5-Omni s2t | Qwen2.5-Omni s2s | FlexiSLM-7B-Stage2 12.5 Hz s2t | FlexiSLM-7B-Stage2 12.5 Hz s2s | FlexiSLM-7B-Stage2 6.25 Hz s2t | FlexiSLM-7B-Stage2 6.25 Hz s2s |
|---|---|---|---|---|---|---|---|
| LibriSpeech | test-clean (WER ↓) | 2.38 | — | 2.14 | — | 3.43 | — |
| test-other (WER ↓) | 4.21 | — | 5.75 | — | 6.15 | — | |
| OpenAudioBench | Llama Questions (Acc ↑) | 76.85 | 72.24 | 80.67 | 74.00 | 80.67 | 72.33 |
| Web Questions (Acc ↑) | 52.4 | 51.5 | 58.2 | 55.2 | 59.0 | 55.1 | |
| TriviaQA (Acc ↑) | 57.6 | 56.16 | 63.3 | 52.5 | 63.3 | 52.6 | |
| VoiceBench | AlpacaEval (Score ↑) | 3.71 | 3.45 | 4.96 | 4.78 | 4.94 | 4.85 |
| CommonEval (Score ↑) | 3.67 | 3.63 | 4.97 | 4.98 | 4.95 | 4.92 | |
| SD-QA (Acc ↑) | 55.88 | 50.99 | 61.84 | 55.88 | 59.67 | 54.07 | |
| AdvBench (Acc ↑) | - | 98.65 | — | 94.04 | — | 94.42 |
The next table uses DeepSeek-V4.1-Flash (4.1 Flash / deepseek-flash) as the judge on the same 12.5 Hz s2s traces, comparing released Stage 2 with Stage 3 (checkpoint-15000, FlexiSLM-7B-Stage3). Absolute scores are not comparable across judges.
| Benchmark | Metric | FlexiSLM-7B-Stage2 12.5 Hz s2t | FlexiSLM-7B-Stage2 12.5 Hz s2s | FlexiSLM-7B-Stage3 12.5 Hz s2t | FlexiSLM-7B-Stage3 12.5 Hz s2s |
|---|---|---|---|---|---|
| LibriSpeech | test-clean (WER ↓) | 2.14 | — | 2.75 | — |
| test-other (WER ↓) | 5.75 | — | 5.95 | — | |
| OpenAudioBench | Llama Questions (Acc ↑) | 80.47 | 72.73 | 82.15 | 74.58 |
| Web Questions (Acc ↑) | 60.33 | 57.64 | 61.39 | 58.63 | |
| TriviaQA (Acc ↑) | 63.89 | 53.23 | 65.69 | 58.97 | |
| VoiceBench | AlpacaEval (Score ↑) | 3.89 | 3.16 | 3.91 | 3.31 |
| CommonEval (Score ↑) | 3.79 | 3.52 | 3.95 | 3.78 | |
| SD-QA (Acc ↑) | 61.84 | 55.88 | 60.94 | 54.79 | |
| AdvBench (Acc ↑) | — | 94.04 | — | 97.88 |
If you find our work useful, please consider citing:
@article{li2026flexislm,
title={FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates},
author={Li, Jiaqi and Wang, Chaoren and Tian, Xiaohai and Chen, Mingjie and Liang, Xinyu and Li, Xu and Lin, Yufan and Qiu, Junwen and Zhang, Jun and Lu, Lu and others},
journal={arXiv preprint arXiv:2606.31247},
year={2026}
}- Our work uses Qwen 2.5 as the backbone and Qwen2.5-Omni as the audio encoder.
- Our training framework is largely based on Transformers.
- Our evaluation uses Kimi-Audio-Evalkit
- Our previous open-source works FlexiCodec and DualCodec are foundational to this work.
This project is licensed under the MIT License.
If you want to understand the project structure, you can refer to the following:
FlexiSLM/
├── assets/ # Documentation and demo assets
├── config/ # Training and dataset YAML configurations
│ └── datasets/ # Dataset recipes used by training
├── data/ # Downloaded data
│ ├── benchmarks/ # VoiceBench, OpenAudioBench, and LibriSpeech
│ └── training/ # Released FlexiSLM training datasets
├── examples/ # Inference notebook and small examples
├── local/ # Data conversion and benchmark request tools
├── Kimi-Audio-Evalkit/ # Evaluation toolkit
├── models/ # Downloaded models
├── scripts/ # Training launchers and runtime setup
├── src/ # Model, training, inference, and evaluation code
│ ├── dataset/ # Dataset loading and collation
│ ├── eval/ # Kimi-Audio-Evalkit adapters
│ ├── infer/ # YAML-driven inference runner
│ ├── models/ # FlexiSLM and vendored FlexiCodec implementation
└── trainer/ # Trainer implementation