Skip to content
AmphionTeamPublic

About

[EMNLP 2026] FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Topics

Resources

Stars

34 stars

Watchers

2 watching

Forks

Repository files navigation

FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

arXiv Paper demo page dataset model WeChat Blog

English | 中文

Overview

This repository contains the code for our paper, "FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates," along with instructions for downloading the released training data.

FlexiSLM is the first spoken language model that supports dynamic and controllable frame rates on both speech input and output. A single trained model can be steered between 12.5 Hz and 4.0 Hz without retraining, while its dynamic frame-rate mechanism adapts to the varying complexity of speech. FlexiSLM matches state-of-the-art 7B models even in reduced 6.25Hz frame rates. It also supports controllable frame rate generation.

News

Installation

git clone --recurse-submodules https://github.com/AmphionTeam/FlexiSLM.git
cd FlexiSLM
pip install -r requirements.txt

Table of Contents

FlexiSLM-Data Details

We open-source the data produced by the following pipeline:

  1. Prompt collection and response generation. Text prompts are collected from public QA, instruction-following, and dialogue datasets. Responses are generated with Qwen3-Omni-30B-A3B. The resulting text pairs are released as dataset.
  2. Speech synthesis. Responses are synthesized with Qwen3-TTS, while prompts are synthesized with Fish-Audio using randomly sampled speaker prompts. The resulting 4.2M samples and approximately 26K hours of audio are released as dataset. Download size is about 2.8TB.
  3. Quality filtering and compression. Stricter filtering is applied and all audio is converted to MP3. The compact release contains 2.43M samples and approximately 14.8K hours of audio in about 385 GB: dataset We use this compact dataset for training.

We believe this is one of the largest open-source datasets for spoken language model training, and we hope this will especially benefit new researchers in this area. For data preview and statistics, please refer to the links above.

Inference

Use the checkpoint flag to select an inference checkpoint. The default is stage2_7B (faster inference). Use stage3_7B for the Stage 3 full-finetuned 7B weights.

Flag Hugging Face repo Will be downloaded to
stage2_7B (default) FlexiSLM/FlexiSLM-7B-Stage2 models/FlexiSLM-7B-Stage2
stage3_7B FlexiSLM/FlexiSLM-7B-Stage3 models/FlexiSLM-7B-Stage3
stage2_0.5B FlexiSLM/FlexiSLM-0_5B-Stage2 models/FlexiSLM-0_5B-Stage2

1. Python API (with Automatic downloading)

Set auto_download=True to download the selected checkpoint (stage2_7B by default, or stage3_7B / stage2_0.5B), plus the Qwen2.5-Omni audio encoder, SenseVoice, FlexiCodec, flow-matching decoder, and vocoder files into models/ on first run. Later runs reuse the local copies.

from pathlib import Path
import soundfile as sf
import torch

from src.inference_flexislm import (
    FlexiSLMInferenceConfig,
    FlexiSLMInference,
)

config = FlexiSLMInferenceConfig(
    auto_download=True,
    checkpoint="stage2_7B",  # or "stage3_7B" / "stage2_0.5B"
    use_flow_matching_decoder=True,
    flow_matching_prompt_audio_path=str(
        Path("examples/input.wav").resolve()
    ),
    enable_flexible_framerate=True,
    default_input_framerate=8.0,
    default_output_framerate=8.0,
    torch_dtype="bfloat16",
    attn_implementation="flash_attention_2",
)
engine = FlexiSLMInference(config, device="cuda:0")


def save_audio(result, output_path):
    waveform = result.get("audio")
    if waveform is None:
        raise RuntimeError("The model did not return decoded audio")
    if torch.is_tensor(waveform):
        waveform = waveform.detach().float().cpu().numpy()
    sf.write(Path(output_path), waveform.squeeze(), 24_000)


# Text-to-speech
result = engine.generate_tts(
    sentence="This is a test sentence.",
    output_framerate=8.0,
)
save_audio(result, "tts.wav")

# Automatic speech recognition
result = engine.generate_from_audio(
    audio_path="examples/input.wav",
    text_query="Please transcribe the audio.",
    input_framerate=8.0,
    output_framerate=8.0,
    output_text_only=True,
)
print(result["text"])

# Audio question answering
result = engine.generate_from_audio(
    audio_path="examples/question.wav",
    text_query="",
    input_framerate=8.0,
    output_framerate=8.0,
    output_text_only=True,
)
print(result["text"])

# Speech-to-speech generation
result = engine.generate_from_audio(
    audio_path="examples/input.wav",
    text_query="",
    input_framerate=8.0,
    output_framerate=8.0,
    output_text_only=False,
)
save_audio(result, "s2s.wav")

2. Python API (Manual downloading)

Download the checkpoint you want to run. Auxiliary encoder and codec files are shared by both sizes.

MODEL_ROOT="$PWD/models"

# if you want to run stage2_7B (default; faster inference)
hf download FlexiSLM/FlexiSLM-7B-Stage2 --local-dir "$MODEL_ROOT/FlexiSLM-7B-Stage2"
# if you want to run stage3_7B
hf download FlexiSLM/FlexiSLM-7B-Stage3 --local-dir "$MODEL_ROOT/FlexiSLM-7B-Stage3"
# if you want to run stage2_0.5B
hf download FlexiSLM/FlexiSLM-0_5B-Stage2 --local-dir "$MODEL_ROOT/FlexiSLM-0_5B-Stage2"

# required files
hf download FlexiSLM/Qwen2_5-Omni-Audio_Encoder --local-dir "$MODEL_ROOT/Qwen2_5-Omni-Audio_Encoder"
hf download FunAudioLLM/SenseVoiceSmall --local-dir "$MODEL_ROOT/SenseVoiceSmall"
hf download jiaqili3/flexicodec \
  12hz_v1_half_config.yaml \
  nartts_flexicodec_only.safetensors \
  nartts.safetensors \
  --local-dir "$MODEL_ROOT/FlexiCodec"
hf download amphion/dualcodec-tts vocos_emilia.safetensors \
  --local-dir "$MODEL_ROOT/FlexiCodec"

Then reuse the Python API example from Section 1, replacing only the config = FlexiSLMInferenceConfig(...) block. Set checkpoint to match the weights you downloaded, or pass model_path directly:

model_root = Path.cwd() / "models"
config = FlexiSLMInferenceConfig(
    checkpoint="stage2_7B",  # or "stage3_7B" / "stage2_0.5B"
    model_path=str(model_root / "FlexiSLM-7B-Stage2"),  # or FlexiSLM-7B-Stage3 / FlexiSLM-0_5B-Stage2
    qwen25o_encoder_path=str(model_root / "Qwen2_5-Omni-Audio_Encoder"),
    qwen25o_encoder_config_path=str(
        model_root / "Qwen2_5-Omni-Audio_Encoder/config.json"
    ),
    flexicodec_ckpt_path=str(
        model_root / "FlexiCodec/nartts_flexicodec_only.safetensors"
    ),
    flexicodec_config_path=str(model_root / "FlexiCodec/12hz_v1_half_config.yaml"),
    sensevoice_path=str(model_root / "SenseVoiceSmall"),
    use_flow_matching_decoder=True,
    flow_matching_ckpt_path=str(model_root / "FlexiCodec/nartts.safetensors"),
    flow_matching_vocoder_path=str(
        model_root / "FlexiCodec/vocos_emilia.safetensors"
    ),
    flow_matching_prompt_audio_path=str(
        Path("examples/input.wav").resolve()
    ),
    enable_flexible_framerate=True,
    default_input_framerate=8.0,
    default_output_framerate=8.0,
    torch_dtype="bfloat16",
    attn_implementation="flash_attention_2",
)

A minimal notebook is available at examples/inference.ipynb.

3. Batch Inference

Batch inference reads requests from JSONL and uses a YAML file for model, input, output, and multi-GPU runtime settings. Committed examples are examples/requests.jsonl and examples/infer_7b.yaml.

examples/requests.jsonl:

{"index": 0, "task": "tts", "input": {"text": "FlexiSLM supports controllable speech generation."}, "metadata": {"sample_id": "tts-demo"}}
{"index": 1, "task": "asr", "input": {"audio_path": "examples/input.wav"}, "metadata": {"sample_id": "asr-demo"}}
{"index": 2, "task": "audio_qa", "input": {"audio_path": "examples/question.wav", "model_prompt": ""}, "metadata": {"sample_id": "qa-demo"}}
{"index": 3, "task": "s2s", "input": {"audio_path": "examples/input.wav"}, "metadata": {"sample_id": "s2s-demo"}}

examples/infer_7b.yaml (abbreviated; see the file for the full config):

engine:
  config:
    checkpoint: stage2_7B  # or stage3_7B / stage2_0.5B
    model_path: models/FlexiSLM-7B-Stage2  # or models/FlexiSLM-7B-Stage3 / models/FlexiSLM-0_5B-Stage2
    qwen25o_encoder_path: models/Qwen2_5-Omni-Audio_Encoder
    # ... encoder / FlexiCodec / SenseVoice / flow-matching paths ...
    use_flow_matching_decoder: true
    enable_flexible_framerate: true
    default_input_framerate: 8.0
    default_output_framerate: 8.0
    output_sample_rate: 24000
    torch_dtype: bfloat16
    attn_implementation: flash_attention_2

# input/output paths are resolved relative to the repository root
input:
  path: examples/requests.jsonl

output:
  trace_path: outputs/inference/traces.jsonl
  audio_dir: outputs/inference/audio
  error_path: outputs/inference/errors.jsonl

inference:
  checkpoint: models/FlexiSLM-7B-Stage2  # or models/FlexiSLM-0_5B-Stage2
  target_framerate_hz: 8.0
  output_sample_rate: 24000

runtime:
  devices: [cuda:0]
  workers_per_device: 1
  fail_fast: false

Run after downloading a released checkpoint and shared encoder/codec files (see Section 2):

python -m src.infer examples/infer_7b.yaml

input/output paths are resolved relative to the repository root. engine.config model paths and JSONL audio_path values are relative to the working directory (run from the repo root). engine.config.checkpoint selects stage2_7B (default), stage3_7B, or stage2_0.5B. inference.checkpoint is the local weights path recorded in traces. To fetch weights automatically instead of setting model_path, use engine.config.auto_download: true with the desired engine.config.checkpoint. Optional inference.transcribe_model_path (for example models/whisper-large-v3) ASR-transcribes generated s2s audio; download Whisper first if you enable it.

The runner writes one unified JSONL trace and stores generated speech under output.audio_dir.

Training Guide

FlexiSLM training has three stages:

  1. Talker and input-module pre-training. Freeze the Qwen backbone and train the Talker, audio embeddings, and input frame-merging module. For better performance, out released checkpoint had the talker module separately trained with TTS-only and merged its talker weights into stage 1 weights. But we confirmed that without talker weights in stage 1, the performance in Stage 2 is still good.
  2. Multi-task LoRA fine-tuning. Train the Talker and input modules while adapting the Thinker with LoRA.
  3. Full fine-tuning. Merge the Stage 2 LoRA weights into the Thinker, enable the Talker-to-Thinker connection, and train all model components.

Our released checkpoints are trained with the same settings using 8 A100 GPUs. The configs here are also adapted for 8 A100 GPUs.

Our training uses on-the-fly data extraction, so it is very convenient to use.

1. Download additional checkpoints and dataset

MODEL_ROOT="$PWD/models"
TRAIN_DATA_ROOT="$PWD/data/training"
BENCHMARK_DATA_ROOT="$PWD/data/benchmarks"

# Previously downloaded in inference guide
hf download FlexiSLM/Qwen2_5-Omni-Audio_Encoder --local-dir "$MODEL_ROOT/Qwen2_5-Omni-Audio_Encoder"
hf download FunAudioLLM/SenseVoiceSmall --local-dir "$MODEL_ROOT/SenseVoiceSmall"
hf download jiaqili3/flexicodec 12hz_v1_half_config.yaml nartts_flexicodec_only.safetensors --local-dir "$MODEL_ROOT/FlexiCodec"

# Required for training
hf download Qwen/Qwen2.5-7B-Instruct --local-dir "$MODEL_ROOT/Qwen2.5-7B-Instruct"
hf download Qwen/Qwen2.5-0.5B-Instruct --local-dir "$MODEL_ROOT/Qwen2.5-0.5B-Instruct"
hf download openai/whisper-large-v3 --local-dir "$MODEL_ROOT/whisper-large-v3"

# S2S Data for Stage 2 and 3 (includes data/ and data_part2_webq_trivia/)
hf download FlexiSLM/FlexiSLM-Data-2M-s2s-compact \
  --repo-type dataset \
  --local-dir "$TRAIN_DATA_ROOT/FlexiSLM-Data-2M-s2s-compact"

# ASR+TTS Data for Stage 1, 2, and 3
hf download FlexiSLM/asrtts_packed_webdataset \
  --repo-type dataset \
  --local-dir "$TRAIN_DATA_ROOT/asrtts_packed_webdataset"

2. Launch Training

Training recipes set report_to: swanlab. Before launching, export your SwanLab API key in the shell:

export SWANLAB_API_KEY="your_swanlab_api_key"

Training arguments are stored in YAML files under config/; launchers live under scripts/:

Stage Configuration Data config Launcher Initialization
Stage 1 (7B) config/train_stage1_7B.yaml config/datasets/train_stage1.yaml scripts/train_stage1_7B.sh Qwen2.5-7B Instruct model
Stage 2 (7B) config/train_stage2_7B.yaml config/datasets/train_stage2_3.yaml scripts/train_stage2_7B.sh released Stage 1 (Hub)
Stage 3 (7B) config/train_stage3_7B.yaml config/datasets/train_stage2_3.yaml scripts/train_stage3_7B.sh released Stage 2 with LoRA merged (Hub)
Stage 1 (0.5B) config/train_stage1_0_5B.yaml config/datasets/train_stage1.yaml scripts/train_stage1_0_5B.sh Qwen2.5-0.5B Instruct model
Stage 2 (0.5B) config/train_stage2_0_5B.yaml config/datasets/train_stage2_3.yaml scripts/train_stage2_0_5B.sh released Stage 1 (Hub)
Stage 3 (0.5B) config/train_stage3_0_5B.yaml config/datasets/train_stage2_3.yaml scripts/train_stage3_0_5B.sh merged 0.5B Stage 2 checkpoint

Stage 2 sets resume_from_checkpoint to the released Stage 1 Hub repo (downloaded into models/ if missing). Stage 3 resumes from Stage 2: the launcher first merges LoRA from the released Stage 2 Hub repo into models/FlexiSLM-7B-Stage2-merged, then runs full-parameter training without LoRA (config/train_stage3_7B.yaml, matching the released Stage 3 recipe). Launch each stage after updating its YAML:

bash scripts/train_stage1_7B.sh
bash scripts/train_stage2_7B.sh
bash scripts/train_stage3_7B.sh
# for 0.5B
bash scripts/train_stage1_0_5B.sh
bash scripts/train_stage2_0_5B.sh
bash scripts/train_stage3_0_5B.sh

YAML values can be overridden on the command line:

bash scripts/train_stage2_7B.sh \
  --resume_from_checkpoint FlexiSLM/FlexiSLM-7B-Stage1 \
  --output_dir outputs/train_stage2_7B \
  --learning_rate 2e-5

bash scripts/train_stage3_7B.sh \
  --output_dir outputs/train_stage3_7B \
  --max_steps 30000

Evaluation with Kimi-Audio-Evalkit

FlexiSLM uses the bundled Kimi-Audio-Evalkit submodule to evaluate VoiceBench, OpenAudioBench, and LibriSpeech. Inference and scoring are separate. Run all commands below from the FlexiSLM repository root.

Use a separate Python environment for Evalkit scoring. The Evalkit dependency set (see Kimi-Audio-Evalkit/requirements.txt) pins packages such as sacrebleu==1.5.1 and an older PyTorch stack that conflict with FlexiSLM training/inference. Keep your training/inference env (for example pip install -r requirements.txt) unchanged, and install Evalkit requirements only in a dedicated env such as kimi-audio-evalkit or eval.

A fresh --recurse-submodules clone already contains the Evalkit. For an existing clone, initialize it with git submodule update --init --recursive. Then create/activate the Evalkit env, install its requirements, and download the benchmark data:

# Example: dedicated conda env (do NOT install this into the FlexiSLM train/infer env)
conda create -n kimi-audio-evalkit python=3.10 -y
conda activate kimi-audio-evalkit
pip install -r Kimi-Audio-Evalkit/requirements.txt

BENCHMARK_DATA_ROOT="$PWD/data/benchmarks"
python Kimi-Audio-Evalkit/data/download_benchmark.py \
  --datasets VoiceBench,OpenAudioBench,LibriSpeech \
  --output-dir "$BENCHMARK_DATA_ROOT"

1. Build Requests and Run Inference

Build requests and run FlexiSLM inference in your training/inference environment (not the Evalkit env). VoiceBench and OpenAudioBench use s2s (spoken answers, 24 kHz WAVs). MMSU and OpenBookQA are omitted. LibriSpeech stays ASR (text only):

# activate the FlexiSLM train/infer env first
export TRANSCRIBE_MODEL_PATH="${TRANSCRIBE_MODEL_PATH:-models/whisper-large-v3}"
python local/build_vb_oab_requests.py \
  --benchmark voicebench \
  --data-root data/benchmarks \
  --out-dir outputs/evaluation/requests/voicebench \
  --task s2s \
  --transcribe-model-path "$TRANSCRIBE_MODEL_PATH" \
  --subsets sd-qa advbench ifeval alpacaeval_full commoneval
python local/build_vb_oab_requests.py \
  --benchmark openaudiobench \
  --data-root data/benchmarks \
  --out-dir outputs/evaluation/requests/openaudiobench \
  --task s2s \
  --transcribe-model-path "$TRANSCRIBE_MODEL_PATH"
python local/build_librispeech_requests.py \
  --data-root data/benchmarks \
  --out-dir outputs/evaluation/requests/librispeech

Run inference with a rate-specific CONFIG. The launcher uses every GPU in CUDA_VISIBLE_DEVICES when set, otherwise all visible GPUs (nvidia-smi), and assigns one job per GPU per wave. Generated s2s WAVs are stored next to each trace under audio/.

# 12.5 Hz in/out → outputs/evaluation/traces/12_5hz/
CONFIG=config/infer_benchmarks_12_5hz.yaml bash scripts/infer_benchmarks.sh

# 6.25 Hz in/out → outputs/evaluation/traces/6_25hz/
CONFIG=config/infer_benchmarks_6_25hz.yaml bash scripts/infer_benchmarks.sh

2. Scoring

Switch to the Evalkit environment, export the DeepSeek API key, and run the eval YAML that matches the inference rate. The YAML lists every job; no --job flags are required.

conda activate kimi-audio-evalkit
export DEEPSEEK_API_KEY="your_deepseek_api_key"
export PYTHONPATH="$PWD:$PWD/Kimi-Audio-Evalkit${PYTHONPATH:+:$PYTHONPATH}"

# After CONFIG=config/infer_benchmarks_12_5hz.yaml
python -m src.eval config/eval_benchmarks_12_5hz.yaml

# After CONFIG=config/infer_benchmarks_6_25hz.yaml
python -m src.eval config/eval_benchmarks_6_25hz.yaml

Results are written under outputs/evaluation/results/{12_5,6_25}hz/. LibriSpeech WER does not need DEEPSEEK_API_KEY; VoiceBench / OpenAudioBench LLM-judge jobs do.

Evaluation Results

The table below uses DeepSeek-V4-Flash (4 Flash / Deepseek-V4-Flash-0731) as the judge. Released checkpoints were evaluated with matching input/output frame rates via the scripts above. For FlexiSLM s2s traces, s2t is the model text channel (output.text) and s2s is Whisper ASR of the spoken response.

Benchmark Metric Qwen2.5-Omni s2t Qwen2.5-Omni s2s FlexiSLM-7B-Stage2 12.5 Hz s2t FlexiSLM-7B-Stage2 12.5 Hz s2s FlexiSLM-7B-Stage2 6.25 Hz s2t FlexiSLM-7B-Stage2 6.25 Hz s2s
LibriSpeech test-clean (WER ↓) 2.38 — 2.14 — 3.43 —
test-other (WER ↓) 4.21 — 5.75 — 6.15 —
OpenAudioBench Llama Questions (Acc ↑) 76.85 72.24 80.67 74.00 80.67 72.33
Web Questions (Acc ↑) 52.4 51.5 58.2 55.2 59.0 55.1
TriviaQA (Acc ↑) 57.6 56.16 63.3 52.5 63.3 52.6
VoiceBench AlpacaEval (Score ↑) 3.71 3.45 4.96 4.78 4.94 4.85
CommonEval (Score ↑) 3.67 3.63 4.97 4.98 4.95 4.92
SD-QA (Acc ↑) 55.88 50.99 61.84 55.88 59.67 54.07
AdvBench (Acc ↑) - 98.65 — 94.04 — 94.42

The next table uses DeepSeek-V4.1-Flash (4.1 Flash / deepseek-flash) as the judge on the same 12.5 Hz s2s traces, comparing released Stage 2 with Stage 3 (checkpoint-15000, FlexiSLM-7B-Stage3). Absolute scores are not comparable across judges.

Benchmark Metric FlexiSLM-7B-Stage2 12.5 Hz s2t FlexiSLM-7B-Stage2 12.5 Hz s2s FlexiSLM-7B-Stage3 12.5 Hz s2t FlexiSLM-7B-Stage3 12.5 Hz s2s
LibriSpeech test-clean (WER ↓) 2.14 — 2.75 —
test-other (WER ↓) 5.75 — 5.95 —
OpenAudioBench Llama Questions (Acc ↑) 80.47 72.73 82.15 74.58
Web Questions (Acc ↑) 60.33 57.64 61.39 58.63
TriviaQA (Acc ↑) 63.89 53.23 65.69 58.97
VoiceBench AlpacaEval (Score ↑) 3.89 3.16 3.91 3.31
CommonEval (Score ↑) 3.79 3.52 3.95 3.78
SD-QA (Acc ↑) 61.84 55.88 60.94 54.79
AdvBench (Acc ↑) — 94.04 — 97.88

Citation

If you find our work useful, please consider citing:

@article{li2026flexislm,
  title={FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates},
  author={Li, Jiaqi and Wang, Chaoren and Tian, Xiaohai and Chen, Mingjie and Liang, Xinyu and Li, Xu and Lin, Yufan and Qiu, Junwen and Zhang, Jun and Lu, Lu and others},
  journal={arXiv preprint arXiv:2606.31247},
  year={2026}
}

Acknowledgements

License

This project is licensed under the MIT License.

Project Structure

If you want to understand the project structure, you can refer to the following:

FlexiSLM/
├── assets/                 # Documentation and demo assets
├── config/                 # Training and dataset YAML configurations
│   └── datasets/           # Dataset recipes used by training
├── data/                   # Downloaded data
│   ├── benchmarks/         # VoiceBench, OpenAudioBench, and LibriSpeech
│   └── training/           # Released FlexiSLM training datasets
├── examples/               # Inference notebook and small examples
├── local/                  # Data conversion and benchmark request tools
├── Kimi-Audio-Evalkit/     # Evaluation toolkit
├── models/                 # Downloaded models
├── scripts/                # Training launchers and runtime setup
├── src/                    # Model, training, inference, and evaluation code
│   ├── dataset/            # Dataset loading and collation
│   ├── eval/               # Kimi-Audio-Evalkit adapters
│   ├── infer/              # YAML-driven inference runner
│   ├── models/             # FlexiSLM and vendored FlexiCodec implementation
    └── trainer/            # Trainer implementation

About

[EMNLP 2026] FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Topics

Resources

Stars

34 stars

Watchers

2 watching

Forks

Used by

Contributors

Languages