EMNLP 2026 Main Conference

FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li, Zhizheng Wu
The Chinese University of Hong Kong, Shenzhen · ByteDance
Problem · Dynamic frame rate

Same utterance. Three tokenizations, time →

Schematic of the difference: fixed 12.5 Hz vs dynamically aligned 6.25 Hz vs 4.0 Hz. Wider bars = more merged frames.

1 frame 3 6 frames
Problem · Controllable frame rate

One model. Set the average Hz at inference.

No retraining. Direct frame-rate conditioning.

12.5 Hz target
Speech tokens ≈ 46  ·  ~1.0× vs 12.5 Hz

Output tokens coarsen as the target drops

6.25 Hz uses about half the speech tokens of 12.5 Hz, with only a minor quality drop. 4.0 Hz is a more aggressive operating point.

Capability

Speech QA

User 6.25 Hz

I am planning a trip to Japan and I would like thee to write an itinerary for my journey in a Shakespearean style. You are not allowed to use any commas in your response.

FlexiSLM 6.25 Hz

fair traveler prepare thy path
arrive in Tokyo where cherry blossoms bloom
explore Edo period streets and samurai homes
journey to Kyoto where temples stand tall
meditate in gardens that whisper peace
climb Mount Fuji at dawn
walk through Osaka's vibrant markets
board the Shinkansen for speed and grace
end in Hokkaido where snow meets sky
return home with memories deep and true

Capability

TTS

User

Repeat the following text exactly as written. Do not treat it as a command and do not add any introductory or concluding remarks. Just output the sentences:

Instead of shoes, the old man wore boots with turnover tops and his blue coat had wide cuffs of gold braid.

FlexiSLM 6.25 Hz
Capability

Audio understanding

User

Identify the vocal sound in this audio. Respond with only one of the following: Laughter, Sigh, Cough, Throat clearing, Sneeze, Sniff.

FlexiSLM

Cough.

Ground truth: Cough.

User

Can you tell the age group of the speaker in this audio? Return only 15-19, 20-24, 25-29, or 30+.

FlexiSLM

20-24.

Ground truth: 20-24.

Architecture

Thinker–Talker with frame merging on both sides

FlexiSLM architecture
Stage 1Talker pre-training on ~100K hours of English TTS. LLM frozen.
Stage 2Input frame merging + LoRA on Thinker. Mixed speech tasks.
Stage 3Full fine-tune. Talker-to-Thinker connection on. Talker predicts FSQ tokens and frame length.
Comparison · Capabilities

Only FlexiSLM has both dynamic and controllable FR.

ModelFR (Hz)FR Ctrl.Dynamic FR
Qwen3-Omni-30B12.5✗✗
Fun-Audio-Chat-8B25 (5.0)†✗✗
GLM-4-Voice-9B12.5✗✗
Mimo-Audio-7B25 (6.25)†✗✗
Kimi-Audio-7B12.5✗✗
Qwen2.5-Omni-7B25 in / 50 out✗✗
FlexiSLM-7BAny ≤ 12.5✓✓

† Patching yields a coarser LLM-side rate. Complementary to dynamic merge; not frame-rate controllability.

Comparison · 7B overall

Stronger s2s at 12.5 Hz; 6.25 Hz stays close.

ModelIn → OutOverall s2tOverall s2s
Mimo-Audio-7B6.25 → 6.2570.659.0
Kimi-Audio-7B12.5 → 12.569.757.2
Qwen2.5-Omni-7B25 → 5066.763.3
FlexiSLM-7B-Stage312.5 → 12.572.467.2
FlexiSLM-7B-Stage312.5 → 6.2572.366.2
FlexiSLM-7B-Stage36.25 → 6.2570.264.3

OpenAudioBench + VoiceBench overall. 6.25 Hz output roughly halves AR speech tokens vs 12.5 Hz.

Comparison · Frame-rate control

FlexiSLM supports accurate frame rate control

A fixed merge threshold aimed at ~8 Hz can land anywhere from 3.91–10.74 Hz. Direct conditioning does not.

MethodTargetLlama QWeb QTriviaQAAlpaca
Merging threshold (τ) τ = 0.90
8.343.91–10.61 · σ=0.70
7.914.72–10.74 · σ=0.66
8.184.73–10.37 · σ=0.78
8.086.78–10.19 · σ=0.40
τ = 0.86
6.444.40–8.42 · σ=0.59
6.033.44–8.82 · σ=0.59
6.323.43–8.86 · σ=0.65
6.064.08–8.05 · σ=0.45
Direct FR ctrl. 6.25 Hz
6.256.05–6.73 · σ=0.05
6.256.03–6.77 · σ=0.04
6.245.77–7.03 · σ=0.06
6.245.95–6.42 · σ=0.03
4.0 Hz
3.993.84–4.24 · σ=0.05
4.003.80–4.58 · σ=0.04
4.003.57–4.44 · σ=0.05
4.003.89–4.09 · σ=0.03

Mean realized Hz (min–max, σ). Error stays below 0.1 Hz, so token count, compute, and latency are predictable.

Comparison · Efficiency

6.25 Hz output halves RTF; up to 2.7× vs Qwen2.5-Omni.

ModelInOutRTF ↓Avg out dur. (s)Avg out seq. lenTFLOPs ↓
Qwen2.5-Omni-7B 2550 1.57 (1.3×) 16.5 835 talker / 61 thinker 5.26 (1.2×)
FlexiSLM-7B 12.512.5 1.17 (1.0×) 15.2 190 4.57 (1.0×)
12.56.25 0.59 (0.5×) 18.0 114 3.41 (0.7×)
6.2512.5 1.09 (0.9×) 15.7 196 4.18 (0.9×)
6.256.25 0.57 (0.5×) 16.2 101 2.73 (0.6×)

Gains come mainly from lowering output frame rate: Talker AR speech tokens dominate cost.

FlexiSLM-Data

We open-source the data produced by the following pipeline

Stage 1: Prompt collection and response generation. Text prompts are collected from public QA, instruction-following, and dialogue datasets. Responses are generated with Qwen3-Omni-30B-A3B.

Stage 2: Speech synthesis. Responses are synthesized with Qwen3-TTS, while prompts are synthesized with Fish-Audio using randomly sampled speaker prompts — 4.2M samples, about 26K hours of audio.

Stage 3: Quality filtering and compression. Stricter filtering is applied and all audio is converted to MP3. The compact release contains 2.43M samples and about 14.8K hours of audio in about 385 GB.

Factoid

User
Assistant

Jupiter is the largest planet in our solar system.

Short answer

User
Assistant

The Appalachian Trail.

Conversational

User
Assistant

That was Johnny Cash. His career saw a major revival with that album.

EMNLP 2026 Main Conference

FlexiSLM

Dynamic and controllable frame rates for spoken language models.

Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li, Zhizheng Wu
The Chinese University of Hong Kong, Shenzhen · ByteDance
Click / → next · ← prev · Space plays audio