Factoid
Jupiter is the largest planet in our solar system.
Schematic of the difference: fixed 12.5 Hz vs dynamically aligned 6.25 Hz vs 4.0 Hz. Wider bars = more merged frames.
No retraining. Direct frame-rate conditioning.
6.25 Hz uses about half the speech tokens of 12.5 Hz, with only a minor quality drop. 4.0 Hz is a more aggressive operating point.
I am planning a trip to Japan and I would like thee to write an itinerary for my journey in a Shakespearean style. You are not allowed to use any commas in your response.
fair traveler prepare thy path
arrive in Tokyo where cherry blossoms bloom
explore Edo period streets and samurai homes
journey to Kyoto where temples stand tall
meditate in gardens that whisper peace
climb Mount Fuji at dawn
walk through Osaka's vibrant markets
board the Shinkansen for speed and grace
end in Hokkaido where snow meets sky
return home with memories deep and true
Repeat the following text exactly as written. Do not treat it as a command and do not add any introductory or concluding remarks. Just output the sentences:
Instead of shoes, the old man wore boots with turnover tops and his blue coat had wide cuffs of gold braid.
Identify the vocal sound in this audio. Respond with only one of the following: Laughter, Sigh, Cough, Throat clearing, Sneeze, Sniff.
Cough.
Ground truth: Cough.
Can you tell the age group of the speaker in this audio? Return only 15-19, 20-24, 25-29, or 30+.
20-24.
Ground truth: 20-24.
| Model | FR (Hz) | FR Ctrl. | Dynamic FR |
|---|---|---|---|
| Qwen3-Omni-30B | 12.5 | ✗ | ✗ |
| Fun-Audio-Chat-8B | 25 (5.0)† | ✗ | ✗ |
| GLM-4-Voice-9B | 12.5 | ✗ | ✗ |
| Mimo-Audio-7B | 25 (6.25)† | ✗ | ✗ |
| Kimi-Audio-7B | 12.5 | ✗ | ✗ |
| Qwen2.5-Omni-7B | 25 in / 50 out | ✗ | ✗ |
| FlexiSLM-7B | Any ≤ 12.5 | ✓ | ✓ |
† Patching yields a coarser LLM-side rate. Complementary to dynamic merge; not frame-rate controllability.
| Model | In → Out | Overall s2t | Overall s2s |
|---|---|---|---|
| Mimo-Audio-7B | 6.25 → 6.25 | 70.6 | 59.0 |
| Kimi-Audio-7B | 12.5 → 12.5 | 69.7 | 57.2 |
| Qwen2.5-Omni-7B | 25 → 50 | 66.7 | 63.3 |
| FlexiSLM-7B-Stage3 | 12.5 → 12.5 | 72.4 | 67.2 |
| FlexiSLM-7B-Stage3 | 12.5 → 6.25 | 72.3 | 66.2 |
| FlexiSLM-7B-Stage3 | 6.25 → 6.25 | 70.2 | 64.3 |
OpenAudioBench + VoiceBench overall. 6.25 Hz output roughly halves AR speech tokens vs 12.5 Hz.
A fixed merge threshold aimed at ~8 Hz can land anywhere from 3.91–10.74 Hz. Direct conditioning does not.
| Method | Target | Llama Q | Web Q | TriviaQA | Alpaca |
|---|---|---|---|---|---|
| Merging threshold (τ) | τ = 0.90 | 8.343.91–10.61 · σ=0.70 |
7.914.72–10.74 · σ=0.66 |
8.184.73–10.37 · σ=0.78 |
8.086.78–10.19 · σ=0.40 |
| τ = 0.86 | 6.444.40–8.42 · σ=0.59 |
6.033.44–8.82 · σ=0.59 |
6.323.43–8.86 · σ=0.65 |
6.064.08–8.05 · σ=0.45 |
|
| Direct FR ctrl. | 6.25 Hz | 6.256.05–6.73 · σ=0.05 |
6.256.03–6.77 · σ=0.04 |
6.245.77–7.03 · σ=0.06 |
6.245.95–6.42 · σ=0.03 |
| 4.0 Hz | 3.993.84–4.24 · σ=0.05 |
4.003.80–4.58 · σ=0.04 |
4.003.57–4.44 · σ=0.05 |
4.003.89–4.09 · σ=0.03 |
Mean realized Hz (min–max, σ). Error stays below 0.1 Hz, so token count, compute, and latency are predictable.
| Model | In | Out | RTF ↓ | Avg out dur. (s) | Avg out seq. len | TFLOPs ↓ |
|---|---|---|---|---|---|---|
| Qwen2.5-Omni-7B | 25 | 50 | 1.57 (1.3×) | 16.5 | 835 talker / 61 thinker | 5.26 (1.2×) |
| FlexiSLM-7B | 12.5 | 12.5 | 1.17 (1.0×) | 15.2 | 190 | 4.57 (1.0×) |
| 12.5 | 6.25 | 0.59 (0.5×) | 18.0 | 114 | 3.41 (0.7×) | |
| 6.25 | 12.5 | 1.09 (0.9×) | 15.7 | 196 | 4.18 (0.9×) | |
| 6.25 | 6.25 | 0.57 (0.5×) | 16.2 | 101 | 2.73 (0.6×) |
Gains come mainly from lowering output frame rate: Talker AR speech tokens dominate cost.
Stage 1: Prompt collection and response generation. Text prompts are collected from public QA, instruction-following, and dialogue datasets. Responses are generated with Qwen3-Omni-30B-A3B.
Stage 2: Speech synthesis. Responses are synthesized with Qwen3-TTS, while prompts are synthesized with Fish-Audio using randomly sampled speaker prompts — 4.2M samples, about 26K hours of audio.
Stage 3: Quality filtering and compression. Stricter filtering is applied and all audio is converted to MP3. The compact release contains 2.43M samples and about 14.8K hours of audio in about 385 GB.
Jupiter is the largest planet in our solar system.
The Appalachian Trail.
That was Johnny Cash. His career saw a major revival with that album.
Dynamic and controllable frame rates for spoken language models.