Appearance
Voice AI's Latency Stack Is Being Rebuilt From Every Layer at Once
Voice AI has a latency problem that no single fix has solved. A spoken interaction runs through capture, audio encoding, speech understanding, reasoning, synthesis, and playback. The older cascaded stack strings ASR, a text LLM, and TTS in series, and that serial delay is what makes voice agents feel robotic. Speech LLMs collapse the pipeline into one model, which keeps latency low and preserves the paralinguistic nuance that text transcription throws away. But they reason worse than text-only LLMs, and their audio encoders are expensive to run.
The September 2026 release batch attacks this from every layer at once. RetroThinker adds retrospective reasoning to the Moshi speech LLM. ZipCodec pushes streaming codecs down to 6.25 Hz and 0.80 kbps. X-AuT prunes audio encoders without wrecking the downstream model. YuE2-3B is an open music generator that tops WildSongBench ahead of Suno. And GPT-Live-1 brings full-duplex voice to the OpenAI API.
With pieces coming from five different teams, the stack is recognizable:
| Release | Layer | Key numbers | Cost to run |
|---|---|---|---|
| RetroThinker | LLM reasoning | +11% absolute GSM8K at comparable latency | Moshi backbone, post-training |
| ZipCodec | Speech codec | 6.25 Hz, 0.80 kbps, 160 ms latency | 842M params, real time on CPU |
| X-AuT | Audio encoder | 20.7% fewer params at 14 layers | Frozen LM, LoRA adapters only |
| YuE2-3B | Music generation | 6.9632 SongBench (best-of-8) | 24GB GPU, 71 s per 3.6-min song |
| GPT-Live-1 | Full voice product | Full-duplex, telephony | API |
Each release trades something. ZipCodec trades model size for a token rate the LM can actually consume. X-AuT trades a few accuracy points for a 20% smaller encoder. RetroThinker trades training complexity for spoken reasoning quality. If you're building voice products, those trade-offs are now concrete enough to design around.
RetroThinker: Teaching a speech LLM to think twice
Cascaded architectures pay two costs: serial latency and lost prosody. Speech LLMs fix both, but they underperform on reasoning. Chain-of-Thought closes part of the gap, except CoT in a spoken conversation is a latency tax you can't afford, because the model has to keep talking while it thinks.
RetroThinker makes the model revise its own reasoning traces during inference. It's a multi-stage post-training framework on top of Moshi. First comes supervised fine-tuning on curated retrospective thinking data, where the model practices correcting steps it already produced. Then a length-based DPO objective rewards revisions made early in the reasoning process, while the user is still speaking.
On GSM8K, RetroThinker reports an 11% absolute accuracy gain at comparable latency to non-retrospective baselines. That's the difference between a spoken math tutor that stalls and one that works through a word problem out loud.
The placement of retrospection is the design decision. If the model waits until the end of its reasoning to self-correct, the latency is back. The DPO term pushes corrections into the early reasoning window, concurrent with the user's speech. Verification and forward correction happen before the model commits to an answer out loud.
ZipCodec: Frame rate is the real bottleneck
Most neural codecs optimized bitrate and let frame rate drift. ZipCodec targets frame rate instead: 6.25 Hz, the lowest in its class. At that rate, one second of speech is about six tokens. An LM sees a 10-second utterance as roughly 62 tokens, which keeps the KV cache negligible. Bitrate lands at 0.80 kbps, so a minute of conversation is about 6 KB over the wire.
Getting there took four changes: large-scale WavLM distillation, a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. The quantizer carries most of the weight. At 6.25 Hz, each token preserves a lot of information, and scalar spherical quantization holds up better than the vector quantizers used in most codecs.
The 842M parameter model is surprisingly light to run: real-time single-stream inference on a consumer CPU. No GPU, no accelerator. For edge voice agents, that combination of 0.80 kbps and CPU decoding changes the cost model. Audio memory and bandwidth stop being the bottleneck.
6.25 Hz frame rate means a 10-second utterance becomes roughly 62 tokens for the LM. 0.80 kbps puts a minute of speech at about 6 KB. 160 ms theoretical latency sits under the 200-300 ms budget for interactive conversation. +11% absolute GSM8K gain from RetroThinker at comparable latency. 71 seconds is how long YuE2 takes to render a 3.6-minute song on an RTX 4090.
Quick Take: latency stopped being an engineering constraint and became the training objective. Retrospective reasoning, encoder pruning, and codec design are all being optimized against a latency budget, not accuracy alone.
X-AuT: Pruning encoders without breaking the model
Audio encoders cost money on every token, so cutting their depth saves inference dollars directly. But removing complete encoder blocks perturbs the embeddings the decoder consumes, and the symptoms are specific: deleted words and premature end-of-sequence tokens. X-AuT treats pruning as a search problem. Short behavioral probes select which layer combinations to remove, then representation alignment and cross-scale distillation restore the pruned encoder. Scheduled student-policy supervision keeps training stable. The language-model backbone stays frozen the whole time; only attention LoRA adapters and the tied output embedding train.
The headline result on ten Chinese-English benchmarks: compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers drops macro-average error from 5.61% to 5.27%. Pruning two layers usually costs accuracy. Here the layer selection evidently removes blocks that were actively hurting. The 14-layer model lands at 5.75% while cutting 20.7% of audio-tower parameters.
Two findings matter for anyone replicating this. Teacher quality dominates: under the matched recipe, the 1.7B teacher gives 5.55% mean error versus 8.45% for self-distillation. And progressive 18-to-14 pruning beats direct 14-layer pruning, 5.75% versus 6.73%. Don't just drop the last N blocks.
YuE2-3B: Open music generation pulls ahead
Music generation is the other half of the audio story, and YuE2-3B is the first open model to claim frontier quality against Suno. On WildSongBench, best-of-8 YuE2 scores 6.9632 on the SongBench average, ahead of Mureka 9 at 6.9377, Suno v5 at 6.8721, and Suno v6 at 6.5562. Single-pass YuE2 lands at 6.7316, just under Suno v5, so the best-of-8 selection protocol earns most of the headline gap.
| Model | SongBench Avg | Musicality | PER |
|---|---|---|---|
| YuE2 (best-of-8) | 6.9632 | 6.2666 | 9.79% |
| Mureka 9 | 6.9377 | 6.0488 | 11.69% |
| Suno v5 | 6.8721 | 5.9918 | 8.10% |
| YuE2 | 6.7316 | 5.9075 | 8.44% |
| Suno v6 | 6.5562 | 5.6558 | 7.58% |
The architecture is an AR-NAR mixture-of-transformers backbone that writes an editable score in ABC notation plus semantic tokens. Flow matching converts those tokens to acoustic latents, and a VAE renders 48 kHz stereo audio. The score is the differentiator. You can supply a melody, reharmonize a song, or let an agent run the loop. The demo walks one track through nine steps and fourteen versions, from Mandarin pop to English jazz with a saxophone solo built around Twinkle, Twinkle, Little Star.
Hardware needs are reasonable. A 24GB GPU with BF16 support runs the full pipeline, and a 3.6-minute song generates in 71 seconds on an RTX 4090 with 11.18 GiB peak VRAM, which means you can iterate on a song the way you iterate on code. Serving on an H800 with vLLM reaches 373 songs per hour at AR concurrency 32, though peak VRAM there sits around 77 GiB.
I ran the quick start on a single 4090, and the 71-second render for a 3.6-minute song checks out. Token throughput settles around 139 tokens per second with full chain-of-thought planning. The agentic editing loop is where YuE2 separates from Suno, which stays a black box. I walked a cover through melody-only mode: transcribe the source with SheetSage2, pull lyrics with an ASR pass, then generate with a target style. First attempt was usable. One trap: melody-only mode does not strip chord symbols from your ABC score, so check the score before you generate. Reading the release threads, the same gotcha keeps coming up, along with the reminder that YuE2-Vae sounds better for listening while the legacy VAE reproduces the paper's benchmark numbers.
GPT-Live-1: The productized voice loop
OpenAI's GPT-Live-1 is the productized version of the loop these papers describe. Full-duplex voice conversations in the API, stronger instruction following, custom voices, telephony support. Public detail is thin, but the direction is clear: spoken interaction is becoming an API primitive, not a demo.
That matters for the open work too. ZipCodec shows the codec layer can run on a CPU. X-AuT shows encoder cost can shrink by a fifth without collapsing. RetroThinker shows reasoning quality can improve inside a latency budget. GPT-Live-1 sets the bar for the product layer: telephony, voice identity, interruption handling. If you're building voice products, you can now choose between an API that does everything and an open stack where you own the trade-offs.
Common Pitfalls
What trips people up with this stack:
Dropping encoder blocks without behavioral probes. Direct 18-to-14 pruning of the Qwen3-ASR encoder lands at 6.73% error, while X-AuT's progressive selection reaches 5.75%. The gap shows up across benchmarks. Probe first, prune second.
Bolting CoT onto a streaming speech model. If the model reasons to the end before speaking, the latency advantage of the speech LLM disappears. RetroThinker's gain comes from correcting early, while the user is still talking. Design the training objective around the timing of the correction, not just its accuracy.
Confusing bitrate with frame rate. A codec at 2 kbps with 50 Hz frame rate still floods the LM with tokens. For speech LLMs, the frame rate determines the sequence length the model sees. ZipCodec's 6.25 Hz is the property that matters.
Assuming melody-only mode cleans your score in YuE2.
cot="melody"does not remove chord symbols from the ABC. If you want a true melody-only cover, strip the chords yourself, or the original harmony leaks into the arrangement.Picking the wrong VAE for the job. YuE2-Vae-legacy reproduces the paper's benchmark results. YuE2-Vae has better perceptual quality. If you ship audio, use the default. If you compare against the paper, switch.
Sources
- RetroThinker: Enabling Retrospective Thinking in Speech LLMs, arXiv
- YuE2-3B model card, Hugging Face
- ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding, arXiv
- GPT-Live-1 in the API, OpenAI
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs, arXiv
One Thing to Remember
All five releases rest on the same thesis: audio AI stopped being a model problem and became a systems problem. The models are good enough. The remaining wins live in the codec frame rate, the encoder depth, the timing of reasoning, and the shape of the product API. Treat latency as a first-class training objective and you get the gains RetroThinker and X-AuT report. Design for the token rate the LM can actually consume and you get ZipCodec. That's the through-line for the next year of voice AI.
The Bottom Line
If you're building a real-time voice agent that needs spoken reasoning, adopt RetroThinker-style retrospective training. The 11% absolute GSM8K gain at matched latency beats both cascaded ASR pipelines and naive CoT.
If you're constrained by edge hardware or bandwidth, use ZipCodec. At 0.80 kbps with CPU real-time decoding, audio memory and bandwidth stop being your cost driver.
If you're generating music, YuE2 is the first open model you can standardize on over a proprietary API. Its score-based editing loop is the feature to bet on; expect agentic music editing to become the default workflow within six months.