Skip to main content
Guides

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

- Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing content type paper published September 2026 Compressing Streaming Neural Audio Encoders via Latent-Space Distillation Authors Prasanth Yadla‡, Mohammad.

By Precis Daily Newsroom2 min read476 words
Illustration for: Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
Illustration
Key points
  • We train only the student encoder to regress the teacher’s per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher–student width mismatch.
  • Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model.
  • At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
  • Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing

content type paper published September 2026

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Authors Prasanth Yadla‡, Mohammad Samragh Razlighi‡, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang†, Yuan Liu§, Zhen Huang, Xiaodan Zhuang

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes—the last representation the two token interfaces share. We train only the student encoder to regress the teacher’s per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher–student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

July 9, 2026 research area Methods and Algorithms , research area Speech and Natural Language Processing

On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically…

July 1, 2025 research area Methods and Algorithms , research area Speech and Natural Language Processing conference ICML

We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a…

Discover opportunities in Machine Learning.

Our research in machine learning breaks new ground every day.

Sources

Summarized from the linked originals.

Related stories

Illustration for: The State of Chinese Physical AI
Models & Research

🎧 Audio version (or download audio for later): Good Morning, Dermot McGrath is an Irish entrepreneur and advisor based in Shanghai with over a decade in China. He founded ZenGen Labs, a research and strategy advisory firm focused on AI, robotics, energy and advanced.

AI Supremacy42 min
Illustration for: Gemini 3.8 text-to-speech says hello
Models & Research

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

Google DeepMind6 min
Illustration for: Building AI to accelerate science and improve lives
Models & Research

Building AI to accelerate science and improve lives We’re asking what’s possible for health, natural disaster and weather resilience, learning, and economic opportunity. SVP, Research, Labs, Technology & Society Your browser does not support the audio element.

Google AI Blog11 min
Illustration for: New insights from Google’s AI & Economy ATLAS
Models & Research

New data visualizations make ATLAS data easier to explore and use, while new research provides insights on how scientists are using AI. AI & Economy Lead, Chief Economist's Office Head of StratOps and Special Projects, Technology & Society Your browser does not support the audio.

Google AI Blog4 min