Skip to main content
Tag

audio

34 stories

Illustration for: The latest AI news we announced in September 2026
Models & Research

The latest AI news we announced in September 2026 Here’s a recap of some of our biggest AI updates from September, including Gemini 4 Argon, new Connected Apps in the Gemini App, and WeatherNext 3. Your browser does not support the audio element.

Google AI Blog8 min
Illustration for: Language Discrimination Improves Linguistic Learning in Mult
Models & Research

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models - Apple Machine Learning Research research area Speech and Natural Language Processing content type paper published October 2026 Language Discrimination Improves Linguistic Learning in Multilingual Speech Models Authors Maureen de Seyssel, Jie Chi*, Zakaria Aldeneh* Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total...

Apple Machine Learning2 min
Illustration for: On the Effectiveness-Fluency Trade-Off in LLM Conditioning:
Models & Research

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference EMNLP content type paper published September 2026 On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study Authors Iuri Macocco†, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni†*, Xavier Suau...

Apple Machine Learning2 min
Illustration for: The Communication Bottleneck: A Round-Trip Study of Tree-Str
Models & Research

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference NeurIPS content type paper published September 2026 The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models Authors Xavier Suau, Alex Ferrando de las Morenas,...

Apple Machine Learning2 min
Illustration for: Eleven v4 and Eleven v4 Turbo Text to Speech models - Eleven
Products & Tools

Eleven v4 and Eleven v4 Turbo Text to Speech models &)]:tw-overflow-x-hidden"> strong]:tw-font-normal tw-text-gray-600"> Introducing Eleven v4 Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12 Create controllable, expressive speech layered with emotion, audio events, and immersive soundscapes.

ElevenLabs8 min
Illustration for: Introducing Gemini 3.8 Live with Live Avatar
Models & Research

Introducing Gemini 3.8 Live with Live Avatar Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users.

Google DeepMind3 min
Illustration for: A Practical Recipe for Semi-Supervised Federated ASR: Online
Models & Research

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing content type paper published September 2026 A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization Authors Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan “Honza” Silovsky,...

Apple Machine Learning2 min
Illustration for: Compressing Streaming Neural Audio Encoders via Latent-Space
Guides

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing content type paper published September 2026 Compressing Streaming Neural Audio Encoders via Latent-Space Distillation Authors Prasanth Yadla‡, Mohammad Samragh Razlighi‡, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang, Chung-Cheng Chiu, Yongqiang Wang†**, Yuan Liu§**, Zhen Huang,...

Apple Machine Learning2 min
Illustration for: Google Beam expands with new regions, partners, and customer
Models & Research

Google Beam expands with new regions, partners, and customers We’re expanding Google Beam to five new countries, partnering with Industrious for an extended network, and proving our impact at Google and beyond. Director, Business Development, Google Beam Your browser does not support the audio element.

Google AI Blog4 min
Illustration for: Gemini 3.8 text-to-speech says hello
Models & Research

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

Google DeepMind6 min
Illustration for: Amazon shuts out Meta's Muse
Guides

PLUS: How to set up ChatGPT to write in your voice Good morning, AI enthusiasts, and welcome to our 4,369 new readers. Meta’s Muse AI assistant climbed to the top of the U.S.

The Rundown AI7 min
Illustration for: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Models & Research

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice.

Google DeepMind5 min
Illustration for: Building AI to accelerate science and improve lives
Models & Research

Building AI to accelerate science and improve lives We’re asking what’s possible for health, natural disaster and weather resilience, learning, and economic opportunity. SVP, Research, Labs, Technology & Society Your browser does not support the audio element.

Google AI Blog11 min
Illustration for: U.S. confirms weapons are in orbit
News

PLUS: A brain implant that decodes your body language Good morning, tech enthusiasts. The U.S. military confirmed for the first time that it has weapons in orbit.

The Rundown AI5 min
Illustration for: New insights from Google’s AI & Economy ATLAS
Models & Research

New insights from Google’s AI & Economy ATLAS New data visualizations make ATLAS data easier to explore and use, while new research provides insights on how scientists are using AI. AI & Economy Lead, Chief Economist's Office Head of StratOps and Special Projects, Technology & Society Your browser does not support the audio element.

Google AI Blog4 min
Illustration for: Proactive cyber defense for governments and enterprises
Models & Research

Proactive cyber defense for governments and enterprises Today, we’re launching our Fairwind Program, a limited access program for governments and trusted partners to use our most advanced cyber defense capabilities. Your browser does not support the audio element.

Google DeepMind3 min
Illustration for: Time to Speak Some Dialects, Qwen-TTS!
Models & Research

Here we introduce the latest update of Qwen-TTS ( qwen-tts-latest or qwen-tts-2025-05-22 ) through Qwen API . Trained on a large-scale dataset encompassing over millions of hours of speech, Qwen-TTS achieves human-level naturalness and expressiveness.

Qwen Blog2 min