
Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.
New models, benchmarks, papers and research breakthroughs.

A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

A new corpus of annotated desktop observations, DeskForge, enables computer-use agents to achieve higher accuracy in complex desktop scenes.

The EditHero benchmark is introduced to evaluate and improve the reliability of long-horizon, part-level 3D editing.

The ComfyUI v0.39.0 release includes dynamic group widgets, new model nodes, and resolved minimax vae offload issue.

The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.

Fulcrum faces an identity crisis, even as it works on creating one by enabling anyone to imitate another person's writing style.

Subtopic Small Models · Vision Language · Quantization Takeaways − Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of Pro View the LinkedIn Profile of Matthias Bastian Starting in October 2026, Google will restrict access to its Gemini models for personal account users.

Subtopic Mixture Of Experts · Long Context Takeaways − Kolibri-1 is a 78B MoE with 3.46B active parameters, Apache 2.0, German and English focus. Context window validated up to 1,048,576 tokens, native 262,144, no position scaling tricks required.

This week, campus IT leaders voted AI their top priority for the first time, Cambridge refused Turnitin's new terms over AI training on student work, Dartmouth opened an investigation into its own provost's writing, and Ken Griffin gave Carnegie Mellon $3 billion. Student AI use is no longer a future scenario.

Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Subtopic Dpo · Fine Tuning · Distillation Takeaways − Cohere released North Small Translate , a 218B / 25B-active MoE translation model with open weights. Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).

Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

Another OpenAI safety departure adds to a pattern of researchers leaving with public warnings View the LinkedIn Profile of Matthias Bastian David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic.

Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity An essay for the Deepmind Institute challenges the familiar image of a lone superintelligence.

Takeaways − DeBERTa-v3-small is a 44M-parameter English encoder with a 128K vocabulary from Microsoft Research. Uses ELECTRA-style replaced-token-detection training instead of masked language modeling for better sample efficiency.

Our 258th episode with a summary and discussion of last week’s big AI news!Recorded on 09/26/2026 ; as usual, I am sorry this is coming out late, next…

Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap The first-of-its-kind Academy trains Frontier Deployed Engineers using the same standard of skills as Anthropic’s own engineers, starting with cohorts from Accenture, Bain, Capgemini, Commonwealth Bank of Australia, Deloitte, McKinsey, Morgan Stanley, Novo Nordisk and others.

Faced with growing opposition to datacenters from residents who don't want the bit barns in their backyards, Amazon plans to invest more than $1 billion over…

Runway said it is testing Praxis-1 on a variety of embodiments and environments to identify and close potential gaps before moving to general availability. | Source: Runway Runway AI Inc.

arXiv, the preprint paper repository, is seeing so many low-quality submissions, with AI tools helping drive the surge, that it has imposed a hard limit of…

Some big AI companies seem to think that the best way to keep models from becoming too chaotic or too mischievous is to keep them locked inside of labs. If only a chosen few can access them, the thinking goes, they can do less damage in the real world while researchers try to understand what they’re capable of.

Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge…

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.

The latest AI news we announced in September 2026 Here’s a recap of some of our biggest AI updates from September, including Gemini 4 Argon, new Connected Apps in the Gemini App, and WeatherNext 3. Your browser does not support the audio element.

Toward provably private learning from federated data Toward provably private learning from federated data Katharine Daly, Software Engineer, and Daniel Ramage, Research Director, Google Research We announce a new Federated Learning system that provides externally verifiable privacy guarantees while shifting computation to the server to improve training speed, accuracy, and device coverage.

Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for…

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models - Apple Machine Learning Research research area Speech and Natural Language Processing content type paper published October 2026 Language Discrimination Improves Linguistic Learning in Multilingual Speech Models Authors Maureen de Seyssel, Jie Chi*, Zakaria Aldeneh* Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total...

Limits of Confidence in Diffusion - Apple Machine Learning Research research area Methods and Algorithms content type paper published October 2026 Authors Russ Webb, Amitis Shidani, Alice Bizeul, Dan Busbridge Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from...

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users.

Quick Apps could already bring live data into an app from connectors and content sources: action connectors (services like Jira, Slack, and Google Drive),…

Barclays scales Claude to upgrade operations and improve client experience Barclays, the British universal bank, is expanding its strategic collaboration with Anthropic to integrate secure, enterprise-grade AI systems across its global operations. Barclays is extending Claude across the bank to accelerate software development, modernize legacy systems, and improve operational efficiency.

Summary: In this guest post, Prof. Matthew Schwartz returns to describe a new approach to AI-accelerated science. In Vibe Physics , Schwartz discussed similarities in capability between Claude and a physics graduate student.

Gemini. It only took Google seven months (February to September) to release their next frontier model.
![AI-generated editorial illustration for: [AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output](/images/articles/ainews-gemini-4-argon-gdm-s-answer-to-astra-fable-with-1m-output.webp)
GDM last shipped a larger-than-Flash model in February (3.1 Pro), and after successive incremental 3.x Flash versions and the big GDM management shakeup last…

Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages | NVIDIA Technical Blog Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages By Imane Khaouja , Amine El Khair , Meshari Alaeena , Zahra Al-Kaf and Abdulrahman Alkhamees NVIDIA Nemotron 3.5 ASR supports multilingual streaming transcription across 40 language-locales, but deployment-specific dialects...

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? - Apple Machine Learning Research research area Methods and Algorithms content type paper published October 2026 How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Runway app for iPhone Runway app for Android An open-weight world action model that turns Runway's video pretraining into control for real robots. “Pick up the tennis ball and put it in the box.” Today we're announcing Praxis-1 , our first open-weight world action model.

Google promised Gemini 3.5 Pro in June, but it spent the summer trotting out smaller Flash models.

Gemini 4 Argon: our next era of frontier intelligence Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. SVP, Google DeepMind and Chief AI Architect, Google Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.

Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK | NVIDIA Technical Blog Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK By Harish Arora , Vikram Sharma Mailthody , Kiran K.

We present a robot exposure index based on how well robots can perform job tasks today. Robots, which we define as autonomous physical machines that sense and act, can perform three-quarters of physical tasks in the US, making up 34% of working hours, but mostly in limited settings.

At a glance End-to-end forecasting: A machine learning pipeline uses forecast-time solar-wind information to generate location-specific risk estimates for…

Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still…

How well does nDCG capture the quality of modern retrieval systems? nDCG is the established yardstick for measuring how well search systems rank results, and is widely used across benchmarks such as MTEB and BEIR .

State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested, particularly across financial datasets, parsed PDFs, and visually rich documents. A new Fast tier: Embed 5 Fast brings strong retrieval quality to latency, and cost-sensitive workloads, at $0.08 per million tokens.

Top NewsAnthropic and OpenAI race to release smarter and cheaper modelsSources:Anthropic launches Claude Opus 5.5 with stricter safeguards for…

🎧 Audio version (or download audio for later): Good Morning, Dermot McGrath is an Irish entrepreneur and advisor based in Shanghai with over a decade…

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference EMNLP content type paper published September 2026 On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study Authors Iuri Macocco†, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni†*, Xavier Suau...

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation - Apple Machine Learning Research research area Methods and Algorithms , research area Tools, Platforms, Frameworks content type paper published September 2026 SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation Authors Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu Continual-learning agents are systems of models, harnesses,...

Salesforce, the #1 AI CRM, today announced it has signed a definitive agreement to acquire Listen Labs, an AI-powered customer research and human simulation…

How Diffusion Controller unifies and simplifies AI image generation How Diffusion Controller unifies and simplifies AI image generation Chih-wei Hsu and Moonkyung Ryu , Software Engineers, Google Research We introduce Diffusion Controller, a lightweight "steering damper" network that precisely steers image generation to achieve significantly better prompt alignment.

We’re launching a new study using Anthropic Interviewer to learn from your experiences with AI, and we’d like you to participate. After you finish, you can decide to make your interview public, so that anyone, not just Anthropic, can read and learn from it.

GLM-5.3 and the spread of advanced cyber capabilities Cole McFaul, Robert Xiao, Tripp Gallagher Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

At a glance Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models,…

Hi folks,I’ll be at OpenAI DevDay today. If you’re there, come say hi - I have pink shoes on.Sam says they've "found a new thing".I have the…

The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.While Jev aims to classify things,…

AI Anthropic's mid-tier Claude climbs the rankings PLUS: Pick the right Claude model with one quick test Good morning, AI enthusiasts, and welcome to our 5,342 new readers. OpenAI takes the stage today for one of its most hyped days of the year.
![AI-generated editorial illustration for: [AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more](/images/articles/ainews-amd-buys-world-labs-for-8-2b-as-atlas-solves-sparse-reconstruct.webp)
The official post is shy, but since AMD is public, we know the purchase price. We covered them less than a year ago:Fei Fei has a lovely reflection blogpost…
![AI-generated editorial illustration for: [AINews] Opus 5.5 is good at explainer videos](/images/articles/ainews-opus-5-5-is-good-at-explainer-videos.webp)
Opus 5.5 shipped this week but the vibes are overwhelmingly positive:And specifically it took over the timeline for explainer videos:AI News for…

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference NeurIPS content type paper published September 2026 The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models Authors Xavier Suau, Alex Ferrando de las Morenas,...

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

Warning : Undefined variable $stocks in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : Undefined variable $funds in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 472 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d...

On July 24, 2025, Microsoft Research Asia – Singapore (MSRA – Singapore) opened its doors as Microsoft’s first research lab in Southeast Asia.

Mistral Opens German Hub in Munich to Advance Industrial AI in Europe’s Largest Economy At Mistral, we have always believed that the most consequential AI applications will be built where real industrial problems are solved. Today, we are putting that conviction into practice by opening our new hub in Munich.

A line of text can change significantly depending on how it’s spoken. “I need you to stay calm" should sound different depending on who's saying it, whether that's a doctor delivering it gently to a frightened patient, or a character in a game shouting to his squad before dropping into battle.

Building Production Agents with Jev and LangGraph Last week, TypeSafe AI released Jev, a new kind of model. Unlike traditional LLMs, Jev doesn't generate text.

Add Runtime Controls to AI Agents with NVIDIA OpenShell | NVIDIA Technical Blog Add Runtime Controls to AI Agents with NVIDIA OpenShell NVIDIA OpenShell 0.1.0 provides an open-source runtime that enforces which systems and data an AI agent can access without rewriting the agent.

Faster Rates for Federated Variational Inequalities - Apple Machine Learning Research research area Methods and Algorithms conference NeurIPS content type paper published September 2026 Faster Rates for Federated Variational Inequalities In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years.

Compass is Cohere’s retrieval platform for developers building AI applications with their enterprise data. It surfaces the most relevant information from your company’s corpus for use in retrieval-augmented generation (RAG), search, and agentic workflows.

Good Morning, As Anthropic gets set for a $2 Trillion IPO in just over a month, it’s getting close to unleashing its version of RSI in biology, drug…
![AI-generated editorial illustration for: Claude’s New addTools() Can Reuse 98.7% of Your Next Request. Editing tools[] Reuses None.](/images/articles/claude-8217-s-new-addtools-can-reuse-98-7-of-your-next-request-editing.webp)
Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI.

Author(s): Quan Huynh Originally published on Towards AI. Build an AI Agent Evaluation with JEV Build a small eval harness for a tool-using AI agent: code…

In this guest post, physicist and science writer Matt von Hippel shares what happened when he issued a challenge to AI companies regarding a problem in his former subfield of theoretical physics. It’s not often that you issue a challenge, only to see it beaten a month later.

Last Updated on September 25, 2026 by Editorial Team Author(s): luisacsfreitas Originally published on Towards AI.

Last Updated on September 25, 2026 by Editorial Team Author(s): luisacsfreitas Originally published on Towards AI.

We’re building personal AI agents for everyone and a whole family of devices that let you connect with them from anywhere.

Automating coherent long-form video generation Automating coherent long-form video generation Yale Song and Yiwen Song, Research Scientists, Google We introduce a unified multi-agent framework that autonomously generates temporally consistent, long-form video narratives, overcoming the identity drift and cascading failures of current linear AI pipelines.

If you're new to fuzzing and want to learn the fundamentals first, check out our Fuzzing 101 course at gh.io/fuzzing101.
Introducing Gemini 3.8 Live with Live Avatar Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users.

Use the LangSmith Fine-Tuning CLI and skill - smithtune - to leverage agent trajectories stored in LangSmith to a fine-tuned model in one end-to-end workflow In addition to driving training, the smithtune CLI handles evaluating the fine-tuned model and uploads eval results to LangSmith for easy analysis We partnered with Fireworks & Baseten to create a seamless link between LangSmith...

When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well…

Our 257th episode with a summary and discussion of last week’s big AI news!Recorded on 09/19/2026 ; as usual, apologies for the none…

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing content type paper published September 2026 A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization Authors Wonho Bae, Zakaria Aldeneh, Martin Pelikan, Jan “Honza” Silovsky,...

Google Beam expands with new regions, partners, and customers We’re expanding Google Beam to five new countries, partnering with Industrious for an extended network, and proving our impact at Google and beyond. Director, Business Development, Google Beam Your browser does not support the audio element.

At a glance Challenges a core assumption in robotics AI: Our research shows that running physical AI inference exclusively on onboard GPUs can limit robot…

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.

NVIDIA AI Day Singapore, which takes place Sept. 22-23 at the Raffles City Convention Centre, is offering attendees opportunities to explore the hands-on…

Jev by Typesafe AI. ShareGood Morning I don’t usually nerd out on non-LLM machine learning announcements. But extraordinary claims have been made.

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

Rollouts move live traffic from your current model to a new checkpoint in gated steps. Health checks always run before traffic moves; on a canary you can also add metric gates (say p95 latency or error rate) that run after each step.

Set up Grok Bot, create your first Bot, and hand it a useful task.

x�eRKN1 ��� &�c;>G@�����Kڂ(i"�>v�:�'�Om����>Z���7z~����BU� �MgpU�d����W|���g `���g�l[Ҝ�@� TwX δmN� > xڵS=o! ��+��` ���j�[�۪�].ST����H�RuH��;����� �����|��X[�O���x3�n�� �j�mU"�3���\&B�� ��pPu�k���^�l���a���5�ػނ��"����V�k11e-E]�j�'�=��fiI �ĵ�S��*���Ul�0U�����i��~ ��{�/��5 x��\Mo�6 ��W�X�o`�C��@o��V�0��t�K�~)��%���f=�L`$���G���#)9��I?�8�_ٜ>5RH���]8�����SM��R���������x���ӵQ�6�������)�v��t`��cע�$��Z0t�d�!�8�?��[�?;���p�s��)H�s���[��Ӎ��Ҟ�Š-��VJ&k� 0��Fӈ�9];�C?��-.���ù>���������W�� �s#Lp��/ ��`��V =��i�h>����2�h6cEP>� @�ݗ"�.�I��@��|�zp�]c3�������Gٷ�']曱SX'���A�W^��׆@ �K;'���6���� �� �K��T��F�tΛ� M� M%����H����%1�r:i�Y5z*�c�A�����i#B(�N ѓJ!]�֦�u�Z� ��TGl���|����}%�z�� �T7U1y1.&FB�U>~����J��6�[���S�b7��Ђ���Թ_��:6����J����/WF)�5_�f��xru|�֘ Wb�%���v�y�e���J�����a5��s"h�����a��p�k��+��w�|�M���|FK�ojFLUV��z�� ��Ϧ�85/�@���|�����A��t�[H�w�7-{Pr`��2�BqUE����� `�f�b����/Po5�PB�� ɷԭ��Zz1�P�)?_4�V�� > x��Zˎ�6��+��v\.���HI���.���U���"~�m�&3-z�3ͨ1�@Q�9uʦ��A'�tV�ٝ�餐��?~�����W��^�_����6}a�O���_�s����Ca:�¨�x������A�z���AY�I�� V��ǎ!�q4�����1N:��>vA�P�koN���[���I�z�H�H᥇�2��{�;�D>�-z�k��� _IF�)4,J���c��Wb�W"=- L� > xڽ[M��6 ��W��+R�-E����[у�ؽt�K�~%K�D�J��d���ر�A=>>R��k���������K'z!

MilleMiglia: A realistic instance generator for middle-mile logistics MilleMiglia: A realistic instance generator for middle-mile logistics Aymane Lotfi, Software Engineer, Ads & Commerce, and Thibaut Cuvelier, Software Engineer, Google Research MilleMiglia bridges the gap between academic theory and industrial logistics by providing open-source, realistic benchmarks that allow researchers to optimize complex middle-mile networks, ultimately leading to more robust and efficient...

PLUS: OpenAI rescues biotech’s lost research Good morning, tech enthusiasts. Meta may have found the easiest way to make smart glasses less creepy: take out the camera.

New experts join Google’s AI & Economy team Nobel Laureate Philippe Aghion, Professor Ajay Agrawal, and leading researchers join Google’s AI & Economy program to expand our scientific understanding of AI’s impact on economic activity worldwide.

Co-creating the future of fashion with Google Two fashion designers created custom tools in Google Flow to help with set design and styling for New York Fashion Week. Your browser does not support the audio element.

Skild AI has made a breakthrough. Good Evening, Visions of AI is a new feature format I’m experimenting with that will amount to a short profile on an…

The future of practice: Enabling teachers to create learning interactives with generative UI The future of practice: Enabling teachers to create learning interactives with generative UI Gal Elidan, Research Scientist, and Yael Haramaty, Product Manager, Google Research We explore how we can harness generative UI with learning design guardrails to give teachers the ability to generate guided, interactive simulations for...

SN50 Premium Inference: A Faster ROI for Neoclouds Faster Returns for Neoclouds: SN50 Delivers ROI in 6 Months AI agents are creating a premium inference business for neoclouds. Coding agents read a repository, generate changes, run tests, and return to the model with the results.

Top NewsOpenAI Says It Has Cracked One of Math’s ‘Millennium Problems’Related:On the Navier–Stokes Millennium Prize ProblemQuartz:…

Introducing the Life Sciences Verification Program Today, we are introducing the Life Sciences Verification Program (LSVP), which gives life science professionals access to our Mythos, Opus, and Sonnet models with a refined set of safeguards more permissive for biology-related work.

Mistral and Mozilla are bringing open, private and multilingual AI to your web browser Today, we are announcing a partnership with Mozilla to bring privacy, control and choice to people using AI to browse online. Firefox Smart Window (beta), Mozilla’s AI browsing assistant, is now powered by Mistral models.

We are announcing the signing of a definitive business combination agreement with Aleph Alpha, following the release of our planned partnership in April of this year. Operating globally as Cohere, the unified company will effectively create the first transatlantic sovereign AI solution.

Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train Pengcheng Jiang, Student Researcher, and Judith Yue Li, Senior Research Engineer, Google Research Instead of relying on expensive inference-time reasoning, the Retrieve-for-Train framework uses reinforcement learning once to train a lightweight diffusion model.

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice.

We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed. SVP, Research, Labs, Technology & Society Google is using AI to help people communicate in hundreds of languages, including those that were previously left out of technology.

Building AI to accelerate science and improve lives We’re asking what’s possible for health, natural disaster and weather resilience, learning, and economic opportunity. SVP, Research, Labs, Technology & Society Your browser does not support the audio element.

From error mitigation to fault-tolerant quantum computing | IBM Quantum Computing Blog The continuous path from error mitigation to fault-tolerant quantum computing A spectrum of error-correcting techniques is enabling useful quantum computation, measured not by logical qubits but by the circuits you can run with them.

New insights from Google’s AI & Economy ATLAS New data visualizations make ATLAS data easier to explore and use, while new research provides insights on how scientists are using AI. AI & Economy Lead, Chief Economist's Office Head of StratOps and Special Projects, Technology & Society Your browser does not support the audio element.

Dario Amodei, “(P)doom vibes” pre IPO must have been his idea. CEO, Anthropic.

Turning an open-weight model into a high-performing model for your task takes a sequence of well-measured experiments. Teams need to understand what the model will train on, follow how each run is progressing, adjust the training recipe, and identify which checkpoint performs best.

ToolGrad: Efficient tool-use dataset generation with textual "gradients" ToolGrad: Efficient tool-use dataset generation with textual "gradients" Zhongyi Zhou, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR ToolGrad is a data generation framework that reverses the traditional paradigm by first generating tool-use answers before user queries. We show this design enables LLMs to achieve better tool-use performance.

Switzerland's first IBM Quantum System Two | IBM Quantum Computing Blog Switzerland will soon get its first IBM Quantum System Two Lockheed Martin and IBM are launching a Swiss quantum innovation hub at ETH Zurich to advance research, industry collaboration, and workforce development.

Multilingual Reasoning Image Inputs Safety Modes Citations Tool Use Structured Outputs For both trial keys and production keys, North Small Translate is free until rate limits are reached. Learn more about rate limits for different models and key types here .

Modernizing complex legacy code with AI agents. Legacy code modernization is challenging, especially when migrating complex Fortran 77 systems to modern C++. Mistral successfully migrated 40,000 lines of a physics-intensive reservoir simulator by building a parity harness for numerical verification, documenting the codebase with AI agents, and using structured workflows with human oversight to ensure high-quality, maintainable code.

A lot has happened in the last few weeks. I am sure that OpenAI’s GPT-6 Astra is top of mind for everyone right now.

SPONSORED BY ODSC AIODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal…

Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

Cleveland Clinic, RIKEN, IBM named Gordon Bell finalists | IBM Quantum Computing Blog Cleveland Clinic, RIKEN, IBM named Gordon Bell finalists Finalist recognition for one of supercomputing’s top prizes arrives as researchers report new progress in automated quantum-HPC chemistry workflows. Cleveland Clinic, RIKEN, and IBM were named 2026 ACM Gordon Bell Prize finalists for breakthrough quantum-HPC chemistry research.

What Is AI Inference? Meaning, Benefits & How It Works The word “inference,” in English, means a conclusion drawn through reasoning and evidence. Similarly, AI inference relates to an AI model’s ability to infer, or extrapolate, conclusions in new situations, using information gained from training, response, and the fine tuning process.

Mistral raises €3B to make sovereign, open-weight AI the technology frontier Mistral today announced that it has raised €3 billion in a Series D funding round at a post-money valuation of more than €21 billion, the largest equity fundraising round ever completed by a European technology company, three years after the company's launch.

Top NewsGPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI EraRelated:OpenAI begins rolling out Astra model after warning of its advanced…

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers.

Providing developers the best model for the task at hand has always been our goal.

Transfer learning for genomic prediction in underrepresented populations Transfer learning for genomic prediction in underrepresented populations Joey Poomarin Phloyphisut, Software Engineer, and Cory McLean, Senior Staff Software Engineer, Google Research We evaluate ways to improve cross-population genetic risk prediction and find that while transfer learning from European cohorts improves prediction in small underrepresented populations, it degrades accuracy once target cohort...

A connectomics milestone: Mapping the complete male fruit fly brain A connectomics milestone: Mapping the complete male fruit fly brain Michał Januszewski and Viren Jain, Research Scientists, Google Research We partnered with HHMI Janelia and collaborators in Cambridge, U.K., to publish a complete map of the male fruit fly’s brain and central nervous system, creating the largest brain map to...

Introducing WeatherNext 3, our most advanced and accurate global weather AI model Our flagship AI weather forecasting model now includes real-time satellite data, hourly refreshes, higher resolution, precise precipitation forecasting, and clean energy variables. It’s now integrated across Search, Gemini, Maps, Google Maps Platform, and Cloud.

Proactive cyber defense for governments and enterprises Today, we’re launching our Fairwind Program, a limited access program for governments and trusted partners to use our most advanced cyber defense capabilities. Your browser does not support the audio element.

Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity.

We’ve built an AI agent that acts as a secondary expert for a given domain, making deep specialist knowledge readily available and preserved for anyone in an…

Mapping global methane emissions from space with deep learning Mapping global methane emissions from space with deep learning Vishal Batchu, Research Engineer, and Michelangelo Conserva, Research Scientist, Google Research The Methane Analysis and Plume Localization with EMIT model is a deep-learning framework that automates the detection, enhancement quantification, and source estimation of methane plumes globally, turning raw satellite data into...

At a glance The Flash family extends GigaPath and GigaTIME with dramatically improved efficiency, making large-scale pathology research more accessible and…

IBM Quantum Nighthawk r2—more circuits, faster | IBM Quantum Computing Blog IBM Quantum Nighthawk r2—more circuits, faster High-speed, independent qubit reset boosts circuit throughput 25x over Heron while enabling accurate computations on circuits with 7,500+ gates. IBM Quantum Nighthawk r2 uses independent, high-speed qubit reset to execute 100,000+ circuits per second—up to 25x higher circuit throughput than IBM Quantum Heron.

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers.

One University of Chicago curriculum is removing AI-assisted writing from the classroom. Alpha School is expanding a model that puts adaptive software at the center of the academic day.

Total funding reaches $232M under new leadership with investor group comprised of Electronic Arts, Sony Music Group, Universal Music Group, Warner Music…

MTIA 300 is the first of Meta’s family of in-house training and inference accelerators optimized for training ranking and recommendation models.

Over the past few weeks, we have announced several important steps to expand sovereign AI by increasing access to local infrastructure capacity. Earlier this summer, we announced an expanded partnership with Microsoft to increase compute capacity in Europe.

At a glance Skala 1.1 demonstrates the continuously improving nature of Microsoft Research’s deep-learning DFT approach: trained on 2.5× more data than…

Agentic Search. More accurate and efficient results from your AI systems. Mistral Agentic Search delivers more accurate search results while reducing turns, token use, and latency against FinanceBench and OfficeQA Pro benchmarks.

If you use AI at work, the tools you rely on could change again before February. OpenAI, Google, Meta, Anthropic, several Chinese labs, and a group of world-model startups are all preparing or rumored to be preparing new releases.

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers.

Building Foundry A practical series on creative workflows, semantic search, and Weaviate.

At a glance MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as…

BMG and Suno have reached a licensing deal allowing BMG s recorded and music publishing works to be used to create derivative works on the platform, the companies jointly announced on Tuesday (Aug. 11).

In-region inference, open models, and new European infrastructure for sovereign AI. Mistral is advancing AI sovereignty by offering enterprises and countries control over AI models, infrastructure, and compute capacity, ensuring regional compliance and reliability.

Scaling test-time compute has become one of the clearest successes of the foundation model era.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining.

SambaRack SN50 Benchmarked on MiniMax M2.7 by SemiAnalysis SemiAnalysis Benchmarks SambaRack SN50 with Fast Inference on MiniMax M2.7 MiniMax M2.7 is a model used by many of our customers around the world that helps augment their coding and agentic workflows using the fast inference speed of SambaNova’s SN40 to accelerate their tasks.

Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied…

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers.

.abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; } .abbel-fig img { display: block; margin: 0.65em auto 0;…

Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell…

Your Prompts and Skills need a system of record. Most enterprises struggle with unmanaged, scattered AI prompts and skills, leading to inconsistent behavior and untraceable issues.

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration…

... government of the people, by the people, for the people ... — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly.

Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D.

As some of you know, I have the long-running habit of keeping a running list of research papers I want to read, revisit, or cite in future articles and…

Meet Stable Audio 3.0, the model family built for artistic experimentation with open-weight models We're releasing Stable Audio 3.0 , a model family with open-weights music models that are trained on fully licensed data.

After a short family break, I am excited to be back and catching up on a busy few weeks of open-weight LLM releases.

Grok Build: SpaceXAI's Coding Agent | SpaceXAI Docs Grok Build is a powerful and extensible coding agent. Use it via an interactive TUI, headlessly in scripts or bots, or through the Agent Client Protocol (ACP) in other apps.

.grasp-results-table table { font-size: 0.875rem; line-height: 1.35; width: 100%; } .grasp-results-table th, .grasp-results-table td { padding: 0.35rem…

Introducing Muse Spark: Scaling Towards Personal Superintelligence Introducing Muse Spark: Scaling Towards Personal Superintelligence Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

In this article, I want to cover the overall design of coding agents and agent harnesses: what they are, how they work, and how the different pieces fit…

--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial…

Written by Nicholas Carlini, a researcher on our Safeguards team. I've been experimenting with a new approach to supervising language models that we’re calling "agent teams." With agent teams, multiple Claude instances work in parallel on a shared codebase without active human intervention.

An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements.

Good evaluations help teams ship AI agents more confidently. Without them, it’s easy to get stuck in reactive loops—catching issues only in production, where fixing one failure creates others.

Editor’s note: The name of NVIDIA DRIVE Hyperion was changed to NVIDIA Hyperion in September 2026. All references to the name have been updated in this blog.

Unveiling what it describes as the most capable model series yet for professional knowledge work, OpenAI launched GPT-5.2 in December.

As AI agents become more capable, developers are increasingly asking them to take on complex tasks requiring work that spans hours, or even days. However, getting agents to make consistent progress across multiple context windows remains an open problem.

The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Connecting agents to tools and data traditionally requires a custom integration for each pairing, creating fragmentation and duplicated effort that makes it difficult to scale truly connected systems.

The partnership will help unlock faster iteration, accelerate workflows, and expand creative possibilities.

Update: We've published Agent Skills as an open standard for cross-platform portability. (December 18, 2025) As model capabilities improve, we can now build general-purpose agents that interact with full-fledged computing environments.

After a few years of prompt engineering being the focus of attention in applied AI, a new term has come to prominence: context engineering .

Between August and early September, three infrastructure bugs intermittently degraded Claude's response quality. We've now resolved these issues and want to explain what happened. In early August, a number of users began reporting degraded responses from Claude.

For more than a century, meteorologists have chased storms with chalkboards, equations, and now, supercomputers.

What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language…

Here we introduce the latest update of Qwen-MT (qwen-mt-turbo) via Qwen API . This update builds upon the powerful Qwen3, leveraging trillions multilingual and translation tokens to comprehensively enhance the model’s multilingual understanding and translation capabilities.

Here we introduce the latest update of Qwen-TTS ( qwen-tts-latest or qwen-tts-2025-05-22 ) through Qwen API . Trained on a large-scale dataset encompassing over millions of hours of speech, Qwen-TTS achieves human-level naturalness and expressiveness.

The evolution of multimodal large models is continually pushing the boundaries of what we believe technology can achieve. From the initial QwenVL to the latest Qwen2.5 VL, we have made progress in enhancing the model's ability to understand image content.

"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry…

Note: Much of the tooling landscape described in this post has changed since December 2024. For our current approach, see how we built Claude Managed Agents and the Managed Agents documentation .

What is the Role of Mathematics in Modern Machine Learning?The past decade has witnessed a shift in how progress is made in machine learning.

Connect 2024: The responsible approach we’re taking to generative AI Connect 2024: The responsible approach we’re taking to generative AI Today at Connect 2024, we shared updates for Meta AI features and released Llama 3.2, a collection of models that includes new vision capabilities as well as lightweight models that can fit on mobile devices.

For an AI model to be useful in specific contexts, it often needs access to background knowledge. For example, customer support chatbots need knowledge about the specific business they're being used for, and legal analyst bots need to know about a vast array of past cases.

LLM-based chatbots’ capabilities have been advancing every month. These improvements are mostly measured by benchmarks like MMLU, HumanEval, and MATH…

IntroductionImagine yourself a decade ago, jumping directly into the present shock of conversing naturally with an encyclopedic AI that crafts images, writes…

The AI revolution drove frenzied investment in both private and public companies and captured the public’s imagination in 2023.

AI models reflect, and often exaggerate, existing gender biases from the real world.

The rise of the vector databaseAs a result of the rapid advancement of generative AI in recent years, many companies are rushing to integrate AI into their…

Skip to content Embed 5 is here: state-of-the-art retrieval - now available in Pro and Fast tiers.

System cards are designed to help people better understand how our artificial intelligence (AI) systems work. Multimodal generative AI systems can accept multiple types of inputs, such as text and images, and produce various forms of output.

Skip to content Embed 5 is here: state-of-the-art retrieval - now available in Pro and Fast tiers.

Meta’s progress and learnings in AI fairness and transparency Meta’s progress and learnings in AI fairness and transparency While AI has brought huge advancements to humanity and our planet, it also has the potential to cause unintended consequences, and technology companies must proactively work to mitigate these issues.

> /W [ 1 3 1 ] /Index [ 232 196 ] /Info 39 0 R /Root 234 0 R /Size 428 /Prev 539374 /ID [ ] >> x�cbd`�g`b``8 "��ٍ ����"�Ad�,��� ���� ��Q)������#X/�(9�I�U�Qr(� � � x�c```b``������� � `6 �)00�Y2�cl)��&��ak@ ӑ���f00��h)uq.3 o[}V�H����&��&�ɮ��0% �@������9��,������֕]� xڍZK��F��W�UR��fw�ͦ┳ym | '�=��ݷ/B�����o�S;f�4̂LŻ���M��L�g�s���~����^��Z�����*0I�K�,�c��/w��ߜ������A��U�x��q���6��}t����t]?ئ�����$�ݩ ��%L��}t�P� w�Ҟ]�#Ho�`|CUp/���mW et�����p�O��.E�+�b�i��*j,T�5�D wI��%éO4(��=�y�a�����jWr�tN~+���,���L�ǃ��,��)13��Ƌ9o�c�/���^�ڎ�jg�F��>���c�� >ϐ+��J~��v�{I/f���e���'[y�;h[��@:݊�j�X,ڭ���U5 t,�TY(,L�:���{�h���u���/�t��v��I�����{�$��}��mC.6H�p�re9]���m�NT� K�x������Ԟ �~'�F%����}�������?��Ѐ@��/�{!����q�� ����C�AVG�6U��4y���X,o2��hsB�r�� �F�TT �RYgWyF˨ ���� f�+Ol2��vDE6�s��v��on-����n�y�V��s�}z�Z���� 6 �=���'.�D�hcW���?�#��֨~� �^fgH=�@3��KP8{��n*2��h�D���g�9�*�}+]��p��J�ѹf�!kH�d X��L�݂�B�wW��� �=�eF�3ʉ�/�7w>���o��LQ���b*�n�5F!�^�z�^G'A�VL~� ��45�[&Bzo�ՈH_���d /�Q��t��u�%e�9����}A�w1�e�֗�C�5� !Q�����.#���Re[ѣ8...

x�ٲ��q�y�O�\$U��A��%�ؐ& -�� �4��.t�Q/1�2�����v�*��"��y�"���X|��n�ߖ�-��߭���n�������mN��r����}����]��q�mv^��N����� ?��_��������z��7˟�G���߁��p��Nǽ}yZo�O�մ\Oǧ�j}X���z�^|���'�����?�����/�������ۯ߾��7_���o>�q'�����ݴ_N�txڟ6�o�������n��a�t�u3 ǧi�!���~wd��;�����{�n5��i�����MO�v��t8 �5�����N���i-��i�������N��͊�����z{|�o�W�G���n�MA"��p(�7��ݴ�8��z���P8�b$?짣 >�5��hۅ0����tZ�kN� y�� E��z��4���n�;@�I�=��z��q�|�[=mv���]����]�SYowڊ�Ԅ�[x��sA�$�������6M,Æe ��b�_�DA �˞x�[|o�Tj�Ԅ��&K��~-���w����]S�ϖ_5� � �x����v��~j�s;o����\�?���8v ����+�ôa +���f���6��͆��F�ܳ�����[��?lWO� ��i{X�����n�[r�v���� ���q�~^&m׃���a==폇'� ��v�]Q�n���k�b���j�1lc�i�� ^C+��a��q�V��i}�Q@|rg�b2�'���3��+6�jo�A��u�1��VB�1wn�hC��v��x9Մ-� Yр$I�%��E��K��"& ��� ��tZ�'w��u�&�)��}���X�E܈1�\�B�~�?� �3}�)��~��������ǫ�i���O���_���"~�[�>Y����e��O�s�0�n2��jz�Jv4�i�a��\��ӑ��LsZ ��CZ\���;M($]��~,�5 O'�CA���5�� ��|������T�b�⅓�����Bqq�݆���O�-�7�T$��ĩ�k��p�틧2�^[c3�����}�6$�F�6uɲc 6��l� N��ns���u���>Gu&��r3H#F ����K]Ȩ*��"N;�T稛���g9*w���G]#jqኡ� �I��3��:t8�DQ��#�em��4� �j�%���T�L�2P$5 �=p?B_E���3�������8 ��`/C�5SB>��OGp����y?]��S�}��~�1�q�Z��*M8�q ��, Q�@`:-k2j�e���a�U�eR� E���>&�mZ�{���#L�Z�r&Ú��/b"�sL��ls�@Ҥ�n��F��m+�v�M���5TlySJ�?f����S��䰊}Ƥa����YR�C���*?_$�����d�J�Id��,[)u�ӆ�0��'T�A-�A����V�䟒r��z��x}t��D��$t� � |�P�(���z��y�ɈY�j���v�l�/( Em����Jm��*�\)�~�>�P��Q=�� A 5���1-��lZy�9��1Q'��tz�O�1+\SK�� �Z�B�!&Z�ߨ����ZZ[�qi��|��y�����k�}�Z�y_���͙(&��]N�S�ֵ�q�Z4��������a%��F}ܘV���$;���q�O���� �> ��I��z���GTLqf�6�q����{ {�WX ��� �a\Q�d>k�_�s� �*4� ��!�-�|�,���%w���V�a�=�|}��P��qQT�%�;Trl��EMj��\�=���t�jO�, Rdk�Yk��Y� ��O�{��2�_�B!�[c��'j2�kjʵ�$S:] ٮM�iD�5Ĭ��=������=���M��\|t��v����'� RHv���� �O�6l1��d�WC/�8L'tMn��J�t����6m�Kȷ]����8�ͬ���Е ��&x��N�~\L����}y'�f7�ʤ�>�U��dW=��u�ޗubG9N+��ʭ��O �9�D�� ���V0�T�m�kf+ �$�E '�~��z�+�M/�I'� {�+A����� �aw� d�=SmpR�`�C?.� Gcm�G;�sZ���X\���3d�sp�2_�6\�1E��-)�n̆���m�U圡�jw@�9;e �CܼPԩ7B?XB�Gr-�m��`�(D��s�6����2q�a�*��K�P�o��z62�i�9����t��Y�␠ �x�ejs�V�Y[q�q+7 �4+�c�j?;��'��Cr6@l-$Rڜ��^��E�ڑE��@cٜؖb��~)�W�+\��9W��;J�D.`�{0*��IvN�x����tң�����t���!��USI���� ɓ]_zNCʤ�mR%�-��L���ڕIw��I7q���M�C�@��Lm*b=�|��w�=5Ɂ�,�#�F E�ijH�Si�HL�x�C�uRJ?���:I�����̆!����Sh����j����σNR�+��[���㨃���f��tE3�Ќ�u`i�l��u&��F��...