
Identifying Interactions at Scale for LLMs
--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial…
New models, benchmarks, papers and research breakthroughs.

--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial…

Written by Nicholas Carlini, a researcher on our Safeguards team. I've been experimenting with a new approach to supervising language models that we’re calling "agent teams." With agent teams, multiple Claude instances work in parallel on a shared codebase without active human intervention.

An encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements.

Good evaluations help teams ship AI agents more confidently. Without them, it’s easy to get stuck in reactive loops—catching issues only in production, where fixing one failure creates others.

Editor’s note: The name of NVIDIA DRIVE Hyperion was changed to NVIDIA Hyperion in September 2026. All references to the name have been updated in this blog.

Unveiling what it describes as the most capable model series yet for professional knowledge work, OpenAI launched GPT-5.2 in December.

As AI agents become more capable, developers are increasingly asking them to take on complex tasks requiring work that spans hours, or even days. However, getting agents to make consistent progress across multiple context windows remains an open problem.

The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Connecting agents to tools and data traditionally requires a custom integration for each pairing, creating fragmentation and duplicated effort that makes it difficult to scale truly connected systems.

The partnership will help unlock faster iteration, accelerate workflows, and expand creative possibilities.

Update: We've published Agent Skills as an open standard for cross-platform portability. (December 18, 2025) As model capabilities improve, we can now build general-purpose agents that interact with full-fledged computing environments.

After a few years of prompt engineering being the focus of attention in applied AI, a new term has come to prominence: context engineering .

Between August and early September, three infrastructure bugs intermittently degraded Claude's response quality. We've now resolved these issues and want to explain what happened. In early August, a number of users began reporting degraded responses from Claude.

For more than a century, meteorologists have chased storms with chalkboards, equations, and now, supercomputers.

What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language…

Here we introduce the latest update of Qwen-MT (qwen-mt-turbo) via Qwen API . This update builds upon the powerful Qwen3, leveraging trillions multilingual and translation tokens to comprehensively enhance the model’s multilingual understanding and translation capabilities.

Here we introduce the latest update of Qwen-TTS ( qwen-tts-latest or qwen-tts-2025-05-22 ) through Qwen API . Trained on a large-scale dataset encompassing over millions of hours of speech, Qwen-TTS achieves human-level naturalness and expressiveness.

The evolution of multimodal large models is continually pushing the boundaries of what we believe technology can achieve. From the initial QwenVL to the latest Qwen2.5 VL, we have made progress in enhancing the model's ability to understand image content.

"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry…

Note: Much of the tooling landscape described in this post has changed since December 2024. For our current approach, see how we built Claude Managed Agents and the Managed Agents documentation .

What is the Role of Mathematics in Modern Machine Learning?The past decade has witnessed a shift in how progress is made in machine learning.

Connect 2024: The responsible approach we’re taking to generative AI Connect 2024: The responsible approach we’re taking to generative AI Today at Connect 2024, we shared updates for Meta AI features and released Llama 3.2, a collection of models that includes new vision capabilities as well as lightweight models that can fit on mobile devices.

For an AI model to be useful in specific contexts, it often needs access to background knowledge. For example, customer support chatbots need knowledge about the specific business they're being used for, and legal analyst bots need to know about a vast array of past cases.

LLM-based chatbots’ capabilities have been advancing every month. These improvements are mostly measured by benchmarks like MMLU, HumanEval, and MATH…

IntroductionImagine yourself a decade ago, jumping directly into the present shock of conversing naturally with an encyclopedic AI that crafts images, writes…