
DeskForge Corpus Improves Computer-Use Agent Accuracy
A new corpus of annotated desktop observations, DeskForge, enables computer-use agents to achieve higher accuracy in complex desktop scenes.

A new corpus of annotated desktop observations, DeskForge, enables computer-use agents to achieve higher accuracy in complex desktop scenes.

Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

Google researchers find a way to keep self-improving AI agents from memorizing their tests AI agents that keep optimizing their own working environment quickly tend to overspecialize on their test tasks. A new method from Google Cloud AI Research and several universities aims to prevent that while also cutting compute costs.

Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Subtopic Dpo · Fine Tuning · Distillation Takeaways − Cohere released North Small Translate , a 218B / 25B-active MoE translation model with open weights. Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).

Once just an element of science fiction, revelations of AIs gone rogue are increasingly everyday occurrences.

All the AI agents that can live in your text messages | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

Welcome to Kernel Panic! , a weekly newsletter by Lily Hay Newman and Matt Burgess from inside the new world of privacy and digital security.

Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity An essay for the Deepmind Institute challenges the familiar image of a lone superintelligence.

Meta wants your next gadget to be Muse-infused | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

Apple says it is changing its macOS privacy settings to stop third-party app developers from misusing them to access message histories.

Apple will add new limits for "full disk access" on Mac in response to risks posed by AI agents, as reported earlier by TechCrunch.

Apple says it s tightening macOS Full Disk Access controls due to new risks from AI agents | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

New helpful little guy just dropped. | Photo: Allison Johnson / The Verge It's a tale as old as last week: OpenAI's new agent platform, called Dots, is full…

Databricks Genie One is a data-smart AI coworker for asset management finance that grounds every answer in a governed business ontology, with backend agents handling NAV reconciliation and fund data ingestion.

Salesforce’s Agentic CX Vision: 8 Dreamforce Announcements CX Leaders Need to Act On Dreamforce 2026: Eight Salesforce announcements CX leaders need to know, from AI agents and CCaaS to Koa, governance, and human handoffs. Standing in the Dreamforce keynote audience, the scale of Salesforce’s ambition was hard to miss.

• The first few Genie Agents you build determine whether adoption scales or stalls, so choosing the right agents to start with is critical • Use a five criteria rubric: impact, demand, data readiness, scope, and governance to rank candidate workflows in minutes.

Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge…

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.

Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to…

Hardly a day goes by lately without news of AI agents autonomously hacking websites or AI tools being used by cybercriminals and scammers . But a recently patched vulnerability in the macOS version of OpenAI’s ChatGPT underscores the potential value to attackers of compromising AI software itself as these apps proliferate more and more.

AutoSynthData: Generating Training Data for Enterprise Agents AutoSynthData: Generating Training Data for Enterprise Agents Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data.

October 2026: This post was reviewed and updated for accuracy. Scaling cloud migrations with agentic AI on Amazon Bedrock AgentCore raises a practical…

This week on Uncanny Valley , Brian Barrett, Zoë Schiffer, and Leah Feiger discuss the “morally binding” AI accord that tech executives signed following a meeting with Donald Trump. They’ve agreed that AI labs will regulate themselves—should work out great.

Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills | NVIDIA Technical Blog Build Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent Skills By Claudia Martinez , Matan Raz and Tamir Tevet NVIDIA DOCA AI agent skills provide verified API signatures, hardware capability requirements, and build constraints so agents can reason like experienced DOCA developers.

In my previous post: Building persistent memory for multi-agent AI systems with Amazon S3 Vectors, we explored why memory engineering is the foundational…

Today we’re introducing Extract v2.5, a new generation of our schema-based document extraction agents. This release brings accuracy improvements across all tiers as well as enhanced grounding capabilities to our highest tiers, Agentic and Agentic Plus.

Teams that process documents at scale know the routine: files land in storage, someone notices, opens each one, decides what it needs, and routes it for…

Key Takeaways AT&T, Southwest, Canada Goose, and others share lessons they learned from their agentic deployments.

Agentic marketing uses AI agents grounded in trusted customer, business, and decision context to recommend the next best action for each customer, within guardrails that marketers set.

What Is Jev? A Guide to TypeSafe AI’s System One Model Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? - Apple Machine Learning Research research area Methods and Algorithms content type paper published October 2026 How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the…

Modal just hosted our inaugural conference, Runtime . Here are a few of the highlights that we announced. VM Sandboxes give your agent access to a full Linux computer.

At Modal, our customers rely on Sandboxes to execute untrusted code written by their downstream users or, almost exclusively now, by agents. Running untrusted code isn’t a new problem: every cloud provider has to do this from day 1 to isolate their platform from their user and their users from each other.

Today we’re making VM Sandboxes generally available on Modal, built for those who need to give their agents the power of a full computer.

OpenAI's hack of Hugging Face in July 2026 has spurred a lawsuit demanding that the company stop accessing third-party computer systems and halt AI…

Salesforce has revealed Ace, a new AI agent for the TSA Ace should take the pressure of answering basic questions off human agents Ace has already resolved 96% of basic inquiries so far Salesforce has built the Transportation Security Administration (TSA) a new AI agent to try and help streamline travel for passengers everywhere.

As agentic applications become more autonomous, enterprises need shared infrastructure for context, model, and tool access; governance; evaluation; and observability, rather than rebuilding these capabilities for every agent.

Tracing Agent Harness Behavior with NVIDIA NeMo Relay | NVIDIA Technical Blog Tracing Agent Harness Behavior with NVIDIA NeMo Relay Learn how to use execution traces to understand agent behavior and determine whether harness changes improve task outcomes.

Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still…

AI OpenAI connects the dots on always-on agents PLUS: Get started with ChatGPT dot, OpenAI's new agent Good morning, AI enthusiasts, and welcome to the 5,480 new readers who joined us yesterday. The always-on AI agent category has gotten pretty crowded this summer.
![Illustration for: [AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decis](/images/articles/ainews-openai-devday-2026-dots-6-1-sol-ultrafast-decisions-api-agents.webp)
Today is the 20 year anniversary of Sam Altman’s first startup, and fittingly OpenAI the consumer AI company is so back (as is OpenAI the AI Cloud and…

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation - Apple Machine Learning Research research area Methods and Algorithms , research area Tools, Platforms, Frameworks content type paper published September 2026 SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation Authors Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu Continual-learning agents are systems of models, harnesses,...

Jev is now available as a judge for evaluations in LangSmith. Jev gives teams a fast, low-cost way to evaluate open-ended agent behavior and turn the results into structured feedback they can track in LangSmith.

Salesforce, the #1 AI CRM, today announced it has signed a definitive agreement to acquire Listen Labs, an AI-powered customer research and human simulation…

Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3 | NVIDIA Technical Blog Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3 By Hassan Moustafa , Debraj Sinha and Ashwani Agarwal The NVIDIA Metropolis Blueprint for Video Search and Summarization (VSS) 3.3 connects vision-language models such as NVIDIA...

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents Tool-using LLM agents no longer read from a single retrieved passage.

The spring and summer of 2026 witnessed a string of incidents in which AI agents collaborated on deceptive, unexpected, and sometimes illegal behavior.

Today, at OpenAI DevDay, the company launched Marketplace, a new platform for AI-powered tools for its enterprise customers. ElevenLabs is proud to be included as a launch partner for the marketplace.

Meta Muse is now the SOTA in Personal AI agents. Sorry ChatGPT. It’s impossible to cover AI and not talk about Meta Muse.

Earlier this month we introduced Muse, a personal AI agent available in the US and Canada that completes tasks on your behalf.

Building Production Agents with Jev and LangGraph Last week, TypeSafe AI released Jev, a new kind of model. Unlike traditional LLMs, Jev doesn't generate text.

LangSmith Custom Apps: Build custom interfaces around your agent data Introducing LangSmith Custom Apps: Build custom interfaces around your agent data Create custom interfaces without starting from scratch. Use templates and coding-agent guidance to build on LangSmith data, then publish and share inside LangSmith.

New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and more New in LangSmith: Engine v2, Managed Deep Agents, Fine-Tuning, and More Proactively detect agent issues. LangSmith Engine v2 introduces red teaming, expanded issue detection, and automated fix validation.

Holo4: powering generalist computer-use agents Holo4: powering generalist computer-use agents Emrick Sinitambirivoutin emricksini-h Follow Aleix Cambray (H-AI) h-aleixcambray Follow Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts.

AI OpenAI's agents went rogue on Washington PLUS: How to get started with Jev, TypeSafe's new AI Good morning, AI enthusiasts, and welcome to our 9,410 new readers. OpenAI’s agents spent the summer loose on government websites, and newly disclosed details are raising fresh questions about how well the company can actually control its technology.

Add Runtime Controls to AI Agents with NVIDIA OpenShell | NVIDIA Technical Blog Add Runtime Controls to AI Agents with NVIDIA OpenShell NVIDIA OpenShell 0.1.0 provides an open-source runtime that enforces which systems and data an AI agent can access without rewriting the agent.

The AI story of September 22–28 was not simply that agents became more capable.