
Model Auditing: Faithfulness of Post-hoc Explanations Investigated
A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

A study published on arXiv cs.CV investigates the faithfulness of post-hoc explanations for a pedestrian detection model across different domains.

Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

The predictive analytics landscape has evolved significantly in recent years, with AI-powered analytics driving growth and innovation across various industries.

Takeaways − ByteShape released ShapeLearn-quantized Qwen3.8-27B GGUFs in five sizes from 2.56 to 3.84 bits per weight. The smallest build fits in 8.8 GB VRAM; the largest IQ4_XS variant is 13.1 GB.

Subtopic Small Models · Vision Language · Quantization Takeaways − Mitsuba is a ternary 1.58-bit quantization of Qwen3.8-27B, shrunk to 7.3 GB for a single 16 GB GPU. Purpose-built for ComfyUI: image to prompt generation for Stable Diffusion, Krea, and video pipelines.

Subtopic Mixture Of Experts · Long Context · Vision Language Takeaways − China Telecom released Xing4.0-29B-A4B , a 29B MoE with only 4B active parameters. Native 256K context extensible to 512K using MLA attention and 64 routed experts.

Chinese AI models parrot state doctrine or refuse to answer on sensitive topics Chinese AI models frequently toe the party line when asked politically sensitive questions, according to a study by Aleph Alpha. The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, giving it a commercial interest in distinguishing its models from Chinese competitors.

Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of Pro Starting in October 2026, Google will restrict access to its Gemini models for personal account users.

Subtopic Mixture Of Experts · Long Context Takeaways − Kolibri-1 is a 78B MoE with 3.46B active parameters, Apache 2.0, German and English focus. Context window validated up to 1,048,576 tokens, native 262,144, no position scaling tricks required.

Takeaways − Red Hat AI released an FP8 quantized build of NVIDIA Nemotron 3.5 Lightning 30B A3B. Cuts GPU memory and disk by roughly 50% versus the BF16 reference weights.

Apparently, OpenAI isn't trying to build "magic intelligence in the sky" anymore OpenAI CEO Sam Altman is pushing back against religious analogies tied to AI models.

Subtopic Dpo · Fine Tuning · Distillation Takeaways − Cohere released North Small Translate , a 218B / 25B-active MoE translation model with open weights. Scores 83.60 on WMT26 across 50+ languages, beating DeepL NextGen (81.37) and Google Translate (68.20).

Subtopic Small Models · Long Context · Inference Optimization Takeaways − OpenBMB released MiniCPM5-2B-DSpark , a 324M-parameter speculative decoding draft model for MiniCPM5-2B. Aggregate accepted length of 5.52 tokens per target forward pass at greedy decoding, 4.05 at temperature 1.0.

Some big AI companies seem to think that the best way to keep models from becoming too chaotic or too mischievous is to keep them locked inside of labs. If only a chosen few can access them, the thinking goes, they can do less damage in the real world while researchers try to understand what they’re capable of.

Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge…

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information.

The latest AI news we announced in September 2026 Here’s a recap of some of our biggest AI updates from September, including Gemini 4 Argon, new Connected Apps in the Gemini App, and WeatherNext 3. Your browser does not support the audio element.

Toward provably private learning from federated data Toward provably private learning from federated data Katharine Daly, Software Engineer, and Daniel Ramage, Research Director, Google Research We announce a new Federated Learning system that provides externally verifiable privacy guarantees while shifting computation to the server to improve training speed, accuracy, and device coverage.

Prior to joining Airbnb as CTO in January, Ahmad Al-Dahle was head of generative AI at Meta and led the launch of its open source Llama models over 2023-2025.

How to Build a Model Router in the Harness How to Build a Model Router in the Harness Many tasks don't need frontier intelligence. In our experiments, model routing cut median cost per coding task by 64% compared with the baseline, with no noticeable change in quality.

Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingly capable open models are shrinking to…

AI Tavus' AI looks, listens, and talks back live PLUS: Is ChatGPT “Dots” worth upgrading to Pro for? Good morning, AI enthusiasts, and welcome to our 6,908 new readers.

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models - Apple Machine Learning Research research area Speech and Natural Language Processing content type paper published October 2026 Language Discrimination Improves Linguistic Learning in Multilingual Speech Models Authors Maureen de Seyssel, Jie Chi*, Zakaria Aldeneh* Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total...

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users.

• Give finance professionals an AI coworker grounded in their business context to accelerate decision-making • Enable teams to analyze performance, model scenarios, investigate variances, and speed up forecasting and financial reporting • Apply consistent guardrails across data access, actions, and AI usage, so every...

“What you see is what you get” is a guiding principle for many software engineers — create programs where the content you’re editing looks the same as the…

Agentic marketing uses AI agents grounded in trusted customer, business, and decision context to recommend the next best action for each customer, within guardrails that marketers set.

Gemini. It only took Google seven months (February to September) to release their next frontier model.

AI Argon aims to return Google to the frontier PLUS: How (and why) to set up Meta’s Muse Desktop Good morning, AI enthusiasts, and welcome to the 5,879 new readers who joined us yesterday.
![Illustration for: [AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1](/images/articles/ainews-gemini-4-argon-gdm-s-answer-to-astra-fable-with-1m-output.webp)
GDM last shipped a larger-than-Flash model in February (3.1 Pro), and after successive incremental 3.x Flash versions and the big GDM management shakeup last…

What Is Jev? A Guide to TypeSafe AI’s System One Model Agents run in a loop: an LLM decides what to do, a tool executes, a model evaluates the results, and then continues in that loop until the task is complete.

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the…

Fetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. The Decisions API is billed at $0.04 per million input tokens.

Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton | NVIDIA Technical Blog Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton By Junyi Qiu , Shijie Liu , J Wyman and Mudit Aggarwal Generative recommender systems reformulate personalization as sequence modeling over user behavior, and NVIDIA recsys-examples now provides an end-to-end HSTU inference workflow with Dynamo-Triton.

Runway app for iPhone Runway app for Android An open-weight world action model that turns Runway's video pretraining into control for real robots. “Pick up the tennis ball and put it in the box.” Today we're announcing Praxis-1 , our first open-weight world action model.

Google promised Gemini 3.5 Pro in June, but it spent the summer trotting out smaller Flash models.

Gemini 4 Argon: our next era of frontier intelligence Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. SVP, Google DeepMind and Chief AI Architect, Google Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.

ai_decide is a new Databricks AI Function that makes fast decisions directly on your data. It takes unstructured text and returns structured decisions in a fraction of a second.

As agentic applications become more autonomous, enterprises need shared infrastructure for context, model, and tool access; governance; evaluation; and observability, rather than rebuilding these capabilities for every agent.

How well does nDCG capture the quality of modern retrieval systems? nDCG is the established yardstick for measuring how well search systems rank results, and is widely used across benchmarks such as MTEB and BEIR .

State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested, particularly across financial datasets, parsed PDFs, and visually rich documents. A new Fast tier: Embed 5 Fast brings strong retrieval quality to latency, and cost-sensitive workloads, at $0.08 per million tokens.

Top NewsAnthropic and OpenAI race to release smarter and cheaper modelsSources:Anthropic launches Claude Opus 5.5 with stricter safeguards for…

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference EMNLP content type paper published September 2026 On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study Authors Iuri Macocco†, Pau Rodríguez Lopez, Arno Blaas, Luca Zappella, Marco Baroni†*, Xavier Suau...

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation - Apple Machine Learning Research research area Methods and Algorithms , research area Tools, Platforms, Frameworks content type paper published September 2026 SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation Authors Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu Continual-learning agents are systems of models, harnesses,...

Jev is now available as a judge for evaluations in LangSmith. Jev gives teams a fast, low-cost way to evaluate open-ended agent behavior and turn the results into structured feedback they can track in LangSmith.

While most tech CEOs take a build it, and they will come mentality to building their products, Stability AI CEO Prem Akkaraju said he doesn t develop any applications or AI models without the experts in the field.

GLM-5.3 and the spread of advanced cyber capabilities Cole McFaul, Robert Xiao, Tripp Gallagher Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits.

At a glance Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models,…

The recently released Jev AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.While Jev aims to classify things,…

AI Anthropic's mid-tier Claude climbs the rankings PLUS: Pick the right Claude model with one quick test Good morning, AI enthusiasts, and welcome to our 5,342 new readers. OpenAI takes the stage today for one of its most hyped days of the year.

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models - Apple Machine Learning Research research area Methods and Algorithms , research area Speech and Natural Language Processing conference NeurIPS content type paper published September 2026 The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models Authors Xavier Suau, Alex Ferrando de las Morenas,...

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

Warning : Undefined variable $stocks in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 448 Warning : Undefined variable $funds in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d code on line 472 Warning : foreach() argument must be of type array|object, null given in /var/www/briefs.co/htdocs/wp-content/plugins/oxygen/component-framework/components/classes/code-block.class.php(133) : eval()'d...

Eleven v4 and Eleven v4 Turbo Text to Speech models &)]:tw-overflow-x-hidden"> strong]:tw-font-normal tw-text-gray-600"> Introducing Eleven v4 Meet Eleven v4, our most emotive model yet. With 3x credits included on Creator+ until October 12 Create controllable, expressive speech layered with emotion, audio events, and immersive soundscapes.

A line of text can change significantly depending on how it’s spoken. “I need you to stay calm" should sound different depending on who's saying it, whether that's a doctor delivering it gently to a frightened patient, or a character in a game shouting to his squad before dropping into battle.

Space was always supposed to be the final frontier of human exploration. It’s shaping up to be the final frontier for artificial intelligence too.Last…

Alibaba s Qwen team shipped Qwen-Image-2.1 on September 20, 2026, and it changes the calculus for anyone running AI image generation on their own hardware in Australia. The 20-billion-parameter model handles text-to-image generation, multi-reference editing with up to 10 source images, and native transparent PNG output from a single checkpoint.

Author(s): Umair Ali Khan, Ph.D. Originally published on Towards AI. How AI decision models offer a fast and cost-effective approach to turning unstructured…

Author(s): Chew Loong Nian – AI ENGINEER Originally published on Towards AI.

Author(s): Dave R | Microsoft Azure & AI MVP ☁️ Originally published on Towards AI.