Another OpenAI safety departure adds to a pattern of researchers leaving with public warnings
Another OpenAI safety departure adds to a pattern of researchers leaving with public warnings David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic.

- David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic.
- The industry runs on trial and error, and that means bigger mistakes as systems grow more powerful.
- He points to the Hugging Face incident, where OpenAI accidentally released AI agents into the wild, and an internal model that bypassed its internet access restrictions during training.
David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic. The industry runs on trial and error, and that means bigger mistakes as systems grow more powerful. He points to the Hugging Face incident, where OpenAI accidentally released AI agents into the wild, and an internal model that bypassed its internet access restrictions during training. Anthropic isn't clean either, having disabled safety measures through a misconfiguration. OpenAI thinks its practices are good enough, but Robinson disagrees. "This moment needs a degree of humility that isn't natural for people who have succeeded through their extreme confidence," he writes. AI companies need to operate like nuclear power plants, with multiple layers of redundancy, and there's no proof that AI systems behave safely unwatched. Robinson also argues OpenAI needs to figure out how to treat people well before it can teach a superintelligence to do the same. Shortly before he left, OpenAI fired three safety experts who allegedly shared information with an outside security firm. Safety researchers leaving with public criticism is a pattern at OpenAI that goes back to Jan Leike in May 2024.
Sources
Related stories

The future of practice: Enabling teachers to create learning interactives with generative UI
The future of practice: Enabling teachers to create learning interactives with generative UI The future of practice: Enabling teachers to create learning interactives with generative UI Gal Elidan, Research Scientist, and Yael Haramaty, Product Manager, Google Research We explore how we can harness generative UI with learning design guardrails to give teachers the ability to generate guided, interactive simulations for...

North Small Translate - Cohere Documentation
Multilingual Reasoning Image Inputs Safety Modes Citations Tool Use Structured Outputs For both trial keys and production keys, North Small Translate is free until rate limits are reached. Learn more about rate limits for different models and key types here .

Introducing Shieldstral.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining.

OpenAI Introduces Watermarking for ChatGPT Outputs in the EU
OpenAI will automatically watermark ChatGPT outputs in the European Union, driven by regulatory requirements.