NVIDIA Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud. It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.

- NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud.
- It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.
- Frontier reasoning, high efficiency: 550B total parameters with only 55B active per token, and optimized for NVFP4—NVIDIA’s 4-bit floating point format that packs the model into less memory and runs faster.
NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud. It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls. Built for long-running agents: Tuned for agent orchestration, coding agents, deep research, and complex enterprise workflows that run across hundreds of steps. 1M token context: Keep entire codebases, long tool histories, and research trails in context without losing the thread. Frontier reasoning, high efficiency: 550B total parameters with only 55B active per token, and optimized for NVFP4—NVIDIA’s 4-bit floating point format that packs the model into less memory and runs faster. Download Ollama , then run Nemotron 3 Ultra with your tool of choice. ollama launch claude --model nemotron-3-ultra:cloud ollama launch hermes --model nemotron-3-ultra:cloud ollama launch openclaw --model nemotron-3-ultra:cloud Nemotron 3 Ultra leads on accuracy across agent productivity, instruction following, and long-context tasks, while delivering leading throughput—saving up to 30% on costs compared to other leading open models. Figure 1: Nemotron 3 Ultra leads among open models on agentic benchmarks for agent productivity, coding, and instruction following. Figure 2: Nemotron 3 Ultra is in the most attractive quadrant with leading accuracy and leading throughput among open models. Figure 3: Nemotron 3 Ultra saves up to 30% in costs and leads on the cost efficiency frontier.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.