Skip to main content
Open Source

NVIDIA Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud. It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.

By Precis Daily Newsroom1 min read221 words
Illustration for: NVIDIA Nemotron 3 Ultra
Illustration
Key points
  • NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud.
  • It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls.
  • Frontier reasoning, high efficiency: 550B total parameters with only 55B active per token, and optimized for NVFP4—NVIDIA’s 4-bit floating point format that packs the model into less memory and runs faster.

NVIDIA Nemotron 3 Ultra is now available on Ollama’s cloud. It’s a 550 billion parameter (55B active) open model from NVIDIA built for long-running, agentic workflows with fast and affordable performance across hundreds of tool calls. Built for long-running agents: Tuned for agent orchestration, coding agents, deep research, and complex enterprise workflows that run across hundreds of steps. 1M token context: Keep entire codebases, long tool histories, and research trails in context without losing the thread. Frontier reasoning, high efficiency: 550B total parameters with only 55B active per token, and optimized for NVFP4—NVIDIA’s 4-bit floating point format that packs the model into less memory and runs faster. Download Ollama , then run Nemotron 3 Ultra with your tool of choice. ollama launch claude --model nemotron-3-ultra:cloud ollama launch hermes --model nemotron-3-ultra:cloud ollama launch openclaw --model nemotron-3-ultra:cloud Nemotron 3 Ultra leads on accuracy across agent productivity, instruction following, and long-context tasks, while delivering leading throughput—saving up to 30% on costs compared to other leading open models. Figure 1: Nemotron 3 Ultra leads among open models on agentic benchmarks for agent productivity, coding, and instruction following. Figure 2: Nemotron 3 Ultra is in the most attractive quadrant with leading accuracy and leading throughput among open models. Figure 3: Nemotron 3 Ultra saves up to 30% in costs and leads on the cost efficiency frontier.

Sources

Summarized from the linked originals.

Related stories

Illustration for: The Agent Said It Was Done. The Database Disagreed.
Open Source

The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.

Hugging Face Blog13 min