NVIDIA Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is now available on Ollama, and it runs completely on your own device. It's a 30 billion parameter (3B active) open model from NVIDIA built for agents that stay running: gathering context, calling tools, and working through multi-step tasks.

- NVIDIA Nemotron 3.5 Lightning is now available on Ollama, and it runs completely on your own device.
- It's a 30 billion parameter (3B active) open model from NVIDIA built for agents that stay running: gathering context, calling tools, and working through multi-step tasks.
- Nemotron 3.5 Lightning is made for agentic tasks such as reading a file, calling a tool, sorting a result, and retrying something that failed.
NVIDIA Nemotron 3.5 Lightning is now available on Ollama, and it runs completely on your own device. It's a 30 billion parameter (3B active) open model from NVIDIA built for agents that stay running: gathering context, calling tools, and working through multi-step tasks. Nemotron 3.5 Lightning is made for agentic tasks such as reading a file, calling a tool, sorting a result, and retrying something that failed. Most of these steps don't need a large model. At 3B active parameters per token, it's built for local systems rather than the datacenter, and running locally means your data stays on your device. Runs where you work: 30B total parameters with only 3B active per token, on a hybrid Mixture-of-Experts architecture. It runs locally on NVIDIA RTX PCs , NVIDIA RTX PRO workstations , NVIDIA DGX Spark and DGX Station , and in the datacenter and cloud. Built for agent harnesses: developed with the Nemotron Coalition and trained for the tools developers already use, across coding, tool calling, instruction following and multi-turn work. 1M token context: a context length up to 1M, which leaves room for long tool histories across multi-turn workflows. Optimized inference: speculative decoding using multi-token prediction (MTP), DFlash or DSpark, offering up to 4x higher throughput than comparable open models. Yours to customize: an open model trained on open datasets. Post-train it for a specific task and run the result anywhere, from edge to datacenter. Nemotron 3.5 Lightning excels on the following workloads: Long-running personal assistants. Email, calendar, projects and bookings. Running locally, the agent can use local context and none of it is sent elsewhere. Coding sub-agents. Running tests, searching the codebase and applying refactors, inside the harnesses you already use. Security operations. Enriching alerts, classifying incidents, querying logs, correlating indicators and preparing structured findings for analysts. A local tier alongside the cloud. Nemotron 3.5 Lightning handles the high-volume steps locally, and a larger hosted model picks up the few that need one. Same CLI, same API. A specialist you train yourself. Open weights and open datasets, so you can post-train Nemotron 3.5 Lightning for one narrow job and run the result locally. Download Ollama , then run Nemotron 3.5 Lightning with your tool of choice. ollama launch claude --model nemotron-3.5-lightning ollama launch openclaw --model nemotron-3.5-lightning ollama launch hermes --model nemotron-3.5-lightning ollama launch opencode --model nemotron-3.5-lightning For users on Apple silicon, Ollama offers the model with state-of-the-art performance: nemotron-3.5-lightning:30b-mlx . The same pattern works for models running in Ollama's cloud, so an agent can send an individual step to a larger model without changing anything else. Nemotron 3.5 Lightning offers 4x higher throughput and 30% faster task completion time compared to other leading open models of similar size and offers leading accuracy across agentic, coding and reasoning tasks. For agents that stay running, throughput is the number that matters most: more steps per minute means long tasks finish. Full results and test configurations are in NVIDIA's launch blog.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.