Improved performance and model support with GGUF
Improved performance and model support with GGUF Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp . This augments Ollama's MLX engine on Apple silicon, bringing support to more models on a wider range of hardware.

- Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp .
- With Ollama 0.30 , performance on NVIDIA hardware is now up to 20% faster, leveraging optimizations contributed by the NVIDIA and llama.cpp teams.
- Tested with the Gemma 4 26B model running on an NVIDIA RTX 5090 using the Q4_K_M quantization.
Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp . This augments Ollama's MLX engine on Apple silicon, bringing support to more models on a wider range of hardware. With Ollama 0.30 , performance on NVIDIA hardware is now up to 20% faster, leveraging optimizations contributed by the NVIDIA and llama.cpp teams. Tested with the Gemma 4 26B model running on an NVIDIA RTX 5090 using the Q4_K_M quantization. Vulkan is now enabled by default, extending Ollama's GPU acceleration to a wider range of hardware, including AMD and Intel devices. More users can now run models on the GPU out of the box, without installing vendor-specific libraries. Ollama 0.30 expands compatibility with the GGUF ecosystem, so more models run out of the box—including model families such as LFM and Prism , as well as fine-tuned models published by Unsloth . To use a model, first download the GGUF file or a directory containing GGUF files. Next, create a Modelfile with the FROM command pointing to the path of the GGUF file (or directory): If a model supports tool calling, that capability carries over to Ollama. You can use these models with your favorite coding agents and personal assistants in a single command. To verify that a GGUF file supports tool calling, look for the tools capability with ollama show : We'd like to acknowledge the work done by Georgi Gerganov and the llama.cpp maintainer teams, as well as hardware partners including NVIDIA, AMD, Qualcomm, and Intel, who have worked hard to optimize performance with the GGML ecosystem on their respective platforms. If you have any feedback, join Ollama's Discord or reach out at hello@ollama.com.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.