Ollama's transparent pricing
Ollama's Pro, Max, and Team plans now use transparent per-token pricing. Based on your feedback, every plan includes a monthly pool of usage credits. If you're on an existing Pro, Max, or Team plan, your plan continues to work as-is.

- Ollama's Pro, Max, and Team plans now use transparent per-token pricing.
- You can upgrade to the new pricing anytime in your Ollama account settings .
- Ollama's new pricing has no service fees and no 5-hour or weekly limits.
Ollama's Pro, Max, and Team plans now use transparent per-token pricing. Based on your feedback, every plan includes a monthly pool of usage credits. If you're on an existing Pro, Max, or Team plan, your plan continues to work as-is. You can upgrade to the new pricing anytime in your Ollama account settings . High-performance access to the latest open models, at published per-token rates Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools Monthly usage credits included with every plan Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models Ollama's new Pro, Max, and Team plans include a pool of usage credits that refreshes every month. Usage is consumed per token, at the rates published on Ollama's pricing page and on each model's page. Pro: $20/month, includes $60 of monthly usage Max: $100/month, includes $300 of monthly usage Team: $500/month, includes $1,000 of shared monthly usage for unlimited users The free plan now includes a small amount of monthly usage for a set of starter models. Add usage credits to access all models on the free plan with pay-as-you-go pricing : no service fees, no subscription required. Ollama's new pricing has no service fees and no 5-hour or weekly limits. Each plan's monthly pool refreshes automatically, and when you use it up, you can keep going at the same per-token rate. Every request runs on dedicated compute in the US and Europe, plus Singapore for a limited set of Qwen models, with zero data retention. We don't log your prompts, and we never train on your data. You can see exactly what each request cost in your account . Ollama's Team plan is now available for signup with introductory pricing of $500/month: $1,000 of shared included monthly usage, at published per-token rates One pool of usage credits shared across your organization with no per-seat limits Invite unlimited users, and view everyone's usage in one place If you're a larger team looking to scale access to open models with Ollama, contact us . Ollama's new plans are available today . Existing subscribers can upgrade in account settings . If you're on an existing Pro, Max, or Team plan, you'll remain on your current plan and can upgrade to the new pricing anytime in billing settings . All new signups start on the new plans. Can I change from a monthly to annual Pro plan and remain on the legacy pricing model? At this time, a change to billing cycle will result in an upgrade to the new pricing model. I changed to the new plan, but now I want to revert back. How can I request this? If you had an active subscription at the time of our pricing launch, we will be able to convert you back to your original plan. This will use the same billing cycle and tier you previously subscribed to. Please reach out to our support team to initiate this request. How long will legacy Ollama Pro and Ollama Max plans be supported? If you continue to auto-renew your existing subscription, you may remain on the legacy pricing model you currently have. However, if you change your billing cycle (monthly/annual) or subscription tier or cancel and re-subscribe later, this will convert your account to the new pricing plan. I'm an existing subscriber and the new pricing plan doesn't fit my needs. How can I get in touch? Please send an email to our support team and we will be happy to take a look at your case! What happens to my usage if I switch from an existing Pro or Max plan to the new one? Your usage is reset: the new plan's full monthly amount is available right away, and the session and weekly limits of the old plans no longer apply. Your monthly reset date stays on your subscription's original start date. Each plan includes a monthly amount of usage, and when you use it up you can keep going at the same per-token rate. Pro: $20/month, includes $60 of monthly usage Max: $100/month, includes $300 of monthly usage Team: $500/month, includes $1,000 of shared monthly usage for unlimited users Free: a starter amount each month for a set of starter models When does my monthly included usage reset? On Pro, Max, and Team plans, usage resets monthly on the same day of the month your plan started, including on annual plans. On the free plan, usage resets monthly from the date you signed up. Does unused included usage roll over to the next month? No. Instead, your included amount refreshes at each monthly reset. With the recent growth of Ollama's cloud, we received feedback that GPU-time based billing was difficult to predict, especially as open models have grown much larger (Kimi K3 has 2.8 trillion parameters). Now, usage is based on industry-standard token pricing.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.