Ollama now supports Jev-style decision models
Ollama now supports Jev-style decision models Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions: Three new decision models available today via Ollama This new API is available as of Ollama 0.35 by using the new /v1/systemone endpoint.

- Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions: Three new decision models available today via Ollama This new API is available as of Ollama 0.35 by using the new /v1/systemone endpoint.
- Send text as state with a set of named questions, and a model running on your machine answers them all in one request.
- Nimble 9B averaged 91ms per decision in the Pac-Man example below when running locally on an M5 Max.
Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions: Three new decision models available today via Ollama This new API is available as of Ollama 0.35 by using the new /v1/systemone endpoint. Send text as state with a set of named questions, and a model running on your machine answers them all in one request. This is great for tasks that require fast decisions, such as ticket triage, model routing, and content or safety moderation. Decision models on Ollama are fast, as requests don't have to travel over a network. Nimble 9B averaged 91ms per decision in the Pac-Man example below when running locally on an M5 Max. That's fast enough to make rapid decisions such as playing a game or processing content in real time: Nimble 9B running on a MacBook Pro M5 Max, replayed at real-time speed. Three new decision models are available to run via Ollama: nimble : open-source 9B parameter decision model developed by Bespoke Labs tev1 : an experimental 4B decision model from Together AI tev1:0.8b : an experimental 0.8B decision model from Together AI More decision models are coming soon, including models served by Ollama's cloud. Bespoke Labs public benchmarks · accuracy, higher is better Mean accuracy across 13 public data sets with human labels, covering 3,880 decisions. Nimble and Tev1 were evaluated on Ollama; Jev 1.13 is from Bespoke Labs' published run on the same decisions. See the Ollama evaluation results and benchmark suite . To get started, first download or upgrade to the latest version of Ollama. Next, download a decision model such as nimble : You can make a request via curl or via TypeSafe's official Python SDK. curl http://localhost:11434/v1/systemone -d '{ "ticket": "I was charged twice. Please refund the extra payment." "instructions": "Which team should handle this ticket?", "instructions": "Does the customer explicitly ask for a refund?" "instructions": "How urgent is this ticket?", "criteria": ["Routine", "Soon", "Urgent"] "probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003}, "refund": {"type": "noul", "noul": 0.997}, "legend": {"0": "Routine", "1": "Soon", "2": "Urgent"}, "probabilities": {"0": 0.378, "1": 0.429, "2": 0.193}, "usage": {"input_tokens": 841, "output_tokens": 4} uv add typesafe-sdk # or: pip install typesafe-sdk export TYPESAFE_BASE_URL=http://localhost:11434 from typesafe_sdk import Choice, Noul, Score, TypeSafeClient ticket = "I was charged twice. Please refund the extra payment." instructions="Which team should handle this ticket?", instructions="Does the customer explicitly ask for a refund?", instructions="How urgent is this ticket?", with TypeSafeClient(timeout=120) as client: print(result.choices["team"].choice) # billing print(result.nouls["refund"].noul) # 0.997 print(result.scores["urgency"].score) # 0.815 This is the first of many releases to come adding decision model support to Ollama. Future updates will include: Faster performance on Apple Silicon powered by MLX More models specializing in different kinds of decision making
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.