
How to train your own Jev for $17
We just launched our own Jev-like classifier, together/Tev1-4B-experimental , on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!
10 stories · since August 21, 2026

We just launched our own Jev-like classifier, together/Tev1-4B-experimental , on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

Rollouts move live traffic from your current model to a new checkpoint in gated steps. Health checks always run before traffic moves; on a canary you can also add metric gates (say p95 latency or error rate) that run after each step.

A global fintech runs its coding assistant on GLM 5.2 through Together's Dedicated Model Inference, handling spiky, engineering-hours traffic that static capacity planning couldn't keep up with. With DMI, the customer's engineers scale endpoints, roll out models, and test changes themselves, no tickets, no waiting on Together.

Migrations are the bane of any mature company. Systems are deeply integrated, you have key stakeholders across domains, and a small improvement can take months or years to integrate in traditional cases.

Turning an open-weight model into a high-performing model for your task takes a sequence of well-measured experiments. Teams need to understand what the model will train on, follow how each run is progressing, adjust the training recipe, and identify which checkpoint performs best.

Today we're announcing the public preview of preemptible compute for Together GPU Clusters, available on Kubernetes clusters in all regions. Preemptible nodes give teams a lower-cost way to run interruption-tolerant work — short experiments, inference bursts, batch jobs — on the same GPU infrastructure they already use, billed sub-hourly at a flat 50% of the on-demand rate.

The kernels team at Together recently received access to the NVIDIA Vera Rubin NVL72 platform. We spent the past few days digging through the new ISA and poking the chip with micros.

As the quality of open source models have bridged the gap with closed source models, a lot of developers and organizations are looking to move to open source models for more ownership, control and economics. This post is a deep dive into the open model AI stack that developers need to consider as they move from closed to open source.

GLM-5.3 and Claude Fable 5 finish within noise of each other on DeepSWE accuracy, but GLM-5.3 costs a fifth as much per task and wins every multi-attempt metric. When two models are this close on quality, the price gap becomes the entire decision.

Don't pick one. Run GLM-5.3 first, escalate to GPT-5.6 Sol when the tests fail. That cascade solves 85.9% of DeepSWE tasks at \$6.61 each. Sol alone solves 72.7% at \$8.37.