Skip to main content
Guides

Running Bigger AI Models, More Efficiently | SambaNova

TL;DR AI has moved from experimentation into production, and that shift exposes three hard constraints: models keep getting bigger, cost climbs steeply at…

By Precis Daily Newsroom1 min read241 words
Illustration for: Running Bigger AI Models, More Efficiently | SambaNova
Illustration
Key points
  • TL;DR AI has moved from experimentation into production, and that shift exposes three hard constraints: models keep getting bigger, cost climbs steeply at scale, and power becomes a physical ceiling.
  • Inference in the agentic era behaves nothing like the single-model throughput problem of the past.
  • SambaNova's focus is premium inference, defined on the panel by speed and size: Running the largest models much more efficiently, at full precision, rather than trading accuracy for speed.

TL;DR AI has moved from experimentation into production, and that shift exposes three hard constraints: models keep getting bigger, cost climbs steeply at scale, and power becomes a physical ceiling. On a recent Fortune panel, SambaNova CEO Rodrigo Liang and Adaption Labs CEO Sara Hooker agreed the industry's central problem is now efficiency, though they approach it from different angles: Liang from the infrastructure side, Hooker from the model-architecture side. Large models are not going away. For the most demanding workloads, they are unavoidable, so the real question is how to run them efficiently rather than whether to run them at all. Inference in the agentic era behaves nothing like the single-model throughput problem of the past. In Hooker's words, “inference is a different beast.” Constant data movement, not raw compute, is the bottleneck. SambaNova's focus is premium inference, defined on the panel by speed and size: Running the largest models much more efficiently, at full precision, rather than trading accuracy for speed. Are Large Models Here to Stay? The question hung over the whole conversation: Are today's largest foundation models where AI is heading, or will we look back on them as a detour? A recent Fortune panel titled “From the AI We Have to the AI We Need” put it directly to two people building very different answers, SambaNova co-founder and CEO Rodrigo Liang and Adaption Labs co-founder and CEO Sara Hooker, moderated by Fortune AI editor Jeremy Kahn.

Sources

Summarized from the linked originals.

Related stories

Illustration for: How to Build a Model Router in the Harness
Guides

How to Build a Model Router in the Harness How to Build a Model Router in the Harness Many tasks don't need frontier intelligence. In our experiments, model routing cut median cost per coding task by 64% compared with the baseline, with no noticeable change in quality.

LangChain Blog11 min