WritingSambaNova SystemsSambaNova Systemspublished Aug 11, 2026seen 2w

Running Bigger AI Models, More Efficiently | SambaNova

Open original ↗

Captured source

source ↗

Running Bigger AI Models, More Efficiently | SambaNova

BACK TO RESOURCES

Blog

Running Bigger AI Models, More Efficiently: The Real Production Challenge

By

SambaNova

--> August 11, 2026

TL;DR

AI has moved from experimentation into production, and that shift exposes three hard constraints: models keep getting bigger, cost climbs steeply at scale, and power becomes a physical ceiling.

On a recent Fortune panel, SambaNova CEO Rodrigo Liang and Adaption Labs CEO Sara Hooker agreed the industry's central problem is now efficiency, though they approach it from different angles: Liang from the infrastructure side, Hooker from the model-architecture side.

Large models are not going away. For the most demanding workloads, they are unavoidable, so the real question is how to run them efficiently rather than whether to run them at all.

Inference in the agentic era behaves nothing like the single-model throughput problem of the past. In Hooker's words, “inference is a different beast.” Constant data movement, not raw compute, is the bottleneck.

SambaNova's focus is premium inference, defined on the panel by speed and size: Running the largest models much more efficiently, at full precision, rather than trading accuracy for speed.

Are Large Models Here to Stay?

The question hung over the whole conversation: Are today's largest foundation models where AI is heading, or will we look back on them as a detour? A recent Fortune panel titled “From the AI We Have to the AI We Need” put it directly to two people building very different answers, SambaNova co-founder and CEO Rodrigo Liang and Adaption Labs co-founder and CEO Sara Hooker, moderated by Fortune AI editor Jeremy Kahn.

They came at it from opposite ends of the stack. Hooker works on model architecture and the way models learn. Liang engineers the chips that run them. But they converged on the same diagnosis: The next phase of AI will be defined less by how large models can get and more by how efficiently we can run the large models we already depend on. Here is what the panel surfaced.

What the Panel Said: AI Has Entered Its Production Phase

Asked whether the industry is stuck with the largest models forever, Liang gave a both/and answer. He explained that the world is heterogeneous, so scale still matters, and large models are going to be here for a while, but there is still a need for a new generation of more efficient models. AI is entering its production phase, and production at scale runs into three constraints at once.

The Three Constraints of Production AI

First, the models keep getting bigger, at least the class of models that attract the most attention. Second, cost becomes a serious problem as you scale, and that cost is largely driven by the infrastructure available to run those models. Third, there is power.

As Liang was careful to point out, the industry's scramble for chips and energy is not caused by small models. It is the large models that consume a disproportionate share of the compute, which is exactly why the pressure lands hardest at the top end.

The constraints do not point toward abandoning scale. They point toward running it far more efficiently..

The Model-Side View: Why the Workload Is Changing

Hooker approached the same problem from the model side. In her words, “Today's models are monolithic. They're stuck in time. A big goal of how we think about what should be fundamentally different is how do we efficiently learn from the environment, because now all models are moving towards interaction. Now it matters more that last mile performance, and for that you need models which can evolve, otherwise you end up with massive inefficiencies.”

She sees an inflection point arriving with real urgency. Part of it is not throwing the largest models at problems that do not need them. She emphasized, “You shouldn't just apply the same model to all problems. Probably 90% of problems are very easy. Many of the things that you do in bulk processing, for example, you shouldn't be throwing a massive model at. 10% you need a ton of compute.”

Her own focus, continuous learning, follows from this. As she defined it on the panel: “Continuous learning is really the goal that you should be able to update model behavior without forgetting what you already know.”

The two perspectives are complementary: efficiency as the key to improving AI. Hooker describes why the workload is changing. Liang describes what it takes to serve that workload.

Why Inference Has Become a Different Beast

For most of the last decade, AI infrastructure was tuned for training: one enormous model, one enormous workload, optimize for throughput and push data through. Hooker's point on the panel was that inference in the agentic era looks nothing like that. What changes everything is that agentic tasks force the system to keep returning to memory, she continues “You have to check in as you go. That is a very painful operation. It means it's basically the whole crisis with memory.”

The Oven Problem: Memory, Not Compute, Is the Bottleneck

She offered an analogy that stuck, “We can build the biggest oven in the world and our only bottleneck is how fast the baker can put stuff in the oven. The baker is memory. What's fascinating now is that we're basically taking things out and putting things in the oven multiple times over the course of the recipe.”

Her conclusion was blunt, “Inference is a different beast,” and the design choices around memory become the critical ones. Liang pushed the same analogy further: “I think we have 50 ovens and we're actually moving from one oven to another to another. And so, ideally, you want to find ways to keep things in the oven and not keep opening. It's like the way I actually do turkeys for Thanksgiving.”

The point underneath the metaphor: Every unnecessary transfer of data costs time and energy. This is the heart of why inference behaves the way it does at scale. The constant shuttling of data between compute and memory increases costs, and agentic AI multiplies that movement.

How SambaNova Approaches It: Premium Inference

If the panel diagnosed the problem, SambaNova's approach is one answer to it. Liang described the company's focus as premium inference, defined by two things: speed and size. Until natively efficient models arrive, he argued, the biggest models are the ones that have to run better, and its Reconfigurable Dataflow Unit (RDU) chip is built for exactly that. He added, “The battleground...

Excerpt shown — open the source for the full document.

Notability

notability 4.0/10

Routine blog post, no traction indicators.