WritingBasetenBasetenpublished Jan 7, 2026seen Jun 26

A Q A From Inference To Training The Inside Story Of Baseten S Newest Product

Open original ↗

Captured source

source ↗

A Q&A From Inference To Training: The Inside Story Of Baseten’s Newest Product Announcing our Series F . Learn more

Product

A Q&A From Inference To Training: The Inside Story Of Baseten’s Newest Product

Baseten launched a new training infrastructure that lets teams bring their existing code and run it on scalable compute with no heavy abstractions.

Authors

Madison Kanna

Raymond Cano

Last updated January 7, 2026

Share

TL;DR We launched a new training infrastructure that lets teams bring their existing code and run it on scalable compute with no heavy abstractions. It’s built for flexibility, custom models, audio models, multi-node jobs rather than point-and-click constraints. The product ties directly into Baseten’s inference stack, making it easy to go from training to deployment seamlessly. Early customers like OpenEvidence and Oxen have already used it to speed up inference, distill models, and even build full platforms on top of Baseten.

We just launched training at Baseten. I sat down with one of the engineers, Raymond, who built it to learn about the product, early customers, and what makes it different. How long have you been at Baseten and what brought you here? I joined in June 2024, so almost a year and a half ago. What really stood out to me was how transparent the founders were about iterating on the product and learning from the market. They talked about how we had "earned the opportunity" to do certain things with our customers. Like how we're an inference company, but we've earned the trust to do training with them. Looking back, it's kind of crazy that a year later I was in the thick of building that training product. We just launched training this week. Can you tell me about it? I'd love to start with the customer journey. A handful of our customers kept asking us, "When are you going to build training?" We had done tons of customer interviews and research, and when I joined the project, they handed me these videos and said, "Watch these and figure out what to build." What we learned was that if you're not an infrastructure company, it's really hard to get the resources to train world-class, open-source models. We heard stories of customers being up at 1 AM clicking the "add machine" button on DWS, waiting and waiting. They just wanted to train on Baseten. A lot of these customers already had training scripts and pipelines that worked. They had ML engineers and researchers who had worked really hard to get something functional. So we asked ourselves: how can we just help them get to Baseten? What did you build? It’s essentially a training infrastructure product. We started with the premise: take all your working code, no frills. Bring whatever image you want, whatever repo you run, and come run it on Baseten. We'll make that really easy. We provide storage primitives that help you iterate quickly with persistent storage, so you don't have to re-download your model and datasets each time. We give you a pipeline for moving your checkpoints into inference so that deploys and evals end to end can be done seamlessly. The goal is to cut out those little 15-30 minute tasks that slow everything down. You talk in the announcement about not wanting to create "yet another training product." How is this different? When we started building the product this year, we saw a lot of point-and-click solutions in the market. You'd have a dropdown of model options—train Qwen 3 70B, train Llama 3 70B—and you'd bring your data and use their training loop. But if you want to do something experimental, something that gives you a competitive edge because you're not doing what everyone else is doing, you might want to come to Baseten. We built a platform that's flexible enough that people training audio models—like Orpheus or Whisper—love coming to Baseten, because these point-and-click solutions don't really cater to different mediums beyond text-based LLMs. Can you talk more about that flexibility? Because we built an infrastructure product and we're not constraining which models you can select, we've opened up more options for customers. If you want to go out of your way and do something that's not in a dropdown menu somewhere, we support that. We also see people who need multi-node training for longer sequence lengths coming to Baseten. In addition, when things go wrong or you want to go deeper, our solution caters to people who are more hands-on. You get into the code, you're able to debug, you can tweak parameters—we're not limiting you to some set of knobs. It's built for developers, essentially. What about migration? What does it take to switch over? Say you're training on GCP today and having a tough time with infrastructure setup. You want to use multi-node on H200s, but DWS doesn't support that. You have an existing stack and want to expand somehow. Bringing that code over, bringing that training pipeline over to Baseten, is really simple. We built this with that person in mind—it's quick to switch over, quick to migrate. We're very unopinionated. You're not going to run into tons of abstractions or SDKs that force you to mold your code to fit our view of the world. You can take the code that works today and bring it to our platform, take advantage of on-demand compute, and use storage primitives that make your iteration loop tighter. How does Baseten's inference expertise play into this? Baseten is incredibly good at inference, and I've never felt like there was more fertile ground to build a new product. Everybody who's on inference at Baseten is looking to train, right? You're going to train that model and where do you want it to go? You want to serve your customers, build a differentiated product. One way we really stand out is our pipeline from training into Baseten's premier inference product. When we go into a call to talk about training, we actually start by asking: What does your inference look like? What do you need for time to first token? What do you need for throughput? What models are you using today? This helps us understand the use case and constrain the number of viable solutions. Before launch, you tested with customers, right? We launched closed beta on May 19th. A big day for Baseten because we also launched Model APIs. Everyone was in SF, we did two product launches. I couldn't believe it—everything...

Excerpt shown — open the source for the full document.

Notability

notability 6.0/10

New product launch from AI infrastructure company.