MTMTS Jul 28, 2026· 14:23

Open Source Can't Be Stopped | Charlie O'Neill Explains

Charlie O'Neill of Baseten argues that open and closed source AI labs can coexist, with Kimi K3 proving open source models can approach frontier capability. He explains the Nvidia open letter's significance as a landmark moment acknowledging this coexistence. O'Neill notes that Kimi K3's training on persistent agent environments—coding, terminals, simulated workplace tools—marks a shift from training on answers to training on work. He claims post-training becomes counterintuitively cheaper when starting from a frontier baseline, as less effort is needed to fix general deficiencies. He predicts 3 to 5 new 2 to 4 trillion-parameter open source models from American labs in the next six months, all engineered for long-running tasks.

Transcript

Opening0:00

Charlie O'Neill0:00

People are now going to start pushing past the trillion-per-minute range, particularly in America. Um, we've been kind of been optimizing around that for the last few years, and we haven't really broken through that. But now the baseline is kind of like, okay, your frontier offering needs to be a multi-trillion-per-minute model that is engineered for long-running work and long-running tasks.

And in order to be, like, useful to the open source ecosystem, it needs to be geared around post-training. And I think we will see, you know, probably at least 3 to 5 new 2 to 4 trillion-per-minute models from open source in the next 6 months, which is just incredibly exciting.

Harvey0:37

We are live with Charlie O'Neill, who is an engineer at Baseten. Baseten is a really interesting inference company. They've done some great parties. I have a copy of the Baseten Inference Engineering book at my house.

Nvidia Open Letter0:37

Host0:52

I do also have one. Yes.

Harvey0:52

Yeah. So that's great. We love Baseten. Charlie, welcome to MTS.

Charlie O'Neill0:57

Thanks, Harvey.

Harvey0:58

Absolutely. So the Open Weight Model Wars have been going on for some time now. Um, Baseten was one of the signatories to the Nvidia open letter on open source. Um, can you tell us a little bit about what this letter means, why it's important?

Charlie O'Neill1:17

Yeah, I think the letter was kind of the landmark moment. I mean, in the history of things, like signing petitions and letters, it doesn't often end up mattering. But I think it was important because it showed that a lot of the key players and the giants in, like, the, I guess, the AI ecosystem are starting to wake up to the fact that having a monopoly or a jobly on intelligence is probably going to be a bad thing.

Um, it's bad, obviously, for the bottom lines, but it's also bad for the state of the world in general. Um, and, like, the fact that people like OpenAI, who obviously their whole business model is based around closed source models, are signatories to the same letter, was really promising.

Um, and to me, it was the first kind of, like, public discourse where we kind of acknowledged that open source doesn't have to entirely win and wipe out, you know, the big frontier closed source labs and vice versa.

Um, but there is a world in which they coexist. And, you know, we have a bit of a bifurcation where, of course, closed source keeps scaling up and they're being used for the most frontier problems and the most frontier challenges, like science.

Um, like, periodic labs is a great example of this. But at the same time, a lot of the economically valuable tasks in the world are going to be done with either off-the-shelf open source as it continues to get better, or specialized open source, because you can specialize open source and post-train it for the particular things you want to do.

And so that's really exciting to us. I think, like, you know, we're very supportive of a world in which both of those players win. And I think, uh, for the first time, we're starting to realize that that is a real possibility and probably the good timeline we want to end up in.

Harvey2:48

Yeah, we were just talking about this earlier, how, uh, the open source maxis are, like, a little bit silly. And the world in which both open and closed source ecosystems thrive is probably the best one.

Charlie O'Neill2:59

Yeah, 100%. And, you know, there's arguments against both sides of the extremes,right? Like, you probably don't want a world in which there isn't this, like, responsible development, um, at the frontier. And you kind of get this canary in the coal mine,right?

Like, of course, closed source labs are going to have more compute and more data and going to be slightly ahead on the scaling front, um, compared to the open source labs. But that could actually be a good thing,right?

Like, if they're seeing what the behaviors that are emerging as these things continue to scale up and they're able to, like, at least provide a warning to the ecosystem about what that looks like, then I think we can do open source development much more responsibly.

Um, and then conversely, of course, like, if you just have the closed source, which is dictating, okay, we're seeing these emerging behaviors and a small group of a few thousand people are deciding who's going to access those behaviors, how they're going to be accessed, um, how much they're going to charge for it, I think that's also a bad play.

Um, so yeah, like, I think the one thing that's been established, though, which is probably now becoming a lot clearer to people, is that despite, you know, OpenAI and Anthropic and the closed source labs having more compute and more data, the recipe is the same.

Like, we're all scaling the same recipe here, and we're at slightly different points of the scaling curve. And it's not like OpenAI or Anthropic have any secret source, um, that really differentiates them scaling up their multi-trillion-per-minute models from the rest of the open source ecosystem.

It really is just a matter of taking the core underlying architecture and training recipe, um, and throwing more compute and data at it, which is really promising. Um, of course, like, these closed source labs probably have a long tail of optimizations they've made, um, which we're going to continue to try and figure out, um, in the open source community.

But the core mechanism is there, and that's really promising. And I think K3 is a great example of, well, if you do scale this recipe up, we're going to get a pretty good model that's very, very close to the frontier, if not frontier itself.

Harvey4:53

Hmm. That sounds kind of surprising if true. Like, uh, the frontier labs really, you know, have no sort of algorithmic advantage over Kimi or DeepSeek or these other companies. Actually, I think that makes sense.

Charlie O'Neill5:08

Yeah, I think there's a few data points here. Um, I don't think the claim would be that they have no algorithmic advantage. I think the claim is that most of the variance is explained by having the roughlyright architecture, which is also servable at scale.

So obviously, Kimi has a lot of optimizations which allow it to serve efficiently, you know, 1 million token context length, and there's a lot of things you have to discover there. But really, like, these neural nets are universal function approximators.

And, like, as long as we get things roughly the same form and you train on roughly the same magnitude of data and compute, you're going to end up with something very similar. So I do think that, you know, the closed source labs probably still do have optimizations and relatively important optimizations we haven't figured out.

Um, but it's not like those optimizations account for a huge variance in the capability or intelligence of the model. It's kind of just like the last few percent on a benchmark that you'd be looking at getting.

Harvey6:00

Hmm. Makes sense.

Host6:02

Yeah. So you guys are day-zero launch partners with Kimi K3, which is awesome. Um, what do you think this means for builders?

Kimi K3 Launch6:02

Charlie O'Neill6:10

I think, yeah, it's twofold. First off is kind of the matrix of, you know, price, capability, um, and control. Like, K3 is probably now the best option in the world. Like, if you can not only, like, host it yourself, but also host a model yourself that is, like, almost indistinguishable from the very, very closed source frontier, like Fable and GPT-5.6 SOL, like, that opens up a whole new world of use cases.

Um, obviously, there's, like, many arguments about when you might want to host a model yourself if you're a large enterprise. Um, there's a lot of security and control and access things that you have to think through. But at least now we have a proof of concept for the fact that an open source model can get to this level of intelligence.

And I think we're going to start seeing a lot of, like, larger and larger enterprises and not just AI-native startups hosting these models themselves, post-training them themselves. I think the second big thing is that post-training, it's a 2.8 trillion-per-minute model.

And I think the first order reaction to that is that, oh, there's only going to be a very small number of players in the world that are able to post-train this or, like, even have the bandwidth or, you know, desire to post-train a model so large, um, given how good it is out of the box.

But I think, like, if you think about that a bit more carefully, what actually what we actually see is that post-training is really about making up for deficiencies in the model and then adapting it to the specialized tasks you're trying to adapt it.

And when you start at a really, really high baseline, we actually don't need to spend anywhere near as much time trying to, um, recover from the general deficiencies in the model. So it's really exciting for us to have a model that is starting at the frontier as a baseline, and then we only have to spend a little bit of effort, like, you know, improving the efficiency or getting a bit of domain knowledge in for the specific thing we're trying to train it for.

So post-training can actually end up being, like, weirdly counterintuitively perhaps cheaper than some of the much smaller models, um, particularly for, like, more complex tasks, um, in this case. So that's another thing we're really excited by. And I think in the next few months, you're going to see a lot of people, um, making small tweaks and small post-training runs.

Um, small being in relative terms, it is still almost a 3 trillion-per-minute model. Um, but we are going to see this burgeoning of, like, you know, an ecosystem of specialized Kimis. And as, you know, Inklings and Nematrons get better from the American open source, we're going to see people, like, start to apply those recipes to those models as well.

K3 Report8:31

Harvey8:31

Right. So you guys have also had, uh, access to the Kimi K3 paper and other technical information for at least a couple days. Um, you know, we've had it for like an hour. What were some of the most interesting things that you saw in the technical report?

Why has Kimi been able to do so well?

Charlie O'Neill8:50

Yeah, I think the big one for me was that they gave Moonshot, like, train K3 in persistent agent environments. And it wasn't just on, like, you know, prompt response status. They gave the model very long-running jobs, like coding tasks, terminals, um, simulated workplace tools.

Like, they've made copies of very, very common workplace tools, which is pretty cool. And it has to act and observe results and, you know, recover from mistakes and so on and get the environment into theright final state. And that's kind of a meaningful shift from, like, training the model on answers to training them on work.

And I think this is the direction that everything is moving towards, which is training on work. And of course, there's architectural changes as well. I think things like attention residuals and, like, the way they've improved, um, the attention mechanisms they're using.

And, like, this is all work that's been carried over from previous lineages of Kimi models and also, like, other open source models, which is great. We're starting to see this, like, kind of information kind of diffuse across the open source ecosystem and people are picking the best parts of the recipe that work.

Um, but for us, it was really about, you know, like, the particularly the post-training and the reinforcement learning they've done has been on really, really long-running work tasks. Um, their approach is generally kind of break work up into distinct kind of areas.

So you might have coding, general reasoning, and then, like, general agentic work. And then within those, they train many different teachers, which is just very good at one particular thing. And thenright at the end, they take all those teachers and use it to teach one student model after the pre-training run.

Um, and this is a recipe that we've seen with other models like GLM. Um, but yeah, it started to really work at scale. And we're now approaching, you know, when you think about the meter, it evolves over a really long time horizons.

Like, these open source models are now gearing around training themselves to be able to do these really long-running tasks. Um, and that is the first class citizen of post-training rather than something that's added in as an afterthought and really focusing on pre-training.

Uh, so that was the exciting thing for me.

New Floor10:45

Host10:45

Yeah, I guess last question before we wrap is, like, do you think this has basically set a new floor for open source models? And then how do you see, you know, like, post-training evolving from here and especially what you guys do at Baseton?

Charlie O'Neill10:58

Yeah, it's definitely set a floor. And I think it's also a very promising proof of concept for all the other model creators in the ecosystem. I think, like, to use an analogy, like, when we know a statement is true or not, it often becomes much easier to prove than if we're not sure if it's true in the first place and we're fumbling around for a proof.

And I think, like, we're starting to see this within particularly the American open source ecosystem, which, you know, for the past few years has been behind, but really is starting to catch up. I mean, thinking machines trained their first model and it was actually a fantastic model.

Um, and, you know, the Nematron models are getting better all the time. So I think, like, what we're going to see is, like, people are now going to start pushing past the trillion-per-minute range, particularly in America. Um, we've kind of been optimizing around that for the last few years and we haven't really broken through that.

But now the baseline is kind of like, okay, your frontier offering needs to be a multi-trillion-per-minute model that is engineered for long-running work and long-running tasks. And in order to be, like, useful to the open source ecosystem, it needs to be geared around post-training.

And I think, you know, Inkling from Thinking Machines is another great example of that. Like, they've explicitly said that the design of the model's capabilities was to make a very good all-round general-purpose reasoning model that you're able to adapt to your own particular task and whatever modality that is.

And that's really exciting for me. And I think we will see, you know, probably at least three to five new two to four trillion-per-minute models from open source in the next six months, which is just incredibly exciting.

Harvey12:25

Yeah, well, we are excited.

Host12:28

We will see what happens.

Harvey12:29

We will. Thanks so much, Charlie, for coming on MTS.

Charlie O'Neill12:32

Thanks for having me.

Outro12:32

Harvey12:36

MTS is an X-Native live streaming news and interview show covering technology, business, politics, and culture as it happens. Catch us every weekday live on X, YouTube, or wherever you get your podcasts.

Scale your startup on Neon, the Postgres backend for apps and agents built by Databricks. neon.com/mts.

ElevenLabs - AI that communicates at human level across every channel and modality. ElevenLabs.io/mts.

Special thanks to our sponsor, Kong. Kong, the AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. Try it at konghq.com.

Thanks to Blitzy - Autonomous software development for enterprise codebases. Ship 5x faster at blitzy.com.

Harvey13:36

Let me tell you about Merge. OpenAI, Dropbox, and Ramp use Merge to get AI to production faster. Merge.dev/mts.

Quick word from AdQuick. Make your brand a billboard. Out-of-home advertising as easy to scale as digital. adquick.com.

Charlie O'Neill13:54

Thanks to our sponsor, Macroscope. AI codebase understanding for engineering leaders. Try it out at macroscope.com/mts.