Tencent's head of AI model training launches BaseLabs to advance open-source continual learning and aggregate RL training environments
Sep 3, 2026 · Full transcript · This transcript is auto-generated and may contain errors.
Featuring Charlie O
Speaker 2: Next, we got Groceries. Space time. People are not talking about space groceries enough. Really make some metaphors there, I think.
Speaker 1: I'm not mistaken.
Speaker 2: The We got Charles from Tencent. The head of AI model training. How are doing, Charles? Welcome to the show. I'm your boy. What's going on? What's the best model you ever trained?
Speaker 3: Best model I ever trained?
Speaker 2: I can't talk about it. Name every name every model. The traditional answer. Didn't realize this. I think Name every has 5,000,000 models trained. Is that enough? Do we need more?
Speaker 3: Well, yeah, I guess our bet is that, you know, we're gonna end up in a world with some great models from the Frontier closed source labs. Sure. But we're gonna end up with, you know, not 5,000,000, but probably hundreds of millions of models, Possibly even one model for each person, possibly,
Speaker 2: you know, tens of models for each person. Like, I guess that's the future that base ten's betting on. Yeah. So what's the what what is the framework or what is the the the the next model, the thesis behind model development that you wanna execute on in this role?
Speaker 3: Yeah. So I think, like, to be honest, and the the closed lives will tell you differently, but everyone is scaling the same recipe right now. Like, there is no difference between OpenAI's RL stack and, you know, the Chinese open source RL stack and the American open source RL stack. Mhmm. We've got the same recipe. People believe it scales. Mhmm. We first scaled pre training and model size. We're now scaling RL. Yeah. And you know, people are gonna keep doing that till the cows come home. And the bet is that, you know, you hit higher and higher levels of intelligence and you'll be able to do more and more economically useful things. I think to us, like we care a lot for instance about continual learning. So we announced BaseLabs today, which is going to be doing like less myopic longer term research around what models can actually do. And continual learning looks very different for us as it does to the big labs. Like the big labs already in this kind of continual learning loop of, you know, they train a model, GPTN or ClaudeN, and they release it into the world, they collect feedback on what it can do, what it can't do. Yeah. Then they build a fuck ton of RL environments to like patch the holes in it. Yeah. And they go back and like from a high level view, bird's eye point of view, that is continual learning, right? Yeah. The continual learning that you or I might imagine is very different. It's where, you know, you have an open source model that you specifically are using either as a firm or a team within a firm or even an individual, and that model is organically adapting to the information that it's learning. It doesn't have to like, you know, write everything down in memory markdown files because it learns about, you know, your business and so on. And then we think there's like major paradigms to be unlocked within that regime, and it's not necessarily gonna get the focus from the big labs because that doesn't necessarily benefit them. Right? Like, wanna serve one big model at scale, and the pitch they've they've sold to their investors is that, you know, you do that at a large enough scale, and you don't need this this kind of continued learning.
Speaker 2: Yeah. How how do you how do you process their claim? Because a lot of the big labs, there's this dance between, oh, well, there's a fine tuned model. It was really good at this one benchmark. Then the next version of the big model that does everything is better than the fine tune, and it feels like this horse race where I can totally see the cost argument, and I can see even, like, the yeah. For three months, this fine tuned model was better. But but if we're talking just, like, raw capability, how how do you how do you interrogate that claim that the big god model will, you know, always be better as long as you give it time?
Speaker 3: This is where I think people are thinking about, like, intelligence capabilities in the wrong way. Okay. Like, people think about intelligence relativistically. They say, okay, the open source gap is like six months behind closed source and you know, GLM 5.3 is at the point that Opus 4.8 was, whatever. I think the best way to think about what models can do for you and for the world is like absolutely. So for any given task that you want to do with an LLM, there is some intelligence threshold that below that you can't do the task and above that you have very diminishing marginal returns to more intelligence on the task. And so when you think about it that way, the game of LLMs over the last five years has been, okay, we have these things we wanna do with them. Closed source hits it first, which to be honest, we think is a good thing for many reasons, which we can get into. But, know, open source eventually, like six months later or nine months later or whatever it is, can then do that task. And then for many reasons, like whether it's to control your own intelligence, whether it's to improve that task specifically, once you've had the base level of intelligence required to do it, you you probably do want to swap to open source. And so it's not really about, you know, the god model being better. Like, if I'm Yeah. You know, filing a tax return, there is just a limit to, like, how much intelligence I need to do that particular thing. And so I think the world is gonna look like, you know, the frontier closed source labs are going to continue to push the frontier. Like we you do want to use the most intelligent model. You have very like, inelastic demand for intelligence when you're doing, like, you know, frontier science or frontier maths. But for a lot of the economically valuable things, it looks a lot like, okay, you know, I'm a cursor or I'm one of these big companies who are now realizing, like, I can't just be a rapper anymore. I've been through the life cycle of building a great product that people love, and they tell me what they love and hate about it. And I should be using that information to, you know, make my model better at the things that I care about and not at anything else. Yeah. And so that that pattern was executed fairly well with Composer. It's obviously early days, the ecosystem isn't mature enough for anyone apart from, you know, the curses of the world and a few other big companies to to go about training. But, you know, I think eventually we'll get there where it's where it's much more organic. Can you talk about the the the goals of this project? I mean, I I I understand the the research direction, continual learning, but is the goal
Speaker 2: more to to advance the research or produce a continual learning model product that is then sold? How are you balancing the discussion that needs to happen within the ecosystem versus productizing something that could be a really meaningful breakthrough?
Speaker 3: Yeah. No. It's a good question. I think, like, the founding kind of mandate of BaseLabs is do whatever we can to make open source models and probably bleeding into closed source models eventually as useful as possible. And so it's very difficult to specify a priori what that looks like. Like one obvious place to bet is yes, this continual learning paradigm and doing research around that, and like thinking about that in a different way to how, you know, the big labs think about continual learning on like, you know, millions of people at once. Yeah. But like, we're not hubristic enough to say that like, that is the only way in which we're going to make models more useful. So for instance, another thing we're thinking a lot about at the moment is everyone talks about aggregating compute for open source to keep up, like you need a certain number of chips to be able to train these things, but not as many people are talking about aggregating data. Like the closed source labs are spending billions a year on RL environments, and again, like RL is the next paradigm that everyone's scaling. You know, distillation and like, you know, cheap data or centralized data, there's a lot of discussions around these things about how the open source labs are currently keeping up, but, you know, what keeps me up at night is I wonder if there's a point where it bifurcates and you can't like get all the way there from, you know, distillation and whatever data you have access to. Like someone needs, without a commercial incentive, needs to be making these incredibly complex RL environments that, you know, normally cost billions of dollars, in the aggregate, like, and open sourcing them to the world. So a key part of, I guess, based labs is like, can we make those environments? And we've been doing this for the last few years, like we're making them for individual people, why don't we just make them for everyone and release them? So, you know, any open source or closed source provider can train on them and you kind of like cut the gap that way. So aggregating data is another big bet that we wanna make. And there's a there's a few other things, know, like one one thing that we're starting to notice a lot is like as the capabilities have risen and, you know, even the open source models have kind of subsumed all the economically valuable tasks that like most people are going after, and it's very difficult to tell the difference between an open source and a closed source model, then you start to think about what the values of these models are, and you can't not have an opinion when you're training a model about what the values and ethics and morality of that model is. Like in your pre training data and your mid training data and your post training and your classifiers and your safety stack you run on top of it, everyone has an implicit or explicit opinion about what the values are. It's a reason why so many American companies don't wanna use Chinese open source is because we're not very clear what those values are. Sure. How do you organically shape those values to what you want them to be? Both as a country, as a firm, and as an individual, like, can we do post post training to elicit the right values and models?
Speaker 1: Trust in benchmarks feels like it's at an all time low and maybe is headed lower. What what are you what what what other do you have any other ideas of how to communicate the sort of may maybe it's still that they they work well enough, but
Speaker 3: any other ideas around how to communicate Yeah. It's the company of certain new models. Bench maxing, bench hacking, and then also just, like, inundation with so many that a lot of consumers and enterprise buyers, sort of just roll their eyes when they see a new one because they're like, I got I've seen a 100 of these. I I I know that you can probably do well on one. But, yeah, I'd love to hear this, that you're you're thinking about I think that let's think about, like, who makes benchmarks. Right? Mhmm. Like, specifically a bunch of, like, you know, research fellows from Stanford or what have you who've pulled together and they're very smart people. Yeah. And they're like, okay, this is what you know, this is this weird way in which we're gonna trip this model up. Like, think ArcGI is, like, a very good example of this, like, weird pocket where the models like traditionally don't do well, and they've obviously updated the benchmark more and more as the labs have kind of caught up. The common saying is like, you know, once you've benchmarked it, you can RL on it. So, you know, releasing these makes the models better. I think the best way to do this is aggregate the economy. And this is where we actually see us having a really big advantage compared to particularly the closed source labs. Like I do believe Anthropic and OpenAI are generally good. And like, you know, they don't look at user data if they say they don't. And like, they have this very like bird's eye view of like, the what models are good at and what they're bad at. It's why they go on Twitter and ask for feedback. We have a real advantage in the sense that we have like, you know, thousands of people using them for different things, and each of those individual companies has tried in some way to build their own evals and benchmarks to test how good a particular model is at the task that they care about. And so rather than, you know, like there's no like silver bullet here, like there is no set of 10 benchmarks that is going to tell you, you know, exactly the core and the jagged frontier of whatever model it is you're testing. The best way to do this is to look at how everyone is using them in the economy for real things that people are paying for. Yep. Not weird, like, niche things like ArcGI and aggregating those benchmarks. And, like, I guess that's part of what we're trying to do with the RL environments and pumping them out at scale. But, yeah, like, we we get a lot of really good signal from our customers. For the record here at a TBPN,
Speaker 2: we had Tyler do ArcGIV three tasks manually as a human. He technically got paid for it. So, you know When did you score a tie? He was he was he was globally ranked. Globally ranked. For for, like, a week. We were very early. Think day day of launch, he was, like, number seven.
Speaker 1: Charlie, what does it mean to feel the whole elephant? Oh,
Speaker 3: Great one. I've I've gotten a lot of flack for this, and people have interpreted it in quite inappropriate ways, including myself, actually. Yeah. But there's this old analogy of, like, you know, blind men touching an elephant, and, like, they don't really know what it is they're looking at, and one's feeling the trunk, one's feeling the leg, one's feeling the tail. I think it is a little bit like that with things like LMs and particularly like continual learning. You know, some people say that continual learning is just like having a god model with, you know, a million tokens as a context and it can search whatever it needs to and organize the information that way. Other people really believe that, no, every single token it should be updating, you know, its information. We know that that fries the model in different ways. I think Thomas Kuhn, wrote a lot about like the structure of scientific revolutions, and I think we are pre paradigm when it comes to things like continual learning. Like no one can agree on what the definitions of these things are. No one can agree on, you know, apart from the core recipe to produce the base LLM in the first place, like what the right ways to be scaling this is, you know, the ecosystem around it. Like compaction is a great example of this. Like OpenAir and Anthropic have gotten really good at compaction in Claude code and Codecs. At the moment, that's just summarizer models. Like there's probably like really, really cool things that you can do if you train like, you know, neural compactors, if you have models themselves doing compaction in KB case space. So I think this is a lot of work to like make this a science because so much of it is inside the closed source labs. And I guess that's mission of BaseLabs is like, how can we bring as much of this into the open as possible without a commercial agenda? Like even the open source labs have a commercial agenda. They need to produce the best model this quarter, and we don't have that pressure. And so, like, I guess we can pick the most interesting scientific problems and work on them for as long as we need to before we feel like we've made traction. What is your Very cool. P,
Speaker 2: brute force, the probability that the answer to continual learning is just brute force. When I think about the what OpenAI and Anthropic do for updating a model, you know, you mapped it out. It's a it's a couple months of data collection, building RL environments, retraining the model. That takes GPU hours. But if you get a 10x speed up in the amount of time it takes to train the actual model versus you can build the RL environments 10 times faster. You can generate the data 10 times faster. And you start working at this over maybe it's a decade. Is there a world where you could just do exactly what we're doing now but every second? And it feels continual because when I talk to anyone, they're continually learning, but they take a second. They say, oh, yeah. Okay. I'm updated. I I now know that fact.
Speaker 3: Yeah. I I think there's a few ways to write this down. From the perspective of RSI, like, recursive self improvement Sure. I actually think that, you know, continual learning is, like, either unnecessary or it is brute force Sure. Like, you know, we're going to asymptote towards models that have access to the whole, you know, OpenAI and Anthropic training stack. They're gonna make slight architectural improvements. They're gonna drop pre training loss. They're gonna be able to pump out our own environments, scale by themselves. And then that process will speed up and speed up and speed up as we get more and more compute. But for the everyday person and the everyday, like, you know, AI native startup or enterprise or whatever, continual learning doesn't look like that. Like at the scale that they're operating at, you can't afford to, like OpenAI and Anthropic, they wash out all the noise and the gradients, they take all the useful stuff and these big pre trained runs makes it work. But when you just focus on like, you know, I have an agent which is a legal associate and I'm trying to fine tune it on all these like complex relationships and all the things that the firm does and this implicit behavior that we want to have, continual learning breaks down. We don't have an answer to it. We can't SFT. It degrades the model. We can't RL because it doesn't give knowledge acquisition in the right way. You know, there's there's a lot of work to be done there. So I guess it depends on, again, which part of the elephant you're touching. Like, what definition of intuition play do you care about? We gotta get an elephant in the onboard.
Speaker 1: Dog. Studio. Horsesh.
Speaker 2: Part of the elephant you're touching. Thank you so much for coming on. This is a fascinating discussion. Congratulations, LaProsse. Yeah. Great to meet you, Charlie. Very excited for you to solve this. Come back soon. This was fun. We'll talk to you soon. Have a great day. Goodbye.