Interview

Tencent's head of AI model training launches BaseLabs to advance open-source continual learning and aggregate RL training environments

Sep 3, 2026 with Charlie O

Key Points

  • Charlie O, former head of AI model training at Tencent, launches BaseLabs as a non-commercial research initiative to build open-source infrastructure that closed-source labs have no incentive to create.
  • BaseLabs is betting continual learning and aggregated RL training environments will let open-source models close the capability gap without requiring every lab to independently replicate billions in closed-source R&D spending.
  • As open-source models reach commercial parity on most tasks, BaseLabs argues model values embedded through training data and safety stacks become the differentiator, and is developing methods to let firms explicitly shape those values.
Tencent's head of AI model training launches BaseLabs to advance open-source continual learning and aggregate RL training environments

Charlie O spent years running AI model training at Tencent before founding BaseLabs, a research initiative he describes as having no commercial agenda — just a mandate to make open-source models as useful as possible. The launch, announced during this conversation, is a bet that the open-source ecosystem needs infrastructure that closed-source labs have little incentive to build.

The central thesis

The frontier labs are all running the same recipe. Charlie's view is that pre-training scaled first, RL is scaling now, and everyone — OpenAI, Chinese open-source, American open-source — is following the same playbook. The differentiation he's chasing sits elsewhere.

The "god model" framing, he argues, misunderstands how intelligence actually maps to economic value. For any given task, there's an intelligence threshold below which you can't complete it and above which returns diminish sharply. Closed-source labs hit that threshold first, open-source follows six to nine months later, and at that point most users have good reasons to switch — control, customization, cost. The frontier labs keep pushing for use cases like frontier science and mathematics, where demand for intelligence is genuinely inelastic. Most commercial tasks aren't those.

We're gonna end up with probably hundreds of millions of models, possibly even one model for each person. We announced BaseLabs today, doing less myopic longer-term research around what models can actually do. The game of LLMs over the last five years: closed source hits capability thresholds first, open source eventually — six to nine months later — can do that task. And for many reasons, once you've had the base level of intelligence required, you probably do want to swap to open source.

Continual learning

This is BaseLabs' primary research bet. Charlie distinguishes between two very different things that get called continual learning. The big labs run a loop — train a model, release it, collect feedback, build RL environments to patch weaknesses, repeat. From a bird's-eye view that's continual learning. What BaseLabs is focused on is much narrower: an open-source model organically adapting to a specific firm, team, or individual over time, without needing to offload memory to markdown files.

The problem is that current tooling breaks down at that scale. Fine-tuning (SFT) degrades the model. RL doesn't deliver knowledge acquisition in the right way. Charlie is candid that BaseLabs doesn't have answers yet — he invokes Thomas Kuhn's framework for scientific revolutions, arguing the field is pre-paradigm on continual learning, with no agreed definitions and no settled approach.

On whether brute force solves it — just running the current training loop faster and faster until it feels continuous — his answer is conditional. At the RSI scale, maybe. For the legal associate fine-tuned on a firm's implicit behaviors and complex relationships, it breaks down long before you get there.

Aggregating RL environments

The second major bet is data, not compute. The open-source community talks constantly about aggregating chips to stay competitive, but closed-source labs spend billions annually building RL environments — and that spend is quietly becoming a structural moat. Charlie argues someone needs to build those environments without a commercial incentive and release them publicly, so any open-source or closed-source provider can train on them.

BaseLabs has been building environments for individual clients at Tencent for years. The plan is to generalize that work and open-source it — cutting the capability gap without requiring every lab to replicate the spend independently.

Model values and trust

As open-source models close in on closed-source capability for most commercial tasks, Charlie argues the differentiating question becomes values — what ethical and political assumptions are embedded in the model through pre-training data, classifiers, and safety stacks. He points to why many American companies won't use Chinese open-source models: the values embedded in them aren't legible. BaseLabs wants to develop post-training methods that let firms, countries, and individuals shape model values explicitly rather than inherit them implicitly.

Benchmarks

On the collapsing credibility of public benchmarks, Charlie's answer is to aggregate real economic usage rather than synthetic tests. The model: every company deploying LLMs has already built its own evals for the tasks it actually cares about. Aggregating those, rather than relying on ARC-AGI-style puzzles that labs can RL their way through once the benchmark is public, gives a truer picture of capability. BaseLabs sees its position — working across many deployments without the commercial pressure of a quarterly model release — as an advantage for building that kind of signal.

The non-commercial structure is the recurring thread. Every open-source lab, Charlie notes, still has investors and a quarterly model to ship. BaseLabs is explicitly designed not to. Whether that freedom translates into scientific traction on problems the big labs are ignoring is the question the initiative is built around.

Every deal, every interview. 5 minutes.

TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.