BaseLabs launches Space Labs research initiative to open-source RL training environments and solve continual learning
Sep 4, 2026 with Charlie O
Key Points
- BaseLabs launches Space Labs, a non-commercial research initiative focused on continual learning and open-source RL training environments that closed-source labs won't prioritize.
- BaseLabs is building public RL environments to close the data gap in open-source AI the way compute aggregation addresses the chip gap.
- Charlie O argues open-source models crossing an intelligence threshold for common tasks become rational choices over frontier models, shifting competition to model values and real-world usage signal over benchmarks.
Summary
BaseLabs launches Space Labs research initiative
BaseLabs is betting that the next meaningful advance in AI isn't a smarter frontier model — it's teaching models to keep learning after they're deployed.
Charlie O, BaseLabs founder, announced Space Labs, a long-term research initiative focused on continual learning and open-sourcing RL training environments. The mandate is deliberately non-commercial: produce research that the big closed-source labs won't prioritize, and make it available to anyone.
What continual learning actually means here
The term means something different at BaseLabs than it does at OpenAI or Anthropic. The frontier labs already run a form of continual learning — train a model, release it, collect feedback, build RL environments to patch the gaps, repeat. Charlie argues that's a bird's-eye process designed to serve one large model at scale.
What BaseLabs is working toward is something more granular: an open-source model that organically adapts to a specific firm, team, or individual without needing to write everything into memory files. A legal associate agent, for example, should absorb the firm's implicit behaviors and relationships over time — not require manual fine-tuning that degrades the model or RL that doesn't deliver knowledge acquisition in the right form. That problem, Charlie says flatly, is unsolved.
“We announced Space Labs today which is going to be doing less myopic longer term research around what models can actually do... Someone needs, without a commercial incentive, to be making these incredibly complex RL environments that normally cost billions of dollars in the aggregate, and open sourcing them to the world. That's a key part of BaseLabs.”
The open-source data gap
The initiative has a second strand that Charlie argues gets less attention than compute aggregation: data. Closed-source labs spend billions building RL environments. Open-source labs survive on distillation and whatever data they can access, which may work until it doesn't. BaseLabs is building complex RL environments and plans to release them publicly so any lab — open or closed — can train on them. The goal is to close the data gap the same way compute aggregation efforts try to close the chip gap.
Intelligence thresholds, not relativism
Charlie pushes back on the standard framing of open-source models being "six months behind" closed-source. His argument is that for most economically valuable tasks, there's an intelligence threshold above which more capability delivers diminishing returns. Once open-source crosses that threshold for a given task — filing a return, drafting a contract, writing code — the rational move for most companies is to switch, because they can then fine-tune for their specific use case rather than paying for general-purpose headroom they don't need. Frontier closed-source stays relevant for frontier science and math, where demand for intelligence is inelastic.
Values as a competitive variable
As open and closed-source models converge on capability for common tasks, Charlie argues the differentiator shifts to model values — the ethics and priorities baked into pre-training data, classifiers, and safety stacks. Every lab has an implicit or explicit position on this. The reason many American companies won't use Chinese open-source, he says, is uncertainty about what those values are. BaseLabs sees post-training methods for shaping model values as an underexplored research area.
Benchmarks
Charlie's view on benchmarks aligns with where the broader market is heading: low trust, diminishing signal. His preferred alternative is aggregate real-world usage data — what thousands of companies are actually paying for models to do — rather than niche academic evaluations. Once a benchmark is published, labs can RL against it and inflate the score. The RL environments BaseLabs is building are partly an attempt to generate more grounded signal at scale.
The absence of commercial pressure is the structural advantage Charlie leans on. Closed-source labs need to ship the best model this quarter. BaseLabs can work on the most interesting unsolved problems for as long as it takes.
Every deal, every interview. 5 minutes.
TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.