Interview

Illuminate raises $30M to build RL training environments that teach frontier AI models non-coding knowledge work

Oct 5, 2026 with Jerry Wu

Key Points

  • Illuminate raises $30M to build reinforcement learning training environments that teach frontier AI models non-coding knowledge work like financial modeling and PowerPoint design.
  • The company's gyms use deterministic scoring for verifiable tasks like spreadsheets and reward models for qualitative work, positioning itself between code compilation checks and human preference signals.
  • Wu observes RL environment complexity doubles every six to eight months, betting Illuminate captures expanding demand as labs push models into new categories of knowledge work.

Illuminate raised $30M to build reinforcement learning training environments for frontier AI labs, targeting the parts of knowledge work that have nothing to do with code.

Jerry Wu, co-founder and CEO, describes the product as Docker-containerized "gyms" — structured environments that give an AI agent a starting set of files, a task, and a scoring mechanism. A financial modeling gym might hand the agent raw data and tell it to build an LBO or a sensitivity table. The score tells the lab's post-training pipeline whether the agent succeeded. Wu frames Illuminate as a game design shop: the company builds the game, the lab runs it, and the model learns from the reward signal.

“What we do at Illuminate is we build the training benchmarks and training environments used by Frontier Labs to improve their models... we call [it] the Moore's Law of RL environments — every six to eight months, the complexity of the environment or gym used to train a frontier model roughly doubles.”

Verification and scoring

How scoring works depends on the task. Spreadsheet outputs can be verified deterministically by checking values and formulas. PowerPoint quality cannot, so those gyms use a reward model to produce more qualitative judgments. The architecture sits somewhere between the clean compile-checks used in code environments and the messier human-preference signals of RLHF.

The demand thesis

Wu argues demand for training environments is derived from a simple question: are there things models can't do today that we want them to do? His answer is that the surface area keeps expanding, and Illuminate's market grows with it. The company tracks what Wu calls a "Moore's Law of RL environments" — the complexity of the gym required to train a frontier model roughly doubles every six to eight months, by his observation.

Extrapolating that curve, Wu sees the product roadmap moving from single-agent spreadsheet tasks to multi-agent collaboration environments, and eventually to simulations of whole companies, governments, and industries. That last step is clearly a long-range vision rather than a near-term product plan, but it reflects where Illuminate thinks the training data problem is headed.

The $30M gives Wu the runway to scale environment creation across more domains. The bet is that as labs push models into more categories of knowledge work, the bottleneck stays on the supply side of well-designed, verifiable training environments — and that Illuminate holds that position.

Every deal, every interview. 5 minutes.

TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.