Illuminate raises $30M to build RL training environments that teach frontier AI models non-coding knowledge work

Oct 5, 2026 · Full transcript · This transcript is auto-generated and may contain errors.

Featuring Jerry Wu

Speaker 2: and security. Our next guest is Jerry Woo from Illuminate. He's the cofounder and CEO. Jerry, welcome to the show. Thank you so much for taking the time to come chat with us. Please introduce yourself.

Speaker 6: Thanks for having me. My name is Jerry. I'm the founder here at Illuminate. What we do at Illuminate is we build the training benchmarks and training environments used by Frontier Labs to improve their models. Specifically, we specialize in improving models on non coding knowledge work Mhmm. Starting with areas like financial services, accounting, consulting. Right? Like, a good way to visualize what we do is if you've ever used a model to not write code, but do things like build a spreadsheet, build a PowerPoint, help with your, you know, financial analysis or investing, a lot of those gains in financial and economic literacy have come from the environments and benchmarks that we've built and scaled with our customers. Yeah. I think we've heard a lot about building RL environments, selling them to the labs. They run them.

Speaker 2: But I'm interested in, like, what does that environment actually look like? Is it, like, screen recordings while you hire a investment banker to build a DCF model in Excel and you're just sending in, like, raw data? Or are you, like, architecting system in Python or something like... What's the actual product look like?

Speaker 6: Yeah. It's a great question. So fundamentally, what an environment is Yeah. I like to use this analogy. It's like a classroom or a game. Mhmm. Right? And what we really are is we're game designers. Okay. And we are building games or environments that are used to teach the models specific tasks in the verticals or domains that we're experts in. So let's take financial modeling. Right? We don't screen record someone to create the gym. That that is a way to collect data. But really, the the product of the gym itself is essentially a Docker container, where within the classroom, give the agent, let's say, some starting Excel files. Yeah. Some starting PowerPoint. Yeah. Some raw data, and you give it a task. For example, build an LBO or build a sensitivity table. Yep. Right? And then just like a game, there's a score. Yeah. Right? This is also known as the verifier where you can essentially get signal on whether or not the agent successfully achieved the task that you designed for it within the game. This signal is then used as a reward signal within the post training pipelines for our customers. Right. So when you think scaling RL Yeah. It's it's not quite like scaling screen recordings or scaling trajectory. It's basically scaling as an industry, the creation of these games or these classrooms on thousands or hundreds of thousands of different domains, tasks, environments to teach these models collectively how to do this work at scale. And is that score

Speaker 2: deterministic? Is that program... Like, programmed by you or is that its own classifier?

Speaker 6: Yeah. It's a great question. It can vary depending on the exact recipe or the approach that we want to take. Mhmm. Right? Traditionally, you know, for example, in code, a lot of verification is deterministic. Right? So you can check whether the code compiles or Yep. For example, with math, there's the lean compile, etcetera. Yep. For things like knowledge work, right, building a spreadsheet, this is actually somewhere in the middle. So you could theoretically verify correctness of a spreadsheet by looking at the values, pulling the formulas. But for something like a PowerPoint, it might not be fully deterministic. You might need to use some more qualitative judgments powered by, say, a reward model Mhmm. Right, to give you a score. So it it completely depends, right, on the task and the gym and the environment that you wanna build. Yeah. You just raised some money. We should hit the gong.

Speaker 4: Tell us how much you raised. Let's talk about that, and then I have a follow-up question.

Speaker 2: Much you raised?

Speaker 6: Raised $30,000,000. Boom.

Speaker 4: Great hit. I think that was the best of the day. Best of the of the day. Hands down. Could you replay? Could you raise a fund? Yeah. Let's let's check the replay. I I wanna I wanna... If we can just confirm the quality of the hit before we move on, I think that'd be great. Little little delay there. Move away. Come come back around quick. Woah. Coming around. There we go. You were kinda angry. Oh, yeah. Oh, yeah. You got it really. Boom. There we go. Great hit for Illuminate. I'm sure you got asked this nonstop during the fundraise, but I'm just curious. Where what what where do you see this? What is this business in in five years? Oh, yeah. Feels like really difficult for me to predict. Like, feels like there will still be work making the models better. I don't know what what form it'll take, but but how would you answer that?

Speaker 6: Exactly. I I think you hit the nail on the head there. The... Ultimately, the value of data research labs like ourselves comes from improving the models on something that it's not good at today. Mhmm. Right? Another way you can think about it is demand for companies and research and services that we provide are derived from the overarching question of like, are there things that models can't do today that we want them to be able to do? Insofar as that is true, and actually the surface area that is growing, the demand for data environments benchmarks will also continue to explode. Another way to think about it is as the surface area of these models increases, the things that we want them to be able to do also increases, and that is what drives a huge demand for a lot of the research around data environments benchmarking. I think another way you could also think about this is what does the actual data product look like? Previously it was human data and trajectories eventually evolved into like RLHF, and now it's like RLVR with these RL environments. Like, another question you could ask is like, you know, what does actual shape of of training look like in the future? We internally at Illuminate have a a phrase we call the Moore's Law of RL environments, which basically means like every six to eight months, we've noticed that the complexity of the environment or gym used to train a frontier model roughly doubles. So if you extrapolate this out, what you realize is that today we're building these gyms for spreadsheet modeling or PowerPoint building. We think in six to eight months, we're going be building gyms of multiple agents collaborating as a team, solving a task together. I think in the future, we're going to be building simulations of whole companies, whole governments, whole industries to teach models how to operate autonomously within our societies, within our companies, etc. That's our vision of what the shape of data will look like, increasingly complex. But, yeah, I mean, I think you guys have probably seen this from your reporting as well. The demand for data continues to skyrocket. Right? And it's driven fundamentally by our

Speaker 2: industry wide need to improve these models on so many categories of work and beyond. Very cool. Well, congratulations on the round. Thank you so much. Yeah. Thanks for breaking it down. For breaking it down for us. Have a great day. Thank you. Great stuff, Jerry. Bye. We'll talk to you soon. Talk soon. Goodbye. Our next guest is Brian Johnson. I think he's rolling any minute now. Got another minute? In the meantime, we can go back to the timeline. There's some all sorts of news all over the place. Some... David Ellison is hiring some folks. The team that leads Skydance

Speaker 4: forward. People are talking about the new Optimus factory Oh, yeah. At Giga Texas. It's progressing rapidly. It'll be 7,000,000 square feet with a targeted capacity of 10,000,000 Optimus per year. Initial production is planned to begin in 2027. David Hole says 10,000,000 humanoid robots is enough physical labor to build New York City in five months, and the first Optimus factory will manufacture that many robots in a single

Speaker 2: year. That's crazy.

Speaker 4: I like... We should build New York

Speaker 2: in California. In Texas. Is it great? I think it's going in Texas. It's gonna be in Texas first. I don't think it's gonna be in California anytime soon. He left. But the the Frenchman over at Atelier Misu followed him to Texas and they put up a statue Mhmm. Which people were going back and forth on saying it was disproportionate. Do you see this? Really? Yes. So apparently, statues should be like one... Like a ratio of head to body of one to eight. They went with one to seven, very controversial in the statue community. They made a lot of people angry. They had... There was a backlash. There was a backlash to the backlash. Lots of people going back and forth on whether or not the statue is aesthetically cool. The big scoop was that they've been accused of sort of LARPing as as like the old school way of making statues. They do in fact use computer aided design CAD tools, c... CNC machines to mill parts of the statue. They use three d software to design the statues. And a lot of people are frustrated by that because they feel like it's not. They didn't fully go back. I don't really care about that. I think the the original question was just Brian just dunked on Ben.

Speaker 4: He just pulled up on him, just dunked on him. Boom. Boom. Uh-oh. Come come on in if if you're ready. I I I wish we had the I wish we had the basketball hoop over here. Yeah. We could just