Ontology wants to automate scientific discovery end-to-end, starting with AI post-training R&D

Aug 3, 2026 · Full transcript · This transcript is auto-generated and may contain errors.

Featuring Ron Arel

Speaker 1: Have a good one. Have a good Cheers, bro. Up next, we have Ron Yarrell from Intology. He's the co founder. And we're talking about RSI, Recursive Self Improvement. Is it here? It might be. We'll get his take. Ron, how are doing? I'm doing well. Thanks for having me on with us. Thanks for hopping on. First time on the show, why don't you introduce yourself and the company a little bit?

Speaker 10: Yeah. Of course. So my name is Ron. I'm the cofounder of Ontology. At Ontology, our mission is to automate scientific discovery, and we're starting with AI r and d. Mhmm. So today, we have some pretty exciting results to talk about regarding automated post training of other language models Yeah. And lots more to discuss, so glad to be on. Yeah. Awesome. How would you characterize

Speaker 1: the the progress, the announcement? Because people can get sort of lost in benchmarks. Are you 20% on this thing, 99% on that thing? Qualitatively, where do you see the technology today?

Speaker 10: Yeah, absolutely. So fundamentally we believe that automating discovery is a domain agnostic problem. What that means is that the structure of discovery problems is pretty similar across domains in the sense of, you know, no matter what problem you're looking at, like, you know, drug discovery or materials discovery, even, like, improving language models, there's always gonna be some sort of process of proposing an experiment, getting feedback from that experiment Yep. Learning from that experiment, and then using that information to propose the next set and continuing until you, you know, make the discovery. Yep. So it's a, you know, a great question because, you know, we think about the problem on that access of how do we build and scale systems to these problems in which the evaluations and experiments become more expensive, more difficult, harder to access. And post training, for example, represents a problem space where experiments are pretty expensive. They can be quite noisy, and they take a long time. Right? So it like, you're trying to build an automated research system that, you know, develops the next state of the art, you know, 70,000,000,000,000 parameter model or whatever, you know, you can't really imagine a system training 15,000 models until it discovers the the the best one because training is expensive. So, you know, you have to think about how do you run efficient experiments? How do you gain knowledge, gain information from those experiments better? And I think that's, you know, today we're showing kind of a closer step in that direction because, you know, in the past we were doing things like kernel optimization, you know, like MLE bench style problems, now post training where, you know, obviously, it's more expensive and takes longer.

Speaker 1: On the cost side, this does sound expensive. If you wanna hammer a bunch of different experiments, what has been your approach? Just sort of suck it up and use venture capital dollars or partner with companies and labs that have big compute allocations? They have the resources. Have them sort of, you know, front the cost, or is there another way to solve it? Because scale seems very important here, and yet your whole job feels like burning compute.

Speaker 10: Yeah. So definitely a little bit of both. Okay. I would say at the beginning, it was a lot of burning venture capital money, and now we have partners and systems in place to run experiments that, don't just burn money for no reason. We have our own cluster, we have our own infrastructure that efficiently utilizes those resources. So we're at a place where we're not just throwing money at the wall for no reason. We're well, I mean, obviously it's an unsolved problem. We will always continue to work on making our system better at this. But I would say it's a mixture of having some great compute partners, you know, spending a lot of time on our infrastructure even before we start running experiments. And now we're at a place where, you know, I feel comfortable throwing, you know, hundreds of thousands of dollars at the wall if we feel like the system can actually make progress. Gary Marcus, he says that doesn't count if it's not a pure LLM,

Speaker 1: that AI is only making progress in verifiable domains. You need a lean output that's fully verifiable. How optimistic are you? I and I'm somewhat sympathetic to it because it does actually seem like we are moving much faster in verifiable domains like math than unverifiable and unverifiable domains like, I don't know, coming up with a script for the next great movie or joke writing or comedy writing or even some of the bio stuff that's maybe verifiable but over a multi year process of going through FDA applications and testing in vitro testing in mice and in monkeys and humans. There's just things that the verification loop isn't just run some lean really quickly and see if it checks out. Right? So, how what do you see the future of transferring all the amazing learnings and the ability for AI to make discoveries in ML, in computer science, in math to anything that's a little bit less verifiable?

Speaker 10: Yeah. For sure. So I think that it really just comes down to I mean, so we focus on verifiable domains, but Okay. I think it really just comes down to, you know, how much you can actually query the evaluator. Right? So back to that example of, you know, post training the next large state of the art model. You could argue that, you know, that system could in theory train the whole model, get the ground truth feedback Mhmm. See how it did, but that's not fully realistic for every run. Right? So it would have to do some sort of experimentation with either its own rewards or, you know, not full ground truth rewards before making that progress. I think it's my my take is I really think it's just about building systems that almost like wean off the requirement of the evaluator. So for example, you know, if you if you're building these kinds of systems for training small models, no problem. You train an infinite amount of models and you find the best one. But, yeah, event like, it's almost like as you scale on this access of difficulty, it's almost like as a forcing function of a compute availability or cost. Mhmm. You start having to think like, Well, now I can't actually have my system train the whole model every time, or I can't run an entire clinical trial every time I wanna test a drug. It's almost like a nature of the research direction. And so I guess that I I don't really think it's that different of a problem. I think it's more that if we continue on this path, like for example, with our system, if we continue on this path where it doesn't need to, you know, query the full evaluate every time, eventually it'll get to the point where it might not even need the evaluator, or it'll only need it at the end when sending a drug to clinical trial or a material in the fabricator. So I think it's really just I I guess we and collectively, the AI for science community, I think we're heading in the right direction. I I don't think I view it as, like, black and white as we're only doing verifiable, and then we're gonna go to unfair viable domains and, like, Interesting. Figure out

Speaker 2: I have one more question. Jordy, do you have anything? Yeah. I was just gonna ask, what do you what do you think your business looks like two years from now? I won't say five or ten, but, like, where where where is this work going?

Speaker 10: Yeah. For sure. So, you know, we I guess we really believe in not building copilots. You know, what we are really trying to do here is build fully autonomous systems that are deployed in r and d environments and just run the entire loop, you know, autonomously perpetually. Obviously, we're, you know, we're we're quite far away from that right now. You know, we have to deploy the system. We have to monitor it. We have to see it work, make sure it's succeeding. I think at the end of the day, definitely two years from now, we wanna be at the point where, you know, our system can basically be deployed as infrastructure in any computational r and d problem, and then it, you know, it gets access to the data on the problem. It gets access to the ability to run experiments on that problem, and then it gets deployed in an environment in which it's, like, you know, hard to hack and you're getting good signal out of the experiments. And then boom, you know, it runs autonomously. It's it's cranking out discoveries. It's shipping them, and, you know, humans can be there to take a look and make sure it's not messing up. But eventually, we want it to, you know, be running end to end. Last question. Very cool. Tell us a little bit about the company, where you're based, how big is the team, who are you hiring? What you were doing before this? Yeah. What you're for the shape of the company. Yeah. For sure. So we're based in San Francisco. We just moved into our new office. That's why my background's pretty boring right now. We're, you know, we're we're growing pretty quickly and been really proud of the team we've been putting together. I mean, we have researchers, you know, coming from DeepMind and Thropic, Factory AI, both in, I guess, industry and in academia. I think this is like a really, you know, fundamental problem that needs to be worked on, and it's not just a research problem. It's not just an infrastructure problem. It's, you know, it's the whole stack, and we've been putting a pretty incredible team together to do so. And I guess before this, my co founder and I ran a pretty large nonprofit research group. We were primarily funded by the National Science Foundation, obviously very different from running a company nowadays, but we did a lot of fundamental research in coding language capabilities, published some of the first work in test time scaling with language models. It kind of felt like a natural progression because, you know, we were really curious about how do we model agentic behavior in this kind of search process. And it kind of just made sense that, you know, this is the time, this is the place.

Speaker 1: Let's make it happen, and that's what we're doing in Ontology. Yeah. Well, thank you so much for coming on the show. I'm sure you'll be back on soon. Have a great rest of update. Well, nice you. Goodbye. Cheers, Ron. Let's go to this Paul Graham post. He had such a wild experience as he bought bought a book. It was awful. Didn't want it on my shelves, but I couldn't throw it away. So it sat on a table near the door. I know someone

Speaker 2: that will rip it apart, feed it to a machine,

Speaker 1: and burn it.

Speaker 2: Know someone I don't know them personally, but I know they they would They love they would love to to take this off your hands.

Speaker 1: Rushing to an appointment this morning, I grabbed it to read it, first mistake, then went to breakfast and had nothing else. So I spent the morning reading the worst book I had found. It's such a funny such a funny like, like I don't know. If it's feels like a Curb Your Enthusiasm episode or something like that. Anyway, thank you so much for tuning into TBPN today. We'll see you tomorrow at 11AM Pacific. Leave a spot for us on Apple Podcast and Spotify. It's been an honor. Sign up for our newsletter. Tbpn.com. Goodbye.