NVIDIA's Ali Golshan on OpenShell: a formal-verification security layer for the AI agent era
Sep 30, 2026 · Full transcript · This transcript is auto-generated and may contain errors.
Featuring Ali Golshan
Speaker 2: Our first guest is from NVIDIA, the senior director of AI Software. Ali, welcome to the show. How are you doing?
Speaker 5: Doing well. Thanks for having me.
Speaker 2: Thank you.
Speaker 1: Great to have
Speaker 2: you. So much for hopping on the show. Since this is the first time on the show, could you tell us a little bit about your role at NVIDIA, what your day to day is like, what projects are you the most focused on these days?
Speaker 5: Sure. So I'm on the product side. Most of my time is spent around security, safety, and infrastructure of AI. Me and my team created this project named OpenShell, which was part of our announcement. We joined about eighteen months ago. We were founders of a company named Gretel. We focused on data privacy, synthetic data, things like that. Founded a few startups and all of us started our careers at security researchers and computer scientists and data intelligence community. So
Speaker 2: Yeah.
Speaker 5: Sort of full circle coming back all the way back to our original work here.
Speaker 2: I like the name Shell. It's better than Sandbox. Sandbox Mhmm. Easy to get out of. Any child can walk out of a Sandbox. The sand goes everywhere when you have a Sandbox. A Shell contains the turtle. The turtle cannot just easily escape the shell. Right?
Speaker 5: Yeah. So I like More importantly, it goes with it wherever it goes
Speaker 1: Yes. To keep
Speaker 5: it safe.
Speaker 2: Yes. And there's little holes that the turtle can poke the arms out of.
Speaker 1: A turtle can be in its shell can still bite you.
Speaker 2: Bite you. Okay. But it's limited. It can't it can't bite you from any angle because there's only Yeah. That's right. One hole that the head We'll comes
Speaker 1: keep working this analogy out while you actually
Speaker 5: Once you perfect it, love to get it back. We'll make it into stickers.
Speaker 2: So OpenShell, is it open source? Are you are you evangelizing this product to labs, neo labs, everyone, businesses? Like, how do you see this product actually rolling out and being used by the AI community, the technology community, the the the real economy, etcetera?
Speaker 5: Yeah. We've been working on it for about a year. We actually released it as an early access in March. It's been Apache two point o since we released it. We're actually on pace to donate it to a foundation. Oh, cool. So we we don't even wanna be sort of the NVIDIA project. We want it to be a true community pro project. Yeah. We intentionally built it open so it becomes a standard. It becomes sort of a trusted layer.
Speaker 4: Mhmm.
Speaker 5: Like we think that what agents are missing is sort of like if you think about like the web from nineties to 2000, there's that trust layer like lights up, SSL is on, you know, your tab is isolated. Like that construct is missing from agents. So our perspective was if you can create a trusted layer, open it so it can be widely adopted and then built on top of, then you can sort of have a trusted layer everybody can work through. So that was sort of the the whole intent and opening that is really about rising tides.
Speaker 2: Yeah. Can you get me up to speed on the state of formal verification? How much optimism should I have towards this as a solution? We had Greg Brockman on this show a few weeks ago, and he was saying that he he he has this point that there is an optimistic scenario where enough software is formally verified, fully secure that certain systems are hardened to a point where it ceases becoming a cat and mouse game. Is that too optimistic? Are we on the path? Are we already seeing glimpses of that? Or will or is the game going to just be played, you know, indefinitely and we just need more good guys than bad guys?
Speaker 5: Yeah. The state of formal methods and what it can do, I think the first substantiations of it, we already have proof of life. Like, if we can demonstrate what a full stack verification looks like. Mhmm. So I think I'm very optimistic because a lot of our thinking that has led to this is is that agent safety trust really needs to be a full stack solution. Like if you think about where models, frontier labs, harness data tools, that, that's really the app player. We build OpenShell because we felt like you needed to be able to move the risk line to a deterministic layer that nothing can bypass. So that's the runtime. You can do formal verification on that and get really good controls. If you then move all that further into hardware, now you can verify the entire stack. So you could be assured down to a single binary decision is anything leaving this system or not leaving this system. So I'm incredibly optimistic about formal verification, but it has to be applied the right places in the stack for it to have the affected needs.
Speaker 2: Mhmm. Talk about your how you interface with NVIDIA's security features on chip. There's been some discussion about Apple advancing this with the secure enclave and partnering with NVIDIA in various ways. And it feels like there's a number of different organizations. They're all swimming in the same direction, maybe talking about things differently, solving problems differently. But how do you actually interface with the the hardware teams or or or what do you what are you seeing from the hardware teams that gives you opt optimism around security?
Speaker 5: Yeah. So let me sort of answer it in two parts. One, what we're doing on the OpenShell side so we can interact with all the hardware, and then what we are seeing on the hardware side. I think those are and they're agnostic because they both should be equally as good without each other, but a much better story together.
Speaker 2: Sure.
Speaker 5: So on the OpenShell side, we've just taken a very agnostic approach. We have this primitive inside the runtime called drivers. It's exactly what it sounds like. You pick your compute under the hood. And the reason it's important to pick the hardware, the compute under the hood is sort of what we announced with our open agent safety platform is we put a reference design out with our BlueField Vera to say if you have hardware, OpenShell can take all these controls at the network for chain of thought and move it all the way into hardware. So now you can do it at machine speed and you can have very deterministic enforcement on it. It was sort of a reference architecture for us. So from the OpenShell side, we've tried to be as agnostic and inclusive as possible by giving a driver that any computer hardware can plug into, which is you probably may have seen this in our announcement why we briefed ARM and Intel and sort of go down the list of hardware providers. Right? Because we want everyone to be able to offer this full stack trust and safety verification. As far as what we are seeing on the hardware side, there are two things we are starting to see sort of generalized patterns. One is the hardware component making primitives available so the higher levels of the stack can offload networking into it. Mhmm. Think of it as a way to, like, manually flip the switch and cut any access. Access.
Speaker 6: Okay.
Speaker 5: So if you're doing that in a hardware, if you're an advanced, for example, researcher or a frontier lab or a large organization doing research or sort of testing of a frontier model, by having that at control at that stage, you know and can verify the boundary or the blast radius you're setting is not going to be breached. Right? It's in the hardware. It can't be bypassed. Agent can't work together to get around it. So that's one area of sort of work that we are seeing, and obviously, this is the work we're doing at NVIDIA too. The other side of it is being able to pass through the way the agent thinks and reasons through silicon. Mhmm. So you can actually see what it's thinking about. The reason this is important is is, like, if you're familiar with security, the holy grail is shifting left. Right? Not responding, but preventing. Like, how do you do that? So if you pass that chain of reason, that thought the agent has through silicon, you can see it's thinking about using a zero day. Then you can check with the runtime. Did it probe anything to see if anything's open? You can come back and you're like, oh, it did probe. Now it's reasoning to say, I actually found a hole. Now let me go build an exploit. Then you can pass that and cut network and say this thing is about to break out of the environment because it's reasoning about using like a zero day or a vulnerability. Mhmm. Or, like most agents, it could just think about these things and basically let it go as a chain of thought and not come back to it. But those are the two really interesting areas of research we're seeing in hardware right now.
Speaker 2: Please.
Speaker 1: Sorry if this is a silly question, but would an agent actually think, should I use a zero day today? Or would it just identify a potential exploit that to a even if you're observing it wouldn't necessarily be immediately obvious that it was an exploit?
Speaker 5: Yeah. So, you know, I'm sort of being overly, you know, rudimentary about saying
Speaker 1: it's yeah. I know you're trying to dumb this down to us.
Speaker 2: Bring you're trying to bring this down, like let
Speaker 1: me bring it down, like, 20 levels. No. We have we
Speaker 5: No. No. Fifteen, sixteen max. Not 20.
Speaker 1: Preschool. Preschool. Please keep it. Don't keep it need preschool, not kindergarten. Thank you.
Speaker 5: For sure. So let let me say it differently. Like, if you look at the types of events that happen with these labs Mhmm. Where the agents did something that was unintended. Yeah. It wasn't a single step function in capability that achieved that. Mhmm. It was doing a subset of things like stealing credentials, communicating with each other in channels that you didn't have, persistent sessions longer. Doing all these things that at the quantum level, we know how to solve these things from a security standpoint. Right? Like we know how to keep credentials away from humans. We can do the same things for agents. We know how to introspect traffic. So it was a whole bunch of techniques uniquely put together and then ran at machine speed. The machine speed part is solved in silicon when you can see it thinking. And then the the path to exploitability and vulnerability exploitation is a set of things and actions that are well understood. Right? You probe the network into an environment that the policy says you're not allowed to. You take that reading and you build a tool against it. So this iterative process is something we can build rules around and then build fail saves as a particular part of the stage.
Speaker 2: Yep. That makes a lot of sense. Sense.
Speaker 1: Yeah. Talk about, when, like, throughout throughout your career in security, there were sort of spikes of panic. Like, I would imagine I would imagine we're we're generally at 10 out of 10 relative to maybe other moments, but I'm curious if there are other moments that you went through on the path to this current moment that were that were any anywhere anywhere close, but maybe they weren't as sort of public and talked about on a national level or in the media as much, but maybe were more kind of internal industry.
Speaker 2: Like SolarWinds?
Speaker 5: Yeah. I mean, so it's funny, like, lot of my panic was either associated into the intelligence community that I can't talk about or as a founder which had to do with funding or board members who
Speaker 1: had less than one conditions. Of
Speaker 5: But I would say that I I wouldn't call call it panic. What I would say is is like ambiguity increases risk. Right? Security is about risk reduction. There's never a perfect system. So I would say the times that created the most level of ambiguity where then sort of opinion fills the void were when we had these massive step functions and form factor. Like, not SolarWinds, but like moving sensitive compute to mobile or to cloud Mhmm. Or making them fully distributed. Mhmm. And now sort of it feels like another level of control given away. Right? Like, you think about cloud, everybody understood hands around things. Right? And now you're like, where is my stuff? Like, that was a very fundamentally, conceptually difficult shift to make for a lot of years. Mhmm. And I think agents AI have a very similar fact factor. Like, we joked about sandbox, but, like, to use one concrete thing. Right? Like, sandbox is meant to constrain something, keep it in or keep it out. These agents by definition have to do both to be productive. Mhmm. So when we talk about, oh, sandboxing is not effective against these agents, we're all panicking. Well, because we're of using a rudimentary version of it. So the same way we didn't take, like, these monolithic applications and shove them in containers and then run them on Kubernetes, the way we redesigned with microservices and immutable ephemeral infrastructure, The same thing needs to be done here. Most of the problems we're seeing is we are running an advanced technology into an environment that is not meant for it or designed for it. So what does a native environment look like for AI basically is where I spend a lot of my time.
Speaker 2: Well, thank you so much for coming on sir. A great conversation.