Positron AI raised $875M total to build memory-focused inference chips and is targeting hyperscalers with plans to tape out this year

Sep 11, 2026 · Full transcript · This transcript is auto-generated and may contain errors.

Featuring Mitesh Agrawal

Speaker 1: And let me also tell you about CrowdStrike. Your business is AI. Their business is securing it. CrowdStrike secures AI and stops breaches. We are joined by Mitesh Agrawal from Positron AI building a new chip for the AI era. Welcome

Speaker 6: to What's going

Speaker 1: on? How are you doing?

Speaker 4: Hey, John Jordy. Pretty good. How are you guys?

Speaker 1: Thanks so much for hopping on the Pretty good. Great to have you here.

Speaker 4: I I just want to start by saying I've been in the background of Steven doing this multiple times now.

Speaker 1: Oh, yeah.

Speaker 4: And Steven Steven Valivan from

Speaker 1: That's right.

Speaker 4: You guys. So this is my first time on, so thanks thanks for having me.

Speaker 1: How how direct is the lineage from from Lambda? You're working at effectively Neo Cloud. You see the problem. You go solve the problem externally with a new startup. Is is the story that simple?

Speaker 4: Yeah. Fairly. I mean, for me, I mean, look, I didn't found Positron. Right? Positron was co founded by Thomas Summers and Edward Kemet.

Speaker 1: Sure.

Speaker 4: Another lineage, Grok lineage from from from before. And, you know, they they designed the actual silicon and and and the system setup. And part of it is just, like, great luck. You know, I've known both Steven and Thomas for over decade. I've been close friends with both of them. I've worked with them, been roommates, everything, all of all of those things. And Thomas has been wanting me to join Positron since day one, since he started the company in 2023. Yeah. But Lambda was just starting on its hockey stick growth then, I was like, look. I'm not leaving Lambda. Started the company with Steven there in Lambda Cloud. But early twenty twenty five, like late twenty four, you know, reasoning models had come out. O one just started to to to get into zeitgeist, and then video generation. You know, Sora, first Sora came out. Not a lot of people saw it, but I I got to see a little bit behind the scenes on on the video generation models, the amount of memory. I remember looking at the Google video model back then, you know, it needed four h one hundreds to run, like, a ten second clip.

Speaker 1: Yeah.

Speaker 4: It's completely memory bound on bandwidth and capacity. Mhmm. And knew what Thomas was building. I was like, this is actually very interesting. Memory is gonna get a big part of the story for inference. Somebody's building something about it. Let me go get and and and work at the at actually the fundamental technology there. Like, Lambda builds technology on the cloud and and services, you know, and I'm chemical engineering, studied fabrication, never used it ever. So I was like, alright, I'm gonna go back a little bit to my roots and then come back to it.

Speaker 2: Amazing.

Speaker 1: How much is AI actually accelerating semiconductor design, semiconductor fabrication? We saw one of your investors, Dylan Patel, at semi analysis talking about the the open air jalapeno chips seemed like ahead of schedule or very, very quick. For a long time, we've been hearing, oh, new chip. That's three years. That's five years. Feels like it's eighteen months now. What are you actually feeling? What are you seeing?

Speaker 4: Yeah. To start from reverse, like, to answer your last part about it, like Yeah. Man, new chip every twelve months. NVIDIA is the absolute king. And if they're coming out with a new silicon every twelve months, you better get in that game or or, you know, like, you like, don't even be part of the conversation kind of thing. Right? So so that's that's for sure. In terms of utilization of AI, mean, look, I'll just, like, focus on Positron itself, but, you

Speaker 1: know Yeah.

Speaker 4: We are a small team. I mean, we we got to we we are right just now. I mean, just yesterday or today, crossed a 100 people, but over 50 of that is over the last three months. So we got our first gen product out with Mhmm. Less than 20 people, and that's built on FPGA, so it's already pre taped out silicon, and we we are deploying and implementing our architecture. And our second gen

Speaker 1: Those are the out end of this year. Are the 50 Atlas racks you have at Oracle?

Speaker 4: Oracle? Yeah. Okay.

Speaker 1: So those are FPGAs. Interesting.

Speaker 4: Yeah. Those those are FPGAs. Cool. And it's kind of like harking back to a little bit of of previous times when you used to design and and build silicon. You would actually test on FPGA before going into the

Speaker 1: Yeah.

Speaker 4: The tape out kind of thing. So we just wanted to get a product out as quickly as possible. Like, that was the whole thing. It's like, you know, from the start of the company, we got our first shipment to a customer in fifteen months, and it was, you know, it was built on FPGAs, obviously. But to get the full, you know, big file ready, getting it deployed, getting models running on it, it was all done in the first fifteen months, and then over the last ten, three months killed it out. But, yeah, I mean, like, look, we have to use a lot of the AI toolkits, especially on verification. Design, less so, I would say. I mean, look, obviously, we use a lot to kind of interact with, like, now Astra, for example, is phenomenal, right, you know, to interact with it, but you're still not going there and saying, hey. Like, come up with this, like, new design yet. Although, like, you know, Ana's computer cursive and others, are they're obviously built out, and, you know, they raised a big round for for that as well. Right? So it's gonna come. You know, you're gonna see very soon. Mhmm. It's like one person an Astra or one person an Astra's Sure. Taped out a chip kind of thing. But we have to use it a lot. I mean, if you think about a 100 people or, you know, very recent until very recently, 50 people. Mhmm. For a company that is targeting tape out end of this year to do with, like, fifty, sixty people, it's it's very tiny amount in a in a Silicon world.

Speaker 1: What does the demand side of the equation work? You're already working with jump trading, i3d.net. Is this something where you like, if you can get capacity, if you can get performance, some solid benchmarks, you think that sales isn't going Or to be a are you going to have to find and work very closely with a customer to sort of co design a solution for a particular problem within the AI stack.

Speaker 4: Yeah. I I don't want to trivialize or make it sound simple like like like that.

Speaker 1: Your sales guys might be listening and they're like, we work very hard. Okay? Shut up.

Speaker 4: But, well, we we we only have one sales person. We only have one sales person, you know, like in that way. Right? But the point I will I will make is really around the way we think about the demand curve is you're kind of hitting the nail on the head in saying that, like, look, if you can make your silicon work, show the performance is comparable, especially in the current ecosystem, even within the niche of doing this ag or something like that, but especially if you can make the entire inference kind of workflow have a good TCO and or, and generally people always assume it's an or that, hey, you have a good TCO or you have a great interactivity curve. Mhmm. But if you can do and or, you know, you're you're gonna bound to have get demand. More importantly, you know, when we when you step into the rooms of, like, not only just jump trading or, you know, hedge funds or or kind of influencer service providers, but, like, really the big labs, the hyperscalers, Kind of two questions that it boils down to is, like, hey. Like, look, guys. Can you fabricate this in enough quantities? Like, you know, is your supply chain and the way that you are using the technology components, is it robust enough that you can fabricate it that we can be interested in it? And then the second thing they're asking is, can we deploy it? You know, is is your power source, like, you know, do you need this kind of liquid pool setup? And and if so, then, you know, we don't we might not have a data center because we've already allocated to GPUs or TPUs. Mhmm. Or can you do something else? So so the questions you can see there they are asking is not that, hey. Like, you know, you know, we we we will see. We don't we we're not showing up at the demand curve. So so from that angle, you're kind of spot on that, look, if you can make the frontier models run, you know, you're you're bound to find kind of adoption in in today's market, And that is a really great like, I mean, I'm so like, you know, as positive from we are so lucky to be building silicon in this environment. And that's kind of what you're seeing for with silicon companies is raising rounds right now.

Speaker 1: Yeah. I think $9,000,000,000 has flowed into silicon companies over just the last twelve months and Yeah. Which we were we were talking yesterday

Speaker 2: feels feels incredibly low relative to the spend like Yeah. Spending category.

Speaker 1: Gavin Baker, one of another one of your investors has this quote. Where he says, I see it as a 1% market share is a $100,000,000,000 opportunity. And that sounds like a crazy bold take and you then you realize like, wait, no. NVIDIA is a $5,000,000,000,000 company. Like, it's gonna be a $10,000,000,000,000 market like any day now. And so, yeah. Actually, 1% should equal a 100,000,000,000. But yeah, anything else? Sorry.

Speaker 4: No. You you like he said that in a board meeting to me like, I don't even know, like a year ago or something. Yeah. There is is basically like, yeah, he's he's like like, look look look guys, like Mitesh Thomas, just just 1% of the market, you know, a 100,000,000,000 enterprise value. Just focus on your architecture where you can do well. You know, like, one of the things that, you know, people always like, whenever a new new chip company raises around the headline, and luckily you guys don't have that, which is like, to rival NVIDIA. It's like, guys. Like, no no one is rivaling NVIDIA. Like, get to at least 10% of their revenue before before putting the tagline on. Right? But but but, like, the point there is just like, look, NVIDIA is everywhere. You gotta work in that to both work with them, but also like having a product that is differentiated enough. Like, you have to have technical innovation, obviously, to stand out, and then show your performance TCOs and interactivity. But then you also got to prove that like, look, in the world of HP and CoWoS constraint, like for us, our big stories are know, like, look, HP and then CoWoS bottleneck. You have NVIDIA, TPUs, AMD's ahead of you in the blind. You know, how do you get around that? Well Mhmm. Again, you know, you you say, okay. We are using commodity memory. Well, pro and con. No no free lunch in Silicon Land. Like, you know, commodity memory is slow. How do you solve that? That's where the technical innovation comes in. And then second thing is, like, okay. It's still not trivial to get commodity memory. It's not like I can just show up to Samsung and be like or Micron, be like, hey. Can you give me LPDDR5x? You know? Yeah. You have to still figure out how to how to get that and plan it. But it is more feasible to to get it, and that becomes a story of that the company can then scale out and saying, not only we're gonna have a product, but we're gonna have a product that will scale with the requirements of hyperscalers and and and kind of the frontier Yeah.

Speaker 2: How do these how do these customers think about, like, the minimum scale when when they're when they're working with you and they're looking at ordering order making orders that will be delivered in, let's say, 2028, 2029. Right? You need to be able to to really be worth a company at that scale's time, you need to be thinking about like it's almost like, hey, the orders we want delivered in 2029 are like a proof of concept for like the 2,032 order which will be, you know Yeah. At some scale to actually impact and and be able to scale the fleet in a meaningful way. But but how are you thinking about like that feels like the biggest challenge is like minimum viable sort of like 100 deployment.

Speaker 4: 100%. I mean, like, literally you kind of circle back on, like, as I said, when we walk in these meetings and like, scale depends on, like, if you're going into hyperscalers and frontier labs, and I said the first question they ask is like, guys, can you fabricate like this? Like, in in enough? And the the question there is, like, in enough quantities that it's, like, worthwhile to us. And and that answer for hyperscalers and and Frontier Labs, like, honestly, they will literally say gigawatt plus. Like, come up to us with a proposal of a gigawatt plus, which is kind of insane. Right? Like, in a gigawatt, like, even at NVIDIA scale, you're talking about $75,000,000,000 worth of of of revenue for them. Right? Even, you know, assume ASIC cheaper, blah blah blah, all those things, you're still talking about tens of billions of dollars. Right? But, like, at least you have to show a plan of, like, how do we get to hundreds of megawatts in you know, to to use your, like, year specific year 2028. You know, for us, we're taping out this year production kinda ramp up in second half of twenty twenty seven. In 2028, we better have a plan of how do we get to, like, hundreds. And I I don't want to just say cop out by saying hundreds as in just a 100 megawatt. Hundreds means truly, like, you know, three, four, five, and above for for those. But

Speaker 2: Yeah. So that's by the next scale up, you're in that gigawatt range.

Speaker 4: Yeah. Exactly. And but also, like, I also don't wanna discount the fact that you have other customers, like, you know, have Inference as service providers, obviously, Sovereign AI Clouds, you know, quantitative finance, quant finance kind of spectrum. And so they have, like, different magnitudes of kind of requirements that that come through with it. So, you know, we although we do internally use kind of go big or go home as as a thing, like, we have to attract one of these large customers to to really be a long term viable company. You know, I I don't wanna just, like, discount the fact that, like, look, you can grow the company through the ranks as well. Like, you can grow the company, you know, get 200,000,000 revenue for $250,000,000, a billion, 2,000,000,000 through through this other kind of channels as well. Right? I think that that that becomes a big part of it. But, yeah, like, if you really wanna get to, like, the frontier labs and hyperscalers, you're really talking about hundreds of megawatts. And that's why, like, you know, like, look, when we, you know, we have Venture Tech Alliance on our kind of cap table, and and when we speak with TSMC for fab capacity, they are also wanting to know kind of like, can you scale, like, know, do you have the balance sheet to do that? Like, one of the reasons we raised, you know, $875,000,000 is not like we need $875,000,000 to spend tomorrow or even in the next six months. I mean, look, we raised $230,000,000 in series b in February of this year. Untouched. Right? We still have all that capital. Part of it is because we have been making revenue this year, but we do have plan to spend that very quickly. Thank you, Joey. That that

Speaker 1: was actually That was one year when you were here.

Speaker 4: Untouched.

Speaker 2: Untouched. Untouched. Untouched. Untouched. Because we're making revenue. Yes.

Speaker 4: The the the main point that is there though is like, look, they they want to know, like, if if you actually get a customer, you have the capital. Yeah. And and even that capital is is that is not enough equity capital to scale out to even, you know, 200 megawatt. Right? And then you have to go to the black stones of the world and figure out how do you how do finance that deal, kind of what Lambda has done. Right? So

Speaker 1: What's the software side of the equation? You're coming for NVIDIA. You're challenging them. You're gonna drive their market cap to zero. Have to drive that end. No. Obviously, this is a market that can sustain multiple players, and there's different tools for the job. But interoperability is important. I'm interested in terms of software development. Are you going to lean more open source with the software side of the business or more integration with just a few buyers and co design on the software side to make sure the integration is really seamless? Is is there even do we even need to be having a software conversation in an era where AI agents can write code?

Speaker 4: Yeah. I mean, you you definitely need the software conversation because, like, you you you have to plan around how people wanna use it, and people wanna use it how they're currently using it and gonna continue to use it, which is based on NVIDIA stack, but also primarily based on PyTorch and then, know, VLLMSG lang as the and and I'm I'm specifically focused on on inference items. Like, look, training, you know, that's such a harder challenge, like, you know, what Jensen says, like, true mode around scale out and everything. Right? Like, that's just that's the only reason you have probably only TPU as a potential, kind of only other silicon that can be used for training. Right?

Speaker 1: Sure. Sure. Sure.

Speaker 4: But but on the inference side of things, for sure, you have to have the conversation. I will say this, in the era of agentic kind of software development, the worries around like, hey, you know, model drops, if you don't have access to it, it takes you days, weeks, months to bring it up. It's it's going away. Like, you know, know, we we had news glimmer drop and within our team, you know, on our Atlas, first gen could get it up and running within hours. Right? And You and that that that yeah. Exactly. Like, I have the same reaction, by the way, when when we had that, and and people are are are making it even faster and more automated too. Like, you don't even have to interact. Model drops, comes in, can can can probably do it in in and that's a very near future of it. But to that point, it doesn't give you the right away the efficiency, the optimization. Sure. Like, know, you you wanna extract every dollar off of it. So to your question around, you know, when when the customer is large enough, you wanna work closely with them to, like, literally extract every single dollar. And also, like, I'll I'll be very frank, like Anthropic, OpenAI, this Frontier Labs, hyperscalers, they are so sophisticated. They kinda wanna come in and be like, look, guys, even if you don't want it, we are working with you to make sure that this is gonna, like, you know, this is optimized to fullest, right? So the answer, as always, in in this scenario is all of the above, you know, even though it might sound like it's like, oh, it's a very cliched answer, but it really is that way.

Speaker 1: Sounds like the mafia coming in, oh, your software stack isn't open source, so you're about to open it for me. I'm gonna make some changes if I The need

Speaker 4: software stack is, like, gonna be built around open source, like, the sense of, like, if if you wanna make, like, every company to to use us for inference Yeah. You know,

Speaker 1: you have

Speaker 4: to build it on SGLAN and and VLM kind of setup. Right? That's okay. But, you know, it's like when you're talking to SpaceX or Anthropic or OpenAI, they're not using the generic SGLAN or VLM, they have all their optimizations built in, they're going to help you do that. Then obviously there's Disag, then within Disag there's all the different domains that they do, and they're going to figure out, it's like, oh, jalapeno is good for this, All the tronia, you're good for this. You And then they're gonna they're gonna say, okay. That's how we're gonna use you guys.

Speaker 1: Amazing. Well, exciting times. I wanna hit the phone for you. You raised 875,000,000. Oh, man.

Speaker 4: That was a solid run.

Speaker 1: Thank you. Job, James. Thank you, John. And thank you

Speaker 2: so Great stuff. Great to meet you.

Speaker 1: Keep an eye out on Instagram. Where you're building. Definitely dropping a NVIDIA challenger slide later today.

Speaker 4: Yeah. Sure. Please do not associate my photo with that.

Speaker 3: Better yes. Well,

Speaker 1: have a great rest of your day.

Speaker 2: Looking forward to the next appearance. Have a good one. Great to hang.

Speaker 1: Goodbye. Cheers. Let me tell you about Railway. Railway is the all in one intelligent cloud provider. Use your favorite agent to deploy web app servers, databases, and more while Railway automatically takes care of scaling, monitoring, and security. Lastly, the New York Stock Exchange. Wanna change the world? Raise capital at the New York Stock Exchange. That's can

Speaker 2: see positron over there pretty soon. Yeah. Pretty soon.

Speaker 1: Picking up some

Speaker 3: fresh ones.

Speaker 2: Wonderful week.

Speaker 3: Wonderful week.

Speaker 2: Short week.

Speaker 1: Monday. Yeah. Monday, 11AM.

Speaker 2: Do us a favor and go ahead and have the best weekend

Speaker 1: Best weekend. Of your

Speaker 2: entire life.

Speaker 1: Have the best weekend of your entire Let's

Speaker 2: do it. Put the pieces together. Make it happen.

Speaker 1: We'll see you on Monday. Leave us five stars in Apple Podcast and Spotify. Sign up for a newsletter at tbpn.goodbye.