Jeff Huber launches Zeitgeist, Chroma's shared AI memory layer for enterprise teams, as cloud hits 1,000 paying customers
Aug 12, 2026 · Full transcript · This transcript is auto-generated and may contain errors.
Featuring Jeff Huber
Speaker 2: We'll talk to you soon. Have a good day.
Speaker 3: Thanks, folks.
Speaker 2: Cheers. Let me tell you about CrowdStrike. Your business is AI. Their business is securing it. CrowdStrike secures AI and stops breaches. Our next guest is Jeff Huber returning to the show. I think it's been over a year. He's the founder and CEO of Chroma. And we got some good news from Jeff Huber. Let's bring him in. I know we were keeping you waiting, Jeff. Sorry for that. Late today. How you doing? We're good.
Speaker 5: Doing good, guys. Good to see you again. Great.
Speaker 2: Good to
Speaker 1: see you too long.
Speaker 2: Give us the latest in your world. Take us through recent announcements, and then we can go all over the place, I'm sure.
Speaker 5: Let's do it. I think it's been a really interesting year in AI, and in particular, you know, seeing how companies are more deeply adopting AI agents. I think, like, as we sort of reflect on, like, the biggest gap in AI, it still feels like memory is the biggest gap.
Speaker 2: I mean,
Speaker 5: I'm sure you all try to use agents all the time, and they're constantly forgetting how to do things, and they're not the right context. And so Yeah. Yeah. We're we're announcing on the show kind of our solution to that.
Speaker 2: Yes. Take us through it. What is zeitgeist?
Speaker 5: Yeah. I think, like, in short form, there's probably a lot of sophisticated approaches to solving memory for AI, but in very simple form, if you can get AI agents just to be able to write things down and then later find those things. Yeah. You know, that's all you need. And in some ways, I think Wikipedia solved it twenty years ago. You can call it a revolutionary data model. It's just a wiki. And so I I think really inspired by like how Open Claw and Hermit Agent have been managing markdown files.
Speaker 2: Yeah. I was about to ask.
Speaker 5: Kind of thinking through
Speaker 2: Why not just a bunch of markdown files? Why not just take the Open Claw vibe coder Mac mini guy and scale it up to the organization of a thousand people in the enterprise?
Speaker 5: It's a great place to start. Okay. You know, as our team was adopting this technology over the last, like, three or four months, you know, we wanted, like, the shared memory layer for our whole team that knows everything about our whole company, really. And you've got to think about concurrency control and access control and versioning and lineage and just a bunch of infrastructure things that matter Yeah. As the dataset scales.
Speaker 2: So so as a concrete example, like, might be someone in human resources that's developed a skill to give people raises, but you don't necessarily want that skill proliferating across the entire organization. So I mean, otherwise, everyone's just giving themselves raises left and right. Is that is that a reasonable example?
Speaker 5: I think so. Yeah. Hopefully, the other access control systems would not allow anybody to give anybody a raise. Right? Sure. But, maybe more grounded, like, what we've seen amongst our own team is that, like, now our engineering team, the first place they go when they have a technical question is they ask Foundation, the product's called, they ask Foundation that question. And Foundation has seen all of our Notion docs, all of our Google Drive docs, all of our coding agent sessions, which actually have so much, like, high value tacit information in them Yeah. About how our team works, how they operate, what we care about.
Speaker 1: Mhmm.
Speaker 5: But now it spreads. The customer support's constantly asking questions. Sales, GTMs asking questions. It really become this, like, repository of everything we know.
Speaker 2: So Yeah. What take us through the the shape of the core business, the the launch of Chroma Cloud, how that business is going, just overall demand for for core products in and around the AI boom because everything is sort of booming and every system's, like, straining. What's your experience been?
Speaker 5: Yes. We launched current cloud in late twenty twenty five. We now have close to a thousand paying customers on the product and tens of thousands of other users who, you know, we'll bring on board soon. You know, I think, like, these lower level infrastructure primitives are cool. Right? A core core database, core search, core operations. They allow developers to pick it up and build useful applications. And we were looking at the market and seeing really what our customers were building. We realized there was an opportunity to meet them where they were. They were all asking us, hey. How do we build do we build Wikis? How do we solve access control? How do we solve versioning just time and time again? And so, you know, us just building those infrastructure for Biz for customers to use allows them to, you know, focus on what actually differentiates their business. Mhmm. So.
Speaker 2: Give me your not your AGI timelines, but your timeline to just good search. I feel like I go into my Gmail and I search for something and it's still bad. And I want an LLM to be involved, but I don't want it to be as slow as a current LLM. I know I can wire it up to codex or whatever and and have it read through everything. But I don't want to wait ten minutes, but I also don't want the current experience. And I feel like there will be a merging of these two, but do we need to bake it into an ASIC or something? Like, how complex is it going be until, like, you just show up to a website and the search function is magical and also as fast as you remember it?
Speaker 5: Yeah. I think there's a few different pieces of the puzzle. So number one is, like, having highly expressive, accurate, and fast search throughout the database layer. Right? And the second piece of that is, like, the agent itself for the model being really fast. You know, one thing that we made a bet on earlier this year was agentic search becoming, like, the default. That really all searches will go through a model. This is what it the bitter I think the bitter lesson means for search is that, you know, you'll spend more tokens, you'll use more inference. That's how you get better at search. And so we released a model in this spring called Context One. It's a fine tune of GPT OSS 20 d, and that model on three bursts today already runs at 3,000 tokens per second. And we were targeting, I think, yeah, Talos got Yeah. Bought last week. I think you guys covered that.
Speaker 1: It's Chad Jimmy. You know
Speaker 2: chadjimmy.ai. Best demo
Speaker 1: that names don't matter when the when the product
Speaker 2: When the product's really technical and good.
Speaker 1: The team has got cracked. You can name your product whatever you want. Chat jimmy.ai.
Speaker 2: We were
Speaker 5: a shared investor. You know, we were targeting context about running at like 15,000 or 20,000 per second on that chip. And I think like if you think about that's gonna come. Right? That's coming this year, and that's gonna change how we really rewire how people think about language models.
Speaker 2: Yeah. Is there still, like even though the the the the token speed is very quick in that example, is there a delay to sort of, like, fire up the chip, like, load load data into memory, build the context window, like do whatever you need to to actually get the once the inference is cooking, I can imagine it being very fast. But is there if you're dealing with something that's like a scaled email product and you have millions of users that that are gonna be searching things randomly, like, you can't store all of that in Yeah. Like, in memory all the time. So, like like, what Yeah. What what like, what is the process there? Is that a meaningful barrier or not?
Speaker 5: I think fast inference sort of what you're implying here, like fast inference does not solve just sort of for you, conduct engineering.
Speaker 6: Yeah.
Speaker 5: Like, you still have to make sure that you're bringing the model the right information, the right granularity of information, letting it forget information such that it can refocus. And I think it's been cool to see like Connus engineering to some degree get folded back into the model, but also as we've seen as of late, like, you know, really good harnesses, like the stuff from PrimeMintilex, like really good harnesses that are targeting tasks, you know, that is in some ways convex engineering and, like, they can outperform frontier models by a huge margin. And so I think there's still huge gains to to convex engineering.
Speaker 1: Mhmm. Jordy, anything else? I got more questions but we're we're running behind. Yeah. Sorry. We Great to great to see you, Jeff. Let's not let it be so long
Speaker 2: Yeah. Let's do it again soon.
Speaker 1: Congrats on all the all the progress.
Speaker 2: Yeah. Thank you so much for coming on the show. We'll talk