Interview

Jeff Huber launches Zeitgeist, Chroma's shared AI memory layer for enterprise teams, as cloud hits 1,000 paying customers

Aug 12, 2026 with Jeff Huber

Key Points

  • Chroma launches Foundation, a shared memory layer that lets AI agents write and retrieve information across enterprise organizations with versioning and access control.
  • Chroma Cloud reaches nearly 1,000 paying customers since launching in late 2025, with Foundation built directly from customer demand for institutional memory infrastructure.
  • CEO Jeff Huber argues all search will eventually run through models, betting that faster inference on specialized chips will reach 15,000 to 20,000 tokens per second this year.
Jeff Huber launches Zeitgeist, Chroma's shared AI memory layer for enterprise teams, as cloud hits 1,000 paying customers

Chroma launches Foundation, bets on agentic search

Jeff Huber, CEO of Chroma, sees memory as the biggest remaining gap in enterprise AI — agents that keep forgetting context, losing institutional knowledge between sessions, and failing to share what they learn across teams.

Foundation, the product Chroma is announcing, is a shared memory layer built for enterprise teams. The concept is deliberately unglamorous: a versioned, access-controlled wiki that AI agents can write to and read from across an organization. Huber says the team was inspired by how tools like Claude and similar coding agents manage markdown files, then asked what it would take to scale that model to an organization of a thousand people. The answer required concurrency control, access control, versioning, and lineage tracking — infrastructure that markdown files alone can't provide.

Chroma's own team is the first real case study. Engineers route technical questions through Foundation before going anywhere else. The system has ingested Notion docs, Google Drive documents, and coding agent session logs — which Huber describes as particularly high-value because they capture tacit knowledge about how the team actually works. Customer support and GTM teams have since followed, and Foundation has effectively become the company's institutional memory.

Chroma Cloud, launched in late 2025, now has close to 1,000 paying customers, with tens of thousands of additional users Huber expects to convert. Foundation grew directly from what those customers were asking for — versioning, access control, wiki-style organization — which Chroma decided to build into the platform rather than leave each customer to solve independently.

Memory is still the biggest gap in AI. In short form, if you can get AI agents just to be able to write things down and then later find those things — that's all you need. Foundation has seen all of our Notion docs, all of our Google Drive docs, all of our coding agent sessions, and now our whole team goes there first with any technical question.

Agentic search

On the broader search question, Huber argues that all search will eventually run through a model. Spending more tokens and more inference is how search gets meaningfully better — the bitter lesson applied to retrieval. Chroma made that bet earlier this year and released Context One, a fine-tune of an open-source model that already runs at 3,000 tokens per second. Huber was targeting 15,000 to 20,000 tokens per second on specialized inference chips, a milestone he believes arrives this year.

Fast inference alone doesn't close the gap, though. Getting the right information to the model at the right granularity — and letting it selectively forget to refocus — remains a distinct engineering problem. Huber points to purpose-built harnesses from companies like PrimeMind that target specific tasks and can outperform frontier models by a wide margin, suggesting context engineering still has substantial room to run even as model speed improves.

Every deal, every interview. 5 minutes.

TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.