Databricks co-founder Patrick Wendell: AI coding costs went exponential, smart routing cuts them 30%
Key Points
- AI coding tool costs spiraled exponentially at Databricks and other large tech firms until smart routing and model substitution brought spending under control.
- Databricks' Unity AI Gateway cuts average task costs by 30% by routing queries to the most efficient model for each workload.
- Databricks sees software engineering as just the opening act; automating knowledge worker tasks like data analysis could unlock vastly larger enterprise markets.
Summary
Read full transcript →Databricks' AI cost problem — and how it solved it
Patrick Wendell, co-founder and VP of Engineering at Databricks, oversees AI product and has been the internal champion for rolling out AI tools across the company's 10,000-person workforce. That position put him in the room when the costs started getting uncomfortable.
The exponential cost curve
The productivity case for AI coding tools was clear early. Wendell says Databricks saw roughly double the engineering output from fixed-size teams, with some highly optimized teams moving even faster. But consumption-based pricing — no seat fees, just usage — meant a single developer could run a loop on the most expensive model and spend an unbounded amount of money. The result was an exponential cost curve that, left unchecked, threatened to wipe out the efficiency gains entirely.
Databricks wasn't alone. Wendell says the company compared notes with Coinbase, Uber, and other large tech companies that had given tens of thousands of employees broad access to coding tools, and collectively they landed on techniques that work.
“Once we got people to use it, we just started seeing this exponential cost curve. These tools all do consumption pricing now. So we're not paying a fixed seat per user. A user can, in principle, spend an unbounded amount of money. They can run a little loop on the most expensive model. And so we started seeing basically this exponential growth curve that, you know, although we were getting the two x or more output from our engineering teams, it's just — if your costs are growing exponentially, you're gonna hit a problem.”
What actually cuts costs
The most impactful lever, Wendell argues, is model substitution. New models — proprietary and open-source — drop roughly weekly, and one every week or two lands on a new cost-efficiency frontier. When that happens, the key is moving traffic to the cheaper model immediately. No behavior change required from engineers, no fancy routing logic. Costs just fall.
Smart routing is the second lever. Wendell says Databricks' Unity AI Gateway cuts average task cost by around 30% by exploiting the fact that different models have different strengths. The routing models themselves have to be small and fast, he notes, because using a large model to route queries defeats the purpose.
The broader deflationary dynamic matters here. Wendell frames the model market as structurally favorable for buyers: margins on the underlying AI models are compressing, which makes the routing and optimization layer above them an increasingly attractive business. Unity AI Gateway now has thousands of customers.
Beyond coding
Software engineering remains the dominant AI cost center at enterprises, Wendell says, because software is a digital artifact that AI can iterate on continuously without waiting for human response time. The contrast with chat-style use cases is direct: when a person asks a question and reads the answer, the meter is bottlenecked on human thinking speed.
The next wave he sees is automating knowledge worker tasks for less technical employees — data workers sifting through tables, running queries, reconciling metrics in spreadsheets. Databricks has products targeting that population, which in a typical enterprise can outnumber software engineers by orders of magnitude.
Routing is an asset-light business model, which means competition will intensify. Wendell's counter is that doing it well is genuinely hard. Databricks has a research team focused on the practical problems of using models at scale, and he argues there is significant IP in doing that optimization correctly across 20,000 customers.
Every deal, every interview. 5 minutes.
TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.