Ramp launches Ramp Router, an AI model routing product built on three years of internal use
Jul 22, 2026 with Veeral Patel
Key Points
- Ramp launches Router, an AI model routing product that dynamically directs inference traffic across providers based on cost, latency, and quality, claiming 30% savings on inference costs.
- The product stems from three years of internal use across Ramp's fintech workloads, giving it production-tested advantages over competing routing gateways.
- Ramp bundles Router with a token spend management dashboard that surfaces AI costs alongside T&E spending, creating a unified expense pane for CFOs and CTOs.
Summary
Read full transcript →Ramp has spent three and a half years routing its own AI inference traffic internally, and this week it made that system available to customers as Ramp Router.
Veeral Patel, director of software engineering and a founding engineer at Ramp, says the product grew out of a real operational need. Ramp was using multiple models for receipt detection, expense parsing, and alcohol detection in its policy agent, and needed a flexible way to choose between them as new models kept arriving. The internal routing layer that solved that problem is now the product.
“We've been using RampRouter internally for the last three and a half, three years. We used all of the models for receipt detection, parsing, alcohol detection on our policy agent. And now we're releasing it and giving everyone access. We've proven it enough times — saving ourselves 30%, maybe even higher soon — and it's just a matter of passing on those same savings now.”
How it works
Companies replace their base OpenAI endpoint with Ramp's and pass in model slugs. From there they can pin all traffic to a single model, shadow multiple models simultaneously, score outputs, and compare results before shifting traffic. Ramp can automate the routing decisions or let teams control them manually.
The product handles sub-model complexity too. For OpenAI's flex and standard service tiers, Ramp tracks observed latency per workload and routes accordingly based on the customer's timeout preference. Patel also points to context-window cost management as a near-term optimization area, citing the pattern where long Claude Code or Codex sessions accumulate expensive context that would be cheaper to compact and restart.
The claimed headline saving is 30% on inference costs, with Patel saying the number could go higher.
CFO and CTO, together
Ramp launched a token spend management product the week before Router, which surfaces token costs alongside T&E spend inside the same dashboard. The pitch to CFOs is a single pane of glass for all AI-related expenditure, the same way Ramp already manages travel and expense budgets. Patel says CFOs and CTOs at customer companies are increasingly in the same room when these conversations happen.
Competitive positioning
Ramp is entering a crowded category. Several infrastructure companies already offer model routing and gateway products. Ramp's differentiation argument is that Router was built under production conditions across real fintech workloads for three years before any external customer touched it, and that the adjacent spend management layer creates a stickier enterprise hook than a standalone routing API.
The longer roadmap Patel describes is routing by workload type rather than just cost and latency, with the example of one model optimized for SDR outbound and another for marketing copy, and Ramp selecting automatically based on business outcome rather than raw benchmark.
Every deal, every interview. 5 minutes.
TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.