Ramp launches Ramp Router, an AI model routing product built on three years of internal use

Jul 22, 2026 · Full transcript · This transcript is auto-generated and may contain errors.

Featuring Veeral Patel

Speaker 2: And so for that crowd, maybe this makes sense, but I you you you introduced this as potential Urus competitor. You thought it was gonna be souped up more like a turbo GT. Yeah.

Speaker 1: I didn't see the EV part.

Speaker 2: But they went EV. I wonder how this will sell. I mean for a lot of Range Rover buyers it's about comfort, it's about quiet, it's about smoothness and EVs can get you there a lot quicker.

Speaker 1: I like the way it looks. It does look beautiful. It's like a good commuter care about autonomous driving.

Speaker 2: Anyway, let me tell you about CrowdStrike. Your business is AI. Their business is securing it. CrowdStrike secures AI and stops breaches now more important than ever. As is our next guest, we have Veeral Patel from Ramp. He's the director of software engineering. He has an exciting announcement for us. How are doing?

Speaker 8: Doing well. How are guys doing?

Speaker 2: We're doing fantastically. Welcome to the show.

Speaker 1: Thank you.

Speaker 2: Give us a little introduction on your background road to Ramp, your how you've ramped up on the team. Yeah. And then and then we can go into the announcement today or this week.

Speaker 8: For sure. Yeah. So I've been at RAMP since the since the beginning. I joined as a founding engineer, worked a lot on our core product team, and more recently have been kind of leading leading the Applied AI team and and launching what we just announced on on Monday, our our RampRouter.

Speaker 2: Yeah. Tell us about the RampRouter. Was this something you built internally first and then sort of productized over time?

Speaker 8: Basically, yeah. So we've been using RampRouter internally for the last three and a half, three years. Mhmm. For like our Three years? Three years, yeah.

Speaker 1: Woah. Okay. So you're using it internally in the product, not even as an organization but deciding when you have Yeah. Basically a task to do

Speaker 2: Yeah. Back then it was identify GPT-four and Gemini

Speaker 1: was like, how do we basically parse a receipt or how do we

Speaker 8: parse Exactly. That We used all of the models for receipt detection, parsing, alcohol detection on our policy agent.

Speaker 2: Oh, sure.

Speaker 8: And we wanted to choose the best models and wanted flexibility. And over time, that's just gotten more and more important. There's new models getting released every other day basically. And so we felt the pain point and we talked to some more customers about it. And now we're releasing it and and and giving everyone access. And so I think it's it's an exciting time to be building AI applications, especially at the application layer. And I think we're we're always have been there for companies to help them save time and money with their TE expenses or their bill pay and now their token costs. So, yeah, it's a really exciting release.

Speaker 2: Yeah. So talk about how the product actually integrates into an enterprise workflow? I mean you can use the receipt processing. Alcohol detection I think is a fun one. Yes. Because I imagine you have to benchmark each model at some point on your workload and then the team can actually understand the trade offs? And then how much of that is driven dynamically based on token price like day to day even?

Speaker 8: Yeah. Exactly. So you would basically replace your base like OpenAI

Speaker 2: Yeah.

Speaker 8: Endpoint with ramps instead, and you can pass in different model slugs. And so you can control if you want to just route all your traffic to to one model

Speaker 1: Yep.

Speaker 8: Or if you want to shadow some models and compare like GPT 5.8 with GLM 5.2 and and you can get the outputs. You can score the results with our with our score. And then in the background, you can actually compare the output and then decide, hey. Do you wanna start moving traffic more traffic over? And Ramp obviously can do this for you automatically, or if you want to control it, you can you can do it yourself too.

Speaker 2: How about walk me through some of the trade offs. Like, if you're on GLM 5.2, are all GLM 5.2 endpoints created equal? Because I imagine that some produce more tokens per second, some might have different prices, also might have different even qualities. You know, you hear about like, oh, this one's been quantized or this one's been nerfed a little bit or they turned down the reasoning on this model post launch. And I imagine that benchmarking is consistent, but then also there's a whole bunch of trade offs that happen even after you've like selected the hot model of the day or the one that makes sense.

Speaker 8: Exactly. Yeah. Beyond just the the model itself, there's different service tiers. So OpenAI, for example, has a flex tier and a standard tier, there's different prices for each. And ramp itself will track what the latency is for this application. You can set a timeout on, like, what you prefer. And and based on that, we'll decide whether to send it to flex tier or standard tier depending on the latency speeds that we're seeing. And so I do think one of the most powerful things here is the fact that we already have, like, these production workloads working for for customers, and it's been really important for us internally. And so we have the proof points of of saving ourselves 30%, maybe even higher soon. And it's just a matter of passing on those same savings now.

Speaker 1: How should how should startups and enterprises, like, think about the significance of this product to Ramp itself? Like, what how much how are what are the resources that you're putting behind it? Because this feels like it feels like deeply aligned to Ramp's mission, but at the same time, going into a category where there's plenty of other companies that want to basically offer this product. Yeah.

Speaker 2: It feels a little bit in the CTO suite as opposed to CFO suite, but they're blending together.

Speaker 8: Yeah. I would say, even even internally, our our our CFOs and CTOs are spending more time together. And Yeah. When we've talked to to more customers, that that story, resonates. And so one of the most interesting things that obviously has been in the news a lot is just how much token costs have become a bigger part of a company's payroll and people have their estimates and budgets. And that's exactly what Ramp has been known for. And so beyond just the router itself on the URL, having all that data flow through and be in Ramp in our token spend management product is, I think, a big part of it. Same way that people have their limits and budgets on their T and E spend where there's been talk about specific companies have token budgets per month or per week. And so we actually launched just last week this product, and you can basically see your token spend alongside your T and E spend. And think, yeah, Eric Eric was on the call last week talking about that. And so it it just makes a lot of sense for those CFOs because they wanna manage that spend better, and then RAMP can be kind of that single pane of glass to do that.

Speaker 2: So how does caching play into this? It feels like that's another way to optimize cost and it would be amazing if it happened sort of more automatically. Yep. What's the future of that look like?

Speaker 8: Yeah. I think one of the I mean, there's there's a bunch of different optimizations we can make if we if we own own the router. As an example, if you're using Cloud Code or Codecs, you'll see as maybe your session is is longer, the the the context loads up and

Speaker 1: Mhmm.

Speaker 8: Your session gets increasingly more expensive. And and sometimes it'd best to just compact that context and start a new session.

Speaker 2: Got

Speaker 8: it. Have it have the model summarize. And so there's interesting experiments like that that we're running internally. And we've we're we're basically gonna do hundreds of these things on behalf of customers and show them exactly what the the before and after kinda kinda looks like here.

Speaker 2: Yeah. How how are you thinking about integrating with tools like Codex and Cloud Code to Yep. Use the UI, UX patterns that users, end users, employees are used to but then still optimize under the hood. There's plenty of situations where you'll give Codex or Clog code just an API key to 11 Labs because 11 Labs can do more efficient, better quality audio generation. Or you might give an API key to all sorts of different things. Is there a world where you can delegate certain tasks to a cheaper GLM 5.2 endpoint, for example, and then have the preferred model and the preferred application still work semi normally?

Speaker 8: Exactly. Yeah. So that's the plan. I mean, it's going to be a partnership with the labs and I the model think one of the interesting things that you see now and will continue to happen is that you'll have kind of jagged capabilities of the models. And maybe one model is really good at writing SDR outbound or another model is really good at writing email copy for the marketing team. And so we'd love to be in a world where RAMP can optimize your use cases for the right kind of business outcome. And I think just be aligned with like, hey, you're just trying to get your work done and then move on with your life and not spend a billion dollars. And so that's that's kinda like what what's really exciting to us. It's beyond just like the starting point, it's like doing this for all types of spend.

Speaker 1: What is Ramp's culture like right now around token consumption? It's, you know Yeah. It's probably the most like aggressively AI native like fintech company or top top three in the world, let's say. But also cost to wear. Yeah. Exactly. It's rare. I would Yeah. Well I would love to see the reaction to one engineer going a little too crazy.

Speaker 8: Yeah. Yeah. No. I mean, it's been fun. I think part of the game and part of what's been fun here is that we were building this product for ourselves.