Ramp launches Router to cut AI inference costs by 40%

Ramp Launches a Model Router

Ramp launches Router to cut AI inference costs by 40%

Ramp introduces Router, a model routing service that sends each request to the lowest-cost model meeting performance needs, cutting inference costs by 40% on average. With one endpoint and one bill, it supports models from OpenAI, Anthropic, and others, and is free through 2026. Ramp's internal use cut AI costs by over 25%, and integration with NVIDIA NeMo Switchyard reduced costs by 59%.

At Ramp, Router cut our overall LLM cost by 30% while making our features smarter and faster.

More from this day

2026-08-19