AI Routers, Explained: The Layer That Decides What Tokens Cost

Rachid Idali

by Rachid Idali

For most of the last two years, "router ai" was a rounding error in our database. Two years ago it did 140 US searches a month. A year ago, 720. Then July 2026 came in at 1,000 and August 2026 came in at 1,600, a 60% jump in a single month and the highest the term has ever been.

We know exactly which Wednesday did that. On August 19, 2026, Stripe said it had agreed to buy OpenRouter, and Ramp launched a free public model router on the same day. An AI router is the proxy that sits between your app and every model you could call, reads each request, and sends it to the cheapest model that can still do the job. It stopped being an infrastructure hobby that week and became a budget line item.

Key takeaways:

  1. "router ai" is up 8.3x in two years, from 140 to 1,600 US monthly searches, classified EXPONENTIAL in the Rising Trends database (data as of Aug 2026).
  2. The whole jump landed in one month. July 2026 was 1,000, August 2026 was 1,600, the month Stripe agreed to buy OpenRouter and Ramp shipped Router.com.
  3. The brand is 154 times bigger than the idea. "openrouter" does 246,000 searches a month. The generic term describing what it is does 1,600.
  4. Cursor says its router hit frontier-quality output at 60% savings in online A/B tests across millions of requests. Ramp says customers cut inference costs 40% on average.
  5. The spread being arbitraged is enormous. On OpenRouter's live price list, output tokens run from $50.00 per million down to $0.28 per million, a 178x range.
  6. Simon Willison, reading the same news: "I haven't seen much evidence that model routing is being widely used yet."

Let's get into it.

The curve, and the month that made it

Here is "router ai" over five years. It is a flat line with a hook on the end.

Line chart of monthly Google search volume for router ai over 60 months, from 10 in Sep 2021 to 1,600 in Aug 2026, peak 1,600 in Aug 2026

Source: Rising Trends database, data as of Aug 2026

The term first crossed a quarter of its current peak in February 2025, which is when it stopped being noise. It then spent eighteen months climbing slowly, and finished with one vertical step: 1,000 in July 2026, 1,600 in August 2026. Our data classifies it EXPONENTIAL with a 1.4 acceleration, meaning the last three months are running 40% hotter than the three before them.

Put it next to the rest of the vocabulary and you can see where it sits.

Horizontal bar chart: Related searches around "router ai". ai inference 2,900, router ai 1,600, llm router 880, model router 720, ai inference platforms 260, ai inference costs 140

Source: Rising Trends database, data as of Aug 2026

Three different names for the same object are all moving at once: router ai at 1,600, llm router at 880 and model router at 720. A category that has not agreed on its own noun yet is a category that is early.

Now the part that reframes all of it. The generic terms are small because the brand ate them. "openrouter" does 246,000 searches a month, which is 154 times the concept term. People do not search for what the thing is. They search for the company that does it.

Horizontal bar chart: Search growth over the last year. openrouter valuation +809%, ai inference costs +250%, ai inference platforms +189%, openrouter +172%, router ai +122%, ai inference +53%, model router +50%, llm router +49%

Source: Rising Trends database, data as of Aug 2026

The growth ranking is the tell. The fastest-moving term in the whole family is "openrouter valuation", up 809% in a year at 1,000 searches a month. Second is ai inference costs, up 250%. Meanwhile the plain, technical "ai inference" is up only 53%.

Nobody is suddenly curious about inference as an engineering topic. They are curious about what it costs and what the company that optimises it is worth.

What an AI router actually does

Strip the marketing off and it is a dispatcher with a price list.

Your application sends one request to one endpoint. Before any model runs, a classifier reads the request and decides which model gets it, based on how hard the task looks, how long the context is, what the task is for, and what each candidate model charges. Simple work goes somewhere cheap. Hard work gets escalated.

Microsoft's Foundry model router documentation is unusually concrete about the trade being made. In its default Balanced mode, the router considers every model "within a small quality range (for example, 1% to 2% compared with the highest-quality model for that prompt)" and picks the cheapest one in that band. Switch to Cost mode and the band widens to "5% to 6%". Quality mode ignores price entirely.

That is the whole product in one sentence. You are choosing how many quality points you will sell for how much money.

OpenRouter's Auto Router, released back on November 8, 2023, does it differently and more strangely. It routes on "what the OpenRouter community collectively spends on for exactly the kind of task your prompt represents", over a trailing seven-day window. It is a crowd-sourced routing table, which means it updates itself the moment a new model gets popular. Multi-turn conversations stick to one model until it stops being a leading choice, which matters for prompt caching more than most people realise: every switch of model is a cache miss you pay for.

Why it broke out in August

Because two things landed on the same day, and one of them had a nine-figure price attached.

Stripe newsroom page dated August 19 2026 announcing it has agreed to acquire OpenRouter, the AI model gateway routing across 400 plus models

Source: stripe.com, captured 2026-09-29

Stripe's August 19, 2026 announcement confirms the deal without naming a price. Press reporting put it above $7 billion. What the release does give you is scale: OpenRouter routes "across 400+ models from more than 80 providers" and is "already used by the likes of NVIDIA, Zoom, and Lovable".

The reasoning is the interesting part. "Tokens are the central currency for companies building with AI," Patrick Collison, Stripe's cofounder and CEO, said in the release. A payments company just bought a routing layer because it decided token spend is a payments problem.

OpenRouter cofounder and CEO Alex Atallah put the thesis more plainly: "We believe intelligence will be multi-model: no single model will be optimal for every task, and developers need a neutral layer to orchestrate and manage them all."

The same day, Ramp made Router.com public. It is a free-through-2026 endpoint that Ramp had been running internally for three years, and Ramp says it cut its own inference bill about 30% and cuts customer bills 40% on average. Ramp's CTO Rahul Sengottuvelu framed it as an accounting problem in the launch release: "AI is the fastest-growing line item at most companies, and the one they can least measure."

That is why the search curve moved. Not a research paper. A spend problem with a bill attached.

The version of this argument getting passed around looks like this clip, posted on September 18.

The claim in it, that routing cuts AI bills by 40% to 70%, is the creator's, not ours. It is worth noting because it is the shape the pitch has settled into: a percentage off a bill, not a capability.

Who is actually shipping one

Four real products, one reported one, and a local-hardware outlier.

WhoWhat it isTheir published number
OpenRouter (Stripe)The neutral gateway, since 2023400+ models, 80+ providers
Cursor RouterCoding-specific router, GA July 22, 202660% savings at frontier quality
Ramp RouterFree public gateway, launched Aug 19, 202640% average cost cut
Microsoft FoundryManaged router inside Azure32 regions, 27-model pool
Meta SwitchboardReported internal projectNot shipped
NVIDIA PAIRRoutes local inference across your own machinesPublic beta, Apache 2.0

Cursor's router post is the most useful document in the category, because it publishes the metric everyone else hides. Cursor trained its classifier on 600k+ live requests and measured it in online A/B tests instead of offline evals, specifically because offline evals "omit the extra cache-miss cost that comes from switching models".

Its finding on why routing pays at all: roughly 60% of developers pick a single model as their daily driver, which bills routine work at frontier prices. Three high-volume early-access accounts saved 30% to 50% versus routing everything to Opus 4.8, with no drop in quality. The same company that built Cursor Origin is now selling the decision layer above the models.

Microsoft's dated release notes show the feature maturing on a normal enterprise schedule. Automatic failover arrived in March 2026, so a wobbling endpoint gets rerouted instead of erroring. In August the routing pool went from two Azure regions to 28. In September it hit 32, and picked up session affinity plus per-request routing metadata in preview, so you can finally see which model served you and why.

That last one matters more than it sounds. A router you cannot audit is a black box between you and your bill.

What the routing decision actually costs

Here is the number nothing else ranking for this term will give you. We read OpenRouter's live model list on September 29, 2026. It carries 460 model entries. These are the list prices per million tokens.

ModelInputOutput
Claude Fable 5$10.00$50.00
Claude Opus 4.8$5.00$25.00
GPT-5.6 Sol$2.00$10.00
Kimi K2.7 Code$0.66$3.30
Qwen3.7 Plus$0.32$1.28
GPT-5.6 Luna$0.20$1.20
DeepSeek V4 Flash$0.14$0.28
Source: OpenRouter model list, list prices read 2026-09-29

The top and bottom of that table are 178x apart on output. That gap is the entire business. A router does not need to be clever to pay for itself, it only needs to be right about which requests are easy, because the penalty for sending an easy request to the top of the list is two orders of magnitude.

Cursor put the same idea in money per unit of shipped work: cost per commit of $6.76 in Intelligence mode and $4.63 in Balance, against $12.69 for Fable 5 and $7.34 for Opus 4.8. Ramp built its own benchmark from production engineering tickets rather than trusting public leaderboards, and routes on that.

This is the same lever as prompt caching and the same lever as the cheap typed-decision models we covered in Jev. Caching cuts what repeated input costs. Small models cut what classification costs. Routing cuts what picking wrong costs. Most teams are still paying all three penalties.

The skeptic's case

It is stronger than the vendors would like.

Simon Willison, responding to the acquisition news on Hacker News: "A bunch of people have been experimenting with automatic model routing recently but mainly as a cost optimization, since tokens for the best models have got expensive once you start piping millions of tokens through them. I haven't seen much evidence that model routing is being widely used yet. I think it's still more of an experimental mechanism right now."

That is worth holding against the search curve. 1,600 searches a month is genuine interest, and it is also a small number. Compare it to the 246,000 for the brand and you get a category where one company has the users and everyone else has a roadmap.

There are three concrete objections.

Every savings number is published by the seller. Cursor's 60%, Ramp's 40%, the 30% internal figure: all vendor-measured, all against baselines the vendor picked. Cursor at least discloses its baseline, which is everything routed to Opus 4.8. That is a generous comparison, because almost nobody actually runs that way.

Switching models breaks your cache. Cursor is explicit that its numbers include cache-miss cost, which tells you the effect is big enough to need disclosing. A router that hops models mid-conversation can spend the savings on re-reading context.

Routing adds a dependency exactly where you cannot afford one. Every request now passes through one more service that can be slow, wrong, or acquired. As one commenter on the same thread put it: "I still find it hilarious that AI is so bad you need something to sit in front of it to pick models for you."

There is also the trust question that routing quietly raises, because a proxy sees every prompt you send. That is the same procurement argument we walked through in ZDR, and it is why OpenRouter's Auto Router explicitly respects retention policies and guardrails when it picks a model.

What to watch next

Whether the concept term catches the brand. Right now people search "openrouter" 154 times more than they search for what it is. If router ai keeps climbing after August while the brand flattens, the category has become bigger than its incumbent, which is when buyers start comparing instead of defaulting.

Whether "ai inference costs" keeps outrunning "ai inference". Up 250% against 53% is a market that has stopped asking how inference works and started asking what it bills. That gap widening is the clearest signal that routing becomes standard plumbing rather than an optimisation project.

Whether Meta ships Switchboard. A reported internal router from a company whose open models lost ground in production would be routing with a thumb on the scale. The first big vendor to ship a router that visibly favours its own models is the moment "neutral layer" stops being the pitch.

The honest summary is that the idea is further along than the adoption. The products are real, the price spread is real, and the savings claims are all marked to the seller's own ruler. What August proved is not that everyone is routing. It is that the people who sell the rails decided routing is where the money sits, and spent $7 billion saying so.


Want to catch infrastructure categories while they are still search curves? Read our guide on how to identify market trends, follow the live router ai trend data, or browse what is breaking out right now on the Rising Trends dashboard.

Unlock More Trends & Insights

Never miss a trend again!

Thousands of Emerging Trends

Thousands of Breakout Apps

Mega Trends

Trend Analysis Tool

Get Access Now
Join +5,000 happy users

Written By

Rachid Idali

Founder of Rising Trends, helping entrepreneurs identify and capitalize on emerging market opportunities through expert trend analysis and insights.

AI Routers, Explained: The Layer That Decides What Tokens Cost