Table of contents
Open Table of contents
The slide
The routing pitch fits on one slide. Most requests don’t need the frontier model. Send the easy ones somewhere cheap, keep the hard ones on the big model, watch the bill fall. RouteLLM put numbers on it back in 2024: over 85% cheaper on MT Bench, while keeping 95% of GPT-4’s performance.
The slide is true. It’s also the easy part.
Model routing is dirty work. Before the first request, every model’s parameters, pricing, capabilities, limits, and tiers have to be right. Then you tinker — classifiers, routing logic, prompts. Then the decision has to actually execute, through a gateway, which is its own pile of dirty work.
Get it right and the savings are huge. Getting it right is the whole job.
The table has to be true
A router is a function over a table. One row per model, and no row is short:
price input · output · cache write · cache read
a second price table past some context length
reasoning tokens: billed as output, mostly unseen
tiers priority · standard · flex · batch
same weights, different price, different latency
rate limits set by your account's tier, per model
limits context window · max output · request size
params which exist, in which ranges, in which combinations
temperature: 0–2, 0–1, or the default only
forced tool choice: yes, not while thinking, or never
caps tools · parallel tools · streamed arguments
vision · structured output · caching
Nothing in there is one number, and nothing is a boolean. “Supports tools” — in parallel? With streamed arguments? Forced? And the same model name on two providers isn’t the same model: different quantization, different context caps, different tool-call parsing. Moonshot ended up publishing a verifier just to measure it: same K2 weights, tool-call schema accuracy anywhere from 100% down to the low 80s, depending on who’s serving.
Then the table rots. Aliases move to new snapshots. Prices drop. Parameters get deprecated. Nobody emails you. A router with a stale table doesn’t throw. It routes, confidently, to the wrong place.
Then you tinker
The classifier’s job sounds simple: guess how hard this request is. But difficulty isn’t in the text. “fix it.” “continue.” “yes.” Any of those can be the easiest turn of the day or the hardest bug of the week. Difficulty lives in the state of the work, not in the last message.
Misroutes aren’t symmetric. Send an easy request to the big model and you overpay a little. Send a hard one to the small model and it doesn’t fail cheaply — it loops, retries, writes code someone has to rewrite. Price per token isn’t price per task.
Prompts don’t transfer. A prompt tuned on one model is merely fine on the next. Either you maintain a prompt per model, or you ship one prompt that’s mediocre everywhere.
And nobody sees the router. When GPT-5’s autoswitcher broke on launch day, nobody saw a router fail. They saw a model that, in Sam Altman’s words, seemed way dumber. A routing bug is experienced as an intelligence bug.
What are you routing?
This is the question I keep coming back to. Per turn? Per session? Per task? Each is a different trade between cost, consistency, and intelligence.
Per turn is the finest grain, and the biggest savings on paper: every turn gets the cheapest model that can handle it. Then the invoice arrives. In an agent loop the context is huge and mostly cached — and a cache belongs to one model. Say 100k tokens of context, warm on the big model, one more turn, input side only:
stay on the big model ($3/M) cache read 100k × $3/M × 0.10 = $0.030
switch to one 3× cheaper ($1/M) cold write 100k × $1/M × 1.25 = $0.125
come back after the cache expires cold write 100k × $3/M × 1.25 = $0.375
On the turn you switch, the cheap model is the expensive one — four times over. Come back after the cache expires, and that one turn costs more than twelve turns of staying put. Providers price it differently, and the newest models discount cache reads even deeper, which only makes leaving pricier. The shape doesn’t change: a switch buys a cold prefix. It pays off only if you stay, and per-turn routing, by construction, doesn’t stay.
Even a switch that isn’t a switch costs you. On some providers, just turning reasoning effort up or down on the same model starts the cache over.
The cache isn’t the only thing left behind. Reasoning comes back as signed or encrypted blobs, and only the provider that issued them will take them back. Switch, and the next model inherits the transcript without the thinking — what was done, not why. Habits don’t carry over either: one model edits with patches, the next rewrites whole files. The session turns into a patchwork, and the user feels every seam.
Per session picks once and sticks. The cache stays warm, the behavior stays consistent, the reasoning stays continuous. But you’re deciding at the moment you know the least: the first message. Sessions drift. They start with “explain this function” and end in a rewrite of the auth layer. A bad pick is locked in for the whole ride.
Per task routes the unit of work: a subagent with its own context, its own goal, its own check. It’s the most natural unit of difficulty, because a task has an outcome you can verify — so routing can learn from outcomes instead of guessing from text. And a fresh context has no cache to lose. The cost moves elsewhere. Someone has to split the work, usually the strong model, and every handoff leaks context. It’s the Uber rule I quoted in On Harness — default subagents to a weaker model — and it’s only as good as the splitter.
cost consistency intelligence
per turn best, on paper worst the classifier sees one line
per session worst best locked in at the first message
per task in between good only as good as the splitter
There’s no obviously correct abstraction. I don’t think one is coming.
Route at the seams
But there is a question that sorts them: what do you lose when you switch?
- the prompt cache
- the reasoning state
- the habits
- the user’s sense of talking to one mind
A switch costs whatever the next model can’t inherit. So route where there’s nothing to inherit — at the seams:
- A new session. Nothing to lose yet.
- A subagent. Fresh context by design.
- Compaction. The history gets rewritten, so the cache is cold anyway. It’s the one moment mid-session where switching is nearly free — and the summary is the best difficulty signal you’ll ever get. It’s the whole session, compressed.
Inside a context, switch one way only: up, on evidence. Stuck is observable — the same test failing three times, the same file edited back and forth. Easy is a prediction. Pay for a cold cache on evidence. For a guess, wait for a seam.
Route at the seams. Upgrade on evidence.
A 400 is a gift
The router decides. The gateway makes the decision executable: one request shape in, a dozen providers out, everything that comes back normalized. Parameters, schemas, capabilities, errors, streaming, tool calls, and every weird edge case — behind one abstraction.
The loud half of the job:
params max_tokens · max_completion_tokens · maxOutputTokens
temperature: clamp it, drop it, or eat the 400
reasoning effort: a word here, a token budget there
schemas system prompt: a message, a top-level param, systemInstruction
tool schemas: JSON Schema here, an OpenAPI subset there
tool arguments: a JSON string here, an object there
tool results: a message per call, or one turn for all, text last
tool call ids: call_…, toolu_…, or exactly nine alphanumerics
errors 429: slow down, or you're out of credit and retries won't help
overloaded: its own status code
context too long: a different string everywhere
streaming ends with [DONE], or message_stop, or the stream just stops
usage: at the end if you asked, in two halves, or somewhere
tool args: JSON fragments, JSONPath pieces, or all at once
tools a tool call ends in tool_calls, or tool_use, or STOP
forced tools: "required" here, "any" there, refused over there
Every line there fails loudly. A 400 is a gift: it tells you exactly which line you got wrong. The failures that cost you come back as 200s.
- Silently dropped. A cache breakpoint lost in translation. Nothing errors. Every turn is a cold prefix, and you find out from the invoice.
- Translated wrong. One provider counts cached tokens inside the input count; another reports them separately. Map one onto the other naively and you drop them or count them twice. Either way the router’s cost data is fiction, and the router learns from fiction.
- Behaves differently. Pass a STOP through on a turn that called a tool, and any agent that trusts the finish reason ends its loop mid-task. It looks like the model gave up.
- State lost. Strip a thinking signature and the next request either 400s or quietly thinks less.
I wrote in On Harness that absence of an error is not evidence of correctness. That was about agents. Gateways are where it bites hardest.
And fallback has a clock on it. Before the first byte, a failure is just a retry somewhere else. After it, you’ve streamed half an answer, and an overloaded error arriving mid-stream can’t turn into a clean answer from another provider. Decide your fallback before the first token, or fail honestly after it.
Nobody defined compatible
Now make every model work inside every coding agent.
Each agent speaks one dialect. Claude Code speaks Anthropic Messages. Codex speaks the Responses API, and nothing else. Most of the rest speak Chat Completions, or bring their own adapters. And each leans on its home provider’s newest features: cache breakpoints, thinking blocks, encrypted reasoning, server-side tools. N agents times M models, and every cell has its own bugs.
There’s barely a guide for what “compatible” even means. So here’s the one I’d write:
L0 it connects auth, base URL, the model name resolves
L1 it round-trips messages, tools, images survive both directions
L2 it streams same events in the same order, usage included
L3 the loop closes stop reasons agree; the agent knows when it's done
L4 state survives cache hits, reasoning signatures, call ids
L5 it's actually good the model handles this agent's tools and prompts
“OpenAI-compatible” usually means L1. Most bugs live in L3 and L4. L5 isn’t the gateway’s job, and the gateway gets blamed for it anyway.
You still discover issues one by one. The trick is to stop discovering them in production. Record real sessions from each agent — requests, streams, tool round-trips. Replay them through the gateway against every model and check the shape of what comes back: event order, stop reasons, ids, usage. Same lesson as the FDE post: the recording is the spec. Every quirk you find becomes a fixture, and never ships twice.
Converge?
Maybe one day models and providers will actually converge on a common interface. lol.
They already do, at the bottom. Chat Completions became the lingua franca; messages, tools, and streaming are roughly settled. But everything worth routing for ships provider-specific first: caching, reasoning state, server-side tools, context management. That’s where providers compete, so that’s where they’ll never agree. The common interface converges on exactly the parts that stopped mattering.
So a gateway is always one release behind the frontier. Normalize everything and you throw away the features agents depend on. Pass everything through and you’re not an abstraction anymore. Real gateways live in between: normalize the core, pass through the rest, and keep a quirks table that only grows.
The game
What makes a game interesting is limitations. Limited time. Limited resources. Limited options. And you still have to find a way to win.
Routing is that game with an invoice attached. A budget, a latency target, a context window, a rate limit, a dozen models each wrong in its own way. The gateway is the fast feedback loop: a request breaks, you find the quirk, the next one passes. Nobody drew the map. You uncover it a tile at a time.
That’s the dirty work. It’s also the game.