Skip to content

On Routing

Published: at 02:41 AM

Table of contents

Open Table of contents

The slide

The routing pitch fits on one slide. Most requests don’t need the frontier model. Send the easy ones somewhere cheap, keep the hard ones on the big model, watch the bill fall. RouteLLM put numbers on it back in 2024: over 85% cheaper on MT Bench, while keeping 95% of GPT-4’s performance.

The slide is true. It’s also the easy part.

Model routing is dirty work. Before the first request, every model’s parameters, pricing, capabilities, limits, and tiers have to be right. Then you tinker — classifiers, routing logic, prompts. Then the decision has to actually execute, through a gateway, which is its own pile of dirty work.

Get it right and the savings are huge. Getting it right is the whole job.

The table has to be true

A router is a function over a table. One row per model, and no row is short:

price    input · output · cache write · cache read
         a second price table past some context length
         reasoning tokens: billed as output, mostly unseen
tiers    priority · standard · flex · batch
         same weights, different price, different latency
         rate limits set by your account's tier, per model
limits   context window · max output · request size
params   which exist, in which ranges, in which combinations
         temperature: 0–2, 0–1, or the default only
         forced tool choice: yes, not while thinking, or never
caps     tools · parallel tools · streamed arguments
         vision · structured output · caching

Nothing in there is one number, and nothing is a boolean. “Supports tools” — in parallel? With streamed arguments? Forced? And the same model name on two providers isn’t the same model: different quantization, different context caps, different tool-call parsing. Moonshot ended up publishing a verifier just to measure it: same K2 weights, tool-call schema accuracy anywhere from 100% down to the low 80s, depending on who’s serving.

Then the table rots. Aliases move to new snapshots. Prices drop. Parameters get deprecated. Nobody emails you. A router with a stale table doesn’t throw. It routes, confidently, to the wrong place.

Then you tinker

The classifier’s job sounds simple: guess how hard this request is. But difficulty isn’t in the text. “fix it.” “continue.” “yes.” Any of those can be the easiest turn of the day or the hardest bug of the week. Difficulty lives in the state of the work, not in the last message.

Misroutes aren’t symmetric. Send an easy request to the big model and you overpay a little. Send a hard one to the small model and it doesn’t fail cheaply — it loops, retries, writes code someone has to rewrite. Price per token isn’t price per task.

Prompts don’t transfer. A prompt tuned on one model is merely fine on the next. Either you maintain a prompt per model, or you ship one prompt that’s mediocre everywhere.

And nobody sees the router. When GPT-5’s autoswitcher broke on launch day, nobody saw a router fail. They saw a model that, in Sam Altman’s words, seemed way dumber. A routing bug is experienced as an intelligence bug.

What are you routing?

This is the question I keep coming back to. Per turn? Per session? Per task? Each is a different trade between cost, consistency, and intelligence.

Per turn is the finest grain, and the biggest savings on paper: every turn gets the cheapest model that can handle it. Then the invoice arrives. In an agent loop the context is huge and mostly cached — and a cache belongs to one model. Say 100k tokens of context, warm on the big model, one more turn, input side only:

stay on the big model ($3/M)         cache read    100k × $3/M × 0.10  =  $0.030
switch to one 3× cheaper ($1/M)      cold write    100k × $1/M × 1.25  =  $0.125
come back after the cache expires    cold write    100k × $3/M × 1.25  =  $0.375

On the turn you switch, the cheap model is the expensive one — four times over. Come back after the cache expires, and that one turn costs more than twelve turns of staying put. Providers price it differently, and the newest models discount cache reads even deeper, which only makes leaving pricier. The shape doesn’t change: a switch buys a cold prefix. It pays off only if you stay, and per-turn routing, by construction, doesn’t stay.

Even a switch that isn’t a switch costs you. On some providers, just turning reasoning effort up or down on the same model starts the cache over.

The cache isn’t the only thing left behind. Reasoning comes back as signed or encrypted blobs, and only the provider that issued them will take them back. Switch, and the next model inherits the transcript without the thinking — what was done, not why. Habits don’t carry over either: one model edits with patches, the next rewrites whole files. The session turns into a patchwork, and the user feels every seam.

Per session picks once and sticks. The cache stays warm, the behavior stays consistent, the reasoning stays continuous. But you’re deciding at the moment you know the least: the first message. Sessions drift. They start with “explain this function” and end in a rewrite of the auth layer. A bad pick is locked in for the whole ride.

Per task routes the unit of work: a subagent with its own context, its own goal, its own check. It’s the most natural unit of difficulty, because a task has an outcome you can verify — so routing can learn from outcomes instead of guessing from text. And a fresh context has no cache to lose. The cost moves elsewhere. Someone has to split the work, usually the strong model, and every handoff leaks context. It’s the Uber rule I quoted in On Harness — default subagents to a weaker model — and it’s only as good as the splitter.

              cost               consistency    intelligence
per turn      best, on paper     worst          the classifier sees one line
per session   worst              best           locked in at the first message
per task      in between         good           only as good as the splitter

There’s no obviously correct abstraction. I don’t think one is coming.

Route at the seams

But there is a question that sorts them: what do you lose when you switch?

A switch costs whatever the next model can’t inherit. So route where there’s nothing to inherit — at the seams:

Inside a context, switch one way only: up, on evidence. Stuck is observable — the same test failing three times, the same file edited back and forth. Easy is a prediction. Pay for a cold cache on evidence. For a guess, wait for a seam.

Route at the seams — one session over time: switching mid-context pays a cold cache and drops the reasoning; switching at a new session, a subagent, or compaction is nearly free; inside a context, upgrade only after the same test fails three times

Route at the seams. Upgrade on evidence.

A 400 is a gift

The router decides. The gateway makes the decision executable: one request shape in, a dozen providers out, everything that comes back normalized. Parameters, schemas, capabilities, errors, streaming, tool calls, and every weird edge case — behind one abstraction.

The loud half of the job:

params     max_tokens · max_completion_tokens · maxOutputTokens
           temperature: clamp it, drop it, or eat the 400
           reasoning effort: a word here, a token budget there
schemas    system prompt: a message, a top-level param, systemInstruction
           tool schemas: JSON Schema here, an OpenAPI subset there
           tool arguments: a JSON string here, an object there
           tool results: a message per call, or one turn for all, text last
           tool call ids: call_…, toolu_…, or exactly nine alphanumerics
errors     429: slow down, or you're out of credit and retries won't help
           overloaded: its own status code
           context too long: a different string everywhere
streaming  ends with [DONE], or message_stop, or the stream just stops
           usage: at the end if you asked, in two halves, or somewhere
           tool args: JSON fragments, JSONPath pieces, or all at once
tools      a tool call ends in tool_calls, or tool_use, or STOP
           forced tools: "required" here, "any" there, refused over there

Every line there fails loudly. A 400 is a gift: it tells you exactly which line you got wrong. The failures that cost you come back as 200s.

I wrote in On Harness that absence of an error is not evidence of correctness. That was about agents. Gateways are where it bites hardest.

And fallback has a clock on it. Before the first byte, a failure is just a retry somewhere else. After it, you’ve streamed half an answer, and an overloaded error arriving mid-stream can’t turn into a clean answer from another provider. Decide your fallback before the first token, or fail honestly after it.

Nobody defined compatible

Now make every model work inside every coding agent.

Each agent speaks one dialect. Claude Code speaks Anthropic Messages. Codex speaks the Responses API, and nothing else. Most of the rest speak Chat Completions, or bring their own adapters. And each leans on its home provider’s newest features: cache breakpoints, thinking blocks, encrypted reasoning, server-side tools. N agents times M models, and every cell has its own bugs.

There’s barely a guide for what “compatible” even means. So here’s the one I’d write:

L0  it connects          auth, base URL, the model name resolves
L1  it round-trips       messages, tools, images survive both directions
L2  it streams           same events in the same order, usage included
L3  the loop closes      stop reasons agree; the agent knows when it's done
L4  state survives       cache hits, reasoning signatures, call ids
L5  it's actually good   the model handles this agent's tools and prompts

“OpenAI-compatible” usually means L1. Most bugs live in L3 and L4. L5 isn’t the gateway’s job, and the gateway gets blamed for it anyway.

You still discover issues one by one. The trick is to stop discovering them in production. Record real sessions from each agent — requests, streams, tool round-trips. Replay them through the gateway against every model and check the shape of what comes back: event order, stop reasons, ids, usage. Same lesson as the FDE post: the recording is the spec. Every quirk you find becomes a fixture, and never ships twice.

Converge?

Maybe one day models and providers will actually converge on a common interface. lol.

They already do, at the bottom. Chat Completions became the lingua franca; messages, tools, and streaming are roughly settled. But everything worth routing for ships provider-specific first: caching, reasoning state, server-side tools, context management. That’s where providers compete, so that’s where they’ll never agree. The common interface converges on exactly the parts that stopped mattering.

So a gateway is always one release behind the frontier. Normalize everything and you throw away the features agents depend on. Pass everything through and you’re not an abstraction anymore. Real gateways live in between: normalize the core, pass through the rest, and keep a quirks table that only grows.

The game

What makes a game interesting is limitations. Limited time. Limited resources. Limited options. And you still have to find a way to win.

Routing is that game with an invoice attached. A budget, a latency target, a context window, a rate limit, a dozen models each wrong in its own way. The gateway is the fast feedback loop: a request breaks, you find the quirk, the next one passes. Nobody drew the map. You uncover it a tile at a time.

That’s the dirty work. It’s also the game.