Auto router
Auto router
Don't want to pick a model yourself? Set model to auto and ZeroToken first estimates how hard the request is, then picks a suitable model from the models available to your API key's group. Simple questions go to inexpensive models and hard ones go to stronger models — nothing else in your code changes.
Choose a mode
Set model to any of these routing modes (case-insensitive):
Lite, Standard, Advanced and Reasoning are model tiers. The site admin can pin candidate models for each tier. When a tier is left empty, the available models are tiered by price automatically: the cheapest third is Lite, the middle third is Standard, the most expensive third is Advanced, and models that support reasoning form the Reasoning tier. Within a tier the cheaper model is tried first; the Reasoning tier tries the most capable model first.
Tip: Modes and tiers follow this site's configuration. The "Auto router" section at the top of the model catalog shows which models are in each tier right now.
How difficulty is estimated
Every request is scored before it is forwarded (no extra model call, practically no added latency). The score runs from 0 to 1 and is split into simple, medium and complex by thresholds. The main signals are:
Length: longer questions and longer conversations score higher; Code: code blocks, stack traces, and coding tasks such as "write a function", "implement", "debug" or "refactor"; Math: formulas, LaTeX, or words like "prove", "derive" or "solve"; Analysis and design: "analyze", "compare", "design", "architecture", "distributed", "concurrency", "consistency" and similar; Multiple requirements: several numbered requirements in the question; Request parameters: tools, structured output, a large max_tokens; Simple intents: greetings, short translations and short "what is …" questions lower the score.
If the request explicitly enables reasoning (reasoning_effort of medium or above in Chat Completions, reasoning.effort in Responses, or Anthropic thinking), it is treated as complex.
Admins can also enable model-based judging: when the rule-based score is borderline, a small model rates the difficulty. If that call fails or times out, the rule-based score is used.
Capability matching
Before choosing, models that can't serve the request are filtered out:
requests with images only go to models that accept image input; requests with tools only go to models that support tool calling; requests that explicitly enable reasoning prefer reasoning models; very long inputs only go to models whose context window fits them (an input estimated above 70% of the window does not fit).
If the target tier has no suitable model, higher tiers are tried first, then lower tiers. If the group has no usable model at all, the request fails with a 400 error.
Supported endpoints
The auto router works on:
POST /v1/chat/completions (and /chat/completions) POST /v1/responses (and /responses) POST /v1/messages (Anthropic format)
The request body must contain exactly one model field.
See which model answered
Response headers report the routing result:
In your usage records, these requests show the routing mode you requested (such as auto) together with the model they were routed to.
Billing
You pay the price of the model that actually served the request. The auto router itself is free.
Examples