Model router

Cerno

Nothing reaches your best model by accident.

Cernere: to separate, to decide. Cerno weighs every request before it is dispatched, by rules, in under a millisecond, and never by paying a second model to read your prompt first.

Your machines are a list, most expensive first. One API key for all of them. While the budget is healthy the first machine takes everything; once the throttle starts, light work is shed down the list, and further down as it tightens.

There is no Cerno token yet. Anything calling itself one is not us.
Weight0.01
Throttle0.00
DearestCheapest
Claude Fable 5.1$10 / $50 per million
Looks easy

62% of the budget is left, so the throttle is off and nothing is shed.

4 words adds 0.04 · 1 word that need typing take off 0.25

Not a mock-up. This panel runs size() and land() from lib/, the same two functions the server runs on a real request.

How it weighs

The throttle tightens as the budget falls

Above half the daily budget the throttle is off and the first machine takes everything. Below half it tightens, and anything weighing less than the throttle is shed down the list, so the last of the budget is spent only on the work too heavy to shed. Drag the budget on the panel above and watch the marker move.

00

Your best model

Plans, hard bugs, long refactors. Too heavy to shed, so it holds this machine whatever the throttle is doing.

Claude Fable 5.1, GPT-6 Astra

01

A mid-size model

Most coding work and most tool calls. Sheds here first, once the budget is past half.

Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro

02

A small, fast model

Renames, summaries, lookups, format changes: work that needs typing rather than judgement, and almost no weight at all.

Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4.1 Flash

03

Cerno's own machines

Last on the list and always switched on. No provider key needed, so it takes whatever the rest are too spent or too rate-limited to hold.

Paid from credit

The house floor

One machine that is always ours

Cerno runs its own models on its own hardware, last on the list and never switched off. They answer with no provider key at all, so when everything else you own is spent or rate-limited the work goes there instead of stopping. Every reply names the model that wrote it, in a header, on every request.

Access is paid for by burning the Cerno token from your wallet, which turns into credit for the API.

Cerno GPU: Qwen3 30B A3B$0.10 / $0.40 per million
65,536 context · replies to 8,192

Not serving on this deployment. CERNO_GPU_URL is not set, so there is no worker to send a request to.

Cerno GPU Pro: gpt-oss-120b$0.25 / $1.00 per million
131,072 context · replies to 32,768

Not serving on this deployment. CERNO_GPU_URL is not set, so there is no worker to send a request to.

Setup

Setup is one setting

Make a key in the app, point your client’s base URL at Cerno, and send any model name you like: the one you send is ignored, because picking it is the job. Add your own provider keys whenever you want Cerno routing between your own models.

OPENAI_BASE_URL=https://cerno.fun/v1 ANTHROPIC_BASE_URL=https://cerno.fun API key: your Cerno key
Read the setup guide
Which models does it work with?

Anything you can reach with an Anthropic, OpenAI, Google or OpenRouter key. OpenRouter covers the rest: Grok, DeepSeek, Mistral, Qwen, Llama, Kimi, GLM and hundreds more. You put the models in order on the Ladder page and Cerno works down that list, and only that list.

How does it decide how hard a request is?

By rules, not by asking another model to read your prompt. Three inputs: how long the last message is, with a ceiling so a pasted log file is not mistaken for architecture; which words are in it, from two lists, one of work that needs judgement and one of work that needs typing; and whether tools are attached. The result is one number between 0 and 1, and the dashboard shows the arithmetic for every request.

How do I pay for Cerno?

By burning the Cerno token on Solana. Burning it from your wallet on the Credits page adds credit to your workspace, and the API only works while there is credit. One million tokens adds $5.00. Cerno's own GPU draws on that credit; routing to your own provider keys does not use it up, because that request is billed to you by the provider already and charging for it twice would be charging twice.

Does burning cost more than the credit is worth?

Sometimes, and the Credits page tells you which. The rate is a fixed number of cents; the token's market price moves. So the page prints what a million tokens is worth on the market right next to what burning a million gives you, and when the market read fails it says the price is unknown rather than showing our half of the comparison on its own.

Which limits does it watch?

The daily budget you set for your nearest model, counted from the cost of the requests Cerno itself routed, and rate-limit errors. When a model returns a 429, Cerno reads the provider's own retry-after header rather than guessing a backoff, shuts that one until the moment it named, and the work is carried past it. It does not change the limits of a ChatGPT, Claude.ai or Gemini subscription, and it cannot see them.

Does a conversation stay on one model?

Yes. A conversation stays on the model it started on and only moves when that model is closed, because thinking blocks and prompt caches belong to one model. Throwing a warm cache away to save a tenth of a cent makes the request more expensive, not less.

What happens to my keys?

Provider keys are encrypted at rest with AES-256-GCM and are only ever used to call that provider for you. There is no route that can hand one back. Cerno records the model, the token counts, the cost and the timing of each request: there is no column in the database that could hold a prompt or a reply.

Why not just lower the effort setting?

Try that first. A lower reasoning effort on one model often costs less and keeps a single prompt cache, which is two advantages Cerno cannot give you. Cerno is for when that still runs out before the day does.

Provenance
Price list
This deployment has never fetched the price list, so no third-party price appears anywhere on this site.
The artwork
Generated for this site with npm run art, not photographed and not borrowed. Eight frames of machines actually under load: fans mid-spin, heat shimmer, amber indicators. Nothing in them is switched off and nothing in them is blue.
Domain
https://cerno.fun