Pricing · the other half
What it costs to run.
Bring-your-own-key means you have a bill — just not from us. That is a real cost and you should see it before you decide, not after. So here is ours — measured, not modelled, from 1,838 real exchanges of somebody genuinely using this thing.
The short version
You choose the model, so you choose the bill.
Companion lets you pick which model answers. That choice moves the cost by more than a factor of twenty-five for an identical piece of work, and it is the biggest lever you have.
| If you pick | Per exchange | 100 / month | 200 / month | 500 / month |
|---|---|---|---|---|
| A fast, cheap model enough for most day-to-day work | $0.0115 | $1.39 | $2.99 | $7.62 |
| A mid-tier model the one most people actually end up on | $0.0596 | $7.30 | $15.51 | $39.99 |
| The most capable model for the work that justifies it | $0.2978 | $36.32 | $77.93 | $200.00 |
Median cost of one exchange, priced onto each model from real recorded usage, assuming the response caching an agent loop should achieve. Two hundred exchanges a month is roughly ten working days of steady use.
Why it costs what it does
You are billed for reading, not writing.
Two facts about agents that nobody puts on a pricing page, and between them they are the whole story.
Across the measured window the ratio of input to output was 154 : 1. An agent that takes seven steps to do something re-sends its instructions and everything it has learned so far on every one of those steps. The answer you read is a rounding error next to the context that produced it.
And the spread between exchanges is enormous. The busiest single exchange in the dataset cost 91× the median one. That is not a billing error — it is one question that needed the agent to open thirty things and think about all of them. It is also exactly why a vendor who buys your inference has to cap you, and why we would rather not be in that position at all.
What this means for you
Pick the cheap model for ordinary work and the expensive one deliberately. The model picker shows each model’s price per million tokens next to its name, so the difference is visible at the moment you choose rather than at the end of the month; what you have actually spent is on the app’s Usage page.
The caveats
What we are not certain about.
- These are reconstructions, not invoices.Every figure here is token counts multiplied by a published rate. Our provider returns the true charged amount on each response and the application currently discards it. We are changing that, and we will reconcile a full month against the real invoice and republish this page with whatever it says — including if it says we were wrong.
- The same model can be priced differently by different providers.Routing one model identifier across sixteen providers, we measured rates from $0.09 to $0.99 per million tokens. Your bill depends on which one you are routed to, which is a setting, not a mystery.
- This is one person's usage.The measured window is a developer building an AI product — an unusually heavy user doing an unusually agentic kind of work. A solicitor reading discovery documents is not that person and should expect the lower rows of the table, not the upper ones.
Window: 15 July 2026 – 2 September 2026 · 1,838 recorded exchanges · 3,744 model calls ·$415.26 of observed spend.