Diagnose
Find out where the bill leaked.
We compare what token pricing implies against what you were actually charged, then break the gap down by cause. Nothing you type is sent to our servers.
New to these terms? 30-second primer
- Token
- The smallest unit an AI reads and writes. Roughly 1-2 tokens per English word; Korean is about 0.3-0.7 tokens per character. Billing is based on this count.
- Input / output
- Input is what you send; output is what the model writes back. Output is usually several times more expensive.
- Reasoning tokens
- What the model thinks through before answering. You never see it, but it is billed as output. This is why bills come in far above expectations.
- Context
- The maximum length of text you can send in one call.
- Cache hit
- The discount you get when the same leading portion is sent repeatedly.
Bill higher than expected?
Enter your actual usage and we point at where it leaked · calculated in your browser only
Copy the numbers straight from your invoice or usage dashboard.
Cost implied by token pricing—
Gap—
Enter your numbers to see the diagnosis
Would self-hosting be cheaper? — break-even
Enter an hourly GPU cost to calculate
Utilisation is almost always overestimated. If you only serve peak hours, real utilisation is often 20-40% — and the GPU bill runs while it sits idle.