The API cost formula
Monthly API cost equals request volume multiplied by the average uncached input, discounted cached input and output token charges. Prices in the calculator are dollars per one million tokens. The defaults are neutral examples, not a quote from any provider.
Caching only changes the result when the provider offers a cached-input price or your own stack avoids repeated prefill. Enter the observed share of eligible input and the actual discount. A claimed cache rate without production evidence can make a cost model look much better than the bill.
The self-hosted cost formula
The calculator estimates monthly hardware cost from GPU count, hourly price and active hours. It estimates output capacity from measured output tokens per second, utilization and active seconds. Dividing cost by delivered tokens produces a hardware-only cost per million output tokens.
This intentionally does not hide software engineering, networking, storage, observability, failed capacity or model-quality work inside the number. Add those costs before making a purchasing decision. For bursty traffic, the utilization assumption is usually the most sensitive variable.
How to use the result
Treat a result within 20 percent as a tie until you have a production-shaped load test. Prefer the option with lower operational risk and better scaling behavior. If self-hosting only wins at nearly perfect utilization, it is not yet a robust economic win.
Recalculate when prompt length, model size, provider price, cache policy or latency target changes. Preserve each scenario with its date and assumptions so future comparisons remain auditable.