Begin with the token bill
For a basic text request, multiply each token count by its corresponding rate. Keep input and output separate, because their prices often differ.
Here is a hypothetical example, not a quote for a particular model:
| Item | Usage | Example rate per million | Cost |
|---|---|---|---|
| Input | 20,000 tokens | $2 | $0.04 |
| Output | 4,000 tokens | $10 | $0.04 |
| Total | $0.08 |
The arithmetic is 20,000 / 1,000,000 × 2, plus 4,000 / 1,000,000 × 10. Any separately charged tools or other services would be added to that total.
This example also shows why output matters. Far fewer output tokens can cost as much as the entire input.
Count the whole attempt
An agent task can involve several model calls, tool calls and intermediate results. A failed attempt may still consume tokens. A second attempt adds another bill, even when only the final answer is shown to the user.
For reasoning models, consult the provider’s documentation about billed reasoning tokens. Do not estimate usage only from the length of the visible answer.
Caching can reduce the cost of eligible repeated input. It does not make every repeated prompt free, and its rules depend on the API. Measure the cache usage reported by the service instead of assuming a cache hit.
Compare the cost of success
A model with a cheaper token rate may need more attempts to produce a usable result. A more expensive model might complete the same task in one call. The opposite can happen too.
The useful quantity is:
Cost per successful task =
total cost of all attempts / number of successful tasks
Define success before running the comparison. For a code change, for example, you could require the patch to pass a meaningful test and a review. A plausible answer alone may not be enough.
Keep a small comparison log
For each model, record actual input, output, cache and reasoning usage; the number of attempts; separately charged tools; time to completion; and whether the result met your criteria.
Use equivalent tasks and budgets. Tokenizers can differ between models, so equal word counts are not necessarily equal token counts.
Compare what it costs to get the job done, with quality held to the same standard.
A token price is a useful starting point. A completed-task measurement is the stronger basis for a decision.




