All articles
PracticeUpdated 2026-10-01Reviewed 2026-10-01

How to Calculate AI API Cost by Tokens

A text request costs input tokens at the input rate plus output tokens at the output rate. For budgets, count the whole workflow — not the visible user message: history, retries and cache behavior decide the total.

Short answer

Count the whole workflow — input plus output at the model rates — not a single message. History and retries multiply the bill silently.

The formula and dialog growth

One request: input tokens × input rate plus output tokens × output rate, divided by one million. Output rates are usually multiples of input, so a chatty bot with short questions and long answers costs notably more than a reserved one at the same dialog count.

History is the hidden multiplier: every request carries the conversation tail, so the tenth message costs multiples of the first. Measure median and upper percentile of full input on real dialogs, not one test question.

  • formula: input + output over one million
  • output usually costs multiples of input
  • late-dialog requests cost multiples of early ones
  • measure median and p95 on real dialogs

Cache, images and retries

Caching cuts repeated input cost but only on confirmed cached usage with a stable prefix. Image models bill per generation, video per second — keep them as separate budget lines, never mixed into token math.

Retries after errors are full new requests: retry transient errors with backoff, never 401/402/403. Unbounded max_tokens plus a loop equals an open bill — cap every call.

  • cache only on confirmed cached usage
  • images per generation, video per second
  • retry transient only, with backoff
  • cap max_tokens on every call

Monthly estimate method

Take three load scenarios — current, 3× and 10×. Multiply each by your median input/output from history and the catalog rates. If even 10× fits the budget with margin, the architecture can stay.

Reconcile monthly totals against request history only: decompose the most expensive calls into input, output and cache. An average-check estimate without context and retries always undershoots.

  • three scenarios: 1×, 3×, 10×
  • catalog rates on one date
  • reconcile against history only

Sources and related pages

Next step

Next articleAI API Errors: Handling 401, 402, 429 and 503

Russian version: все статьи на русском.