What can the model work with?

A context window describes the token budget available to a model during a request. Depending on the API, that budget can include instructions, conversation history, retrieved text, tool results, generated output and reasoning tokens.

The exact accounting is model-specific. Always check the provider’s documentation before assuming that the advertised context window is entirely available for a document you want to send.

How much can it generate?

A maximum output limit sets a separate ceiling on the generated part of a request. Some APIs count internal reasoning against that ceiling; the visible answer can therefore be shorter than the number suggests.

For a concrete specification, OpenAI’s GPT-6 Astra reference lists a 1,050,000-token context window and a 128,000-token maximum output. Those two fields describe different capacities.

By contrast, the one-million-token figure in Google’s Gemini 4 Argon announcement refers to the expanded output limit. It should not be relabeled as an input-context specification.

A small example

Imagine a document review. You need the model to read a large contract but return a one-page summary. A generous input budget matters more than an unusually large output cap.

Now imagine generating a long technical report from a short brief. Your input is small, but the answer may need much more output headroom.

RequirementCheck first
Read a large body of materialContext budget and input constraints
Produce a long answerMaximum output and reasoning accounting
Reuse a long conversationHistory, tool-result and truncation behavior

Neither number tells you whether the answer will be accurate. Capacity and quality are separate questions.

Read the whole specification

Before choosing a model, check:

  • Whether input and output share a total budget.
  • How reasoning tokens are counted and billed.
  • Any per-file, image, audio or video constraints.
  • The behavior when a request exceeds a limit.
  • The specific model version and API you will use.

A large token limit is room to work. It is not a promise that every token will be used well.

Try a representative document, inspect the result and record the actual token usage. That experiment tells you more about your workflow than the largest number on a specification sheet.