What can the model work with?
A context window describes the token budget available to a model during a request. Depending on the API, that budget can include instructions, conversation history, retrieved text, tool results, generated output and reasoning tokens.
The exact accounting is model-specific. Always check the provider’s documentation before assuming that the advertised context window is entirely available for a document you want to send.
How much can it generate?
A maximum output limit sets a separate ceiling on the generated part of a request. Some APIs count internal reasoning against that ceiling; the visible answer can therefore be shorter than the number suggests.
For a concrete specification, OpenAI’s GPT-6 Astra reference lists a 1,050,000-token context window and a 128,000-token maximum output. Those two fields describe different capacities.
By contrast, the one-million-token figure in Google’s Gemini 4 Argon announcement refers to the expanded output limit. It should not be relabeled as an input-context specification.
A small example
Imagine a document review. You need the model to read a large contract but return a one-page summary. A generous input budget matters more than an unusually large output cap.
Now imagine generating a long technical report from a short brief. Your input is small, but the answer may need much more output headroom.
| Requirement | Check first |
|---|---|
| Read a large body of material | Context budget and input constraints |
| Produce a long answer | Maximum output and reasoning accounting |
| Reuse a long conversation | History, tool-result and truncation behavior |
Neither number tells you whether the answer will be accurate. Capacity and quality are separate questions.
Read the whole specification
Before choosing a model, check:
- Whether input and output share a total budget.
- How reasoning tokens are counted and billed.
- Any per-file, image, audio or video constraints.
- The behavior when a request exceeds a limit.
- The specific model version and API you will use.
A large token limit is room to work. It is not a promise that every token will be used well.
Try a representative document, inspect the result and record the actual token usage. That experiment tells you more about your workflow than the largest number on a specification sheet.




