Verified limit lookup

Mistral API Context Windows by Model

Mistral API Context Windows by Model, verified against Mistral API's official documentation with scope, implementation impact, caveats, and a direct check.

Verified Aug 22, 20261 official source
Quick answer

Mistral API documents current model context range as 128K or 256K tokens for documented current models. Input and output tokens together count toward the model context window.

Verified Aug 22, 2026Official source

Current limits

ConstraintCurrent valueScopeVerified source
current model context rangeInput and output tokens together count toward the model context window.128K or 256K tokens for documented current modelsModel specificMistral AIAug 22, 2026

Why does this limit matter?

Reserving no output space can turn an otherwise valid input into a 400 response.

This value is scoped to Model specific; a different plan, runtime, model, endpoint, region, or account can produce a different effective constraint.

What should you check?

  1. Match the deployed model ID to the limitations table and reserve max_tokens.
  2. Confirm the exact plan, model, runtime, endpoint, region, and account that serve the failing workload.
  3. Record the observed value, response headers or configuration, timestamp, and source without logging secrets.

Important caveats

  • The model cards are the authoritative current list and may add models with different windows.
  • Treat the official source and live account configuration as authoritative if they differ from this verified snapshot.
HyperObserve reports the documented platform constraint. Your application, SDK, gateway, provider, region, or account can impose a lower effective limit.
Related

Related references and tools

Found an outdated limit? Report it.