Verified limit lookup

Mistral API Rate Limits

Mistral API Rate Limits, verified against Mistral API's official documentation with scope, implementation impact, caveats, and a direct check.

Verified Aug 22, 20261 official source
Quick answer

Mistral API documents organization rate dimensions as Requests/second and tokens/minute vary by tier and model. The two dimensions are enforced independently and 429 is returned when exceeded.

Verified Aug 22, 2026Official source

Current limits

ConstraintCurrent valueScopeVerified source
organization rate dimensionsThe two dimensions are enforced independently and 429 is returned when exceeded.Requests/second and tokens/minute vary by tier and modelPer organizationMistral AIAug 22, 2026

Why does this limit matter?

Short high-frequency calls and long token-heavy calls can hit different bottlenecks.

This value is scoped to Per organization; a different plan, runtime, model, endpoint, region, or account can produce a different effective constraint.

What should you check?

  1. Use the Admin rate-limit endpoint or workspace settings for current organization values.
  2. Confirm the exact plan, model, runtime, endpoint, region, and account that serve the failing workload.
  3. Record the observed value, response headers or configuration, timestamp, and source without logging secrets.

Important caveats

  • Batch processing does not count against the real-time rate limits described on the source page.
  • Treat the official source and live account configuration as authoritative if they differ from this verified snapshot.
HyperObserve reports the documented platform constraint. Your application, SDK, gateway, provider, region, or account can impose a lower effective limit.
Related

Related references and tools

Found an outdated limit? Report it.