Quick answer
Cursor documents model context window as Up to 1M tokens with extended or long context on listed models. Some models bill long context above 272K tokens at 2× input pricing while others keep the same per-token rate; Max Mode extends context only on legacy request-based plans.
Verified Sep 17, 2026Official source
Current limits
| Constraint | Current value | Scope | Verified source |
|---|---|---|---|
| model context windowSome models bill long context above 272K tokens at 2× input pricing while others keep the same per-token rate; Max Mode extends context only on legacy request-based plans. | Up to 1M tokens with extended or long context on listed models | Usage-based plans; model specific | CursorSep 17, 2026 |
Why does this limit matter?
Attached files, prompts, replies, and tool results compete for the same model conversation budget.
This value is scoped to Usage-based plans; model specific; a different plan, runtime, model, endpoint, region, or account can produce a different effective constraint.
What should you check?
- Confirm the selected model and whether Max mode is active before sizing the task.
- Confirm the exact plan, model, runtime, endpoint, region, and account that serve the failing workload.
- Record the observed value, response headers or configuration, timestamp, and source without logging secrets.
Important caveats
- The selected model determines the available context and pricing tier; Cursor no longer publishes a single 200K figure.
- Treat the official source and live account configuration as authoritative if they differ from this verified snapshot.
HyperObserve reports the documented platform constraint. Your application, SDK, gateway, provider, region, or account can impose a lower effective limit.
Related