Quick answer
Perplexity API documents agent api tier range as Tier 0: 1 QPS / 50 RPM; Tier 4: 33 QPS / 4,000 RPM; Tier 5: 33 QPS / 8,000 RPM. QPS and RPM are independent limits and a request must fit both: Tier 1 is 3 QPS / 150 RPM, Tier 2 8 QPS / 500 RPM, Tier 3 17 QPS / 1,000 RPM.
Verified Sep 17, 2026Official source
Current limits
| Constraint | Current value | Scope | Verified source |
|---|---|---|---|
| Agent API tier rangeQPS and RPM are independent limits and a request must fit both: Tier 1 is 3 QPS / 150 RPM, Tier 2 8 QPS / 500 RPM, Tier 3 17 QPS / 1,000 RPM. | Tier 0: 1 QPS / 50 RPM; Tier 4: 33 QPS / 4,000 RPM; Tier 5: 33 QPS / 8,000 RPM | Permanent cumulative-spend tiers | PerplexitySep 17, 2026 |
Why does this limit matter?
Both sustained minute capacity and per-second pacing must fit the tier.
This value is scoped to Permanent cumulative-spend tiers; a different plan, runtime, model, endpoint, region, or account can produce a different effective constraint.
What should you check?
- Read the current usage tier in the API Platform console before sizing worker concurrency.
- Confirm the exact plan, model, runtime, endpoint, region, and account that serve the failing workload.
- Record the observed value, response headers or configuration, timestamp, and source without logging secrets.
Important caveats
- Enterprise or custom capacity can differ from the public Tier 0–5 table.
- Treat the official source and live account configuration as authoritative if they differ from this verified snapshot.
HyperObserve reports the documented platform constraint. Your application, SDK, gateway, provider, region, or account can impose a lower effective limit.
Related