Rate limits
The per-key request limit, the headers every response carries, and how an agent should pace itself.
UnifAPI applies one request limit: 60 requests per 60 seconds, per API key. It is scoped to the key, not to the workspace or the endpoint — two keys in the same workspace do not compete for the same budget, and no operation has a tighter limit of its own.
Headers
Every response — 2xx, 4xx, or 5xx — carries the quota that applies to it, so you never need a probe request to learn where you stand.
| Header | Meaning |
|---|---|
RateLimit-Policy | The policy as an IETF structured field: "api-key";q=60;w=60 (q requests, w window seconds) |
RateLimit | Current state, same form: "api-key";r=<remaining>;t=<seconds until reset> |
RateLimit-Limit | Requests allowed in a window |
RateLimit-Remaining | Requests left in the current window |
RateLimit-Reset | Seconds until the window resets — a delta, not a Unix timestamp |
X-RateLimit-Limit / -Remaining / -Reset | Aliases of the three above, for clients that only read the X- spelling |
Retry-After | Seconds to wait, sent on a 429 UnifAPI itself raised |
A response that never resolved an API key — a 401, for example — reports the default policy with a full budget. It describes the quota you will face once authenticated, not usage you have spent.
Browser callers can read all of these: they are listed in Access-Control-Expose-Headers on cross-origin responses.
HTTP/1.1 200 OK
RateLimit-Policy: "api-key";q=60;w=60
RateLimit: "api-key";r=57;t=60
RateLimit-Limit: 60
RateLimit-Remaining: 57
RateLimit-Reset: 60The window
The window is anchored on the key's most recent request: the counter clears once a full window passes with no calls from that key. There is no burst allowance, so design for the steady-state rate rather than for a saved-up balance.
Handling 429
async function call(url: string, init: RequestInit, attempt = 0): Promise<Response> {
const res = await fetch(url, init);
if (res.status !== 429 || attempt >= 5) return res;
const wait = Number(res.headers.get("Retry-After") ?? res.headers.get("RateLimit-Reset") ?? 1);
await new Promise((r) => setTimeout(r, wait * 1000));
return call(url, init, attempt + 1);
}Two rules to live by:
- Pace before you are throttled. Read
RateLimit-Remainingon every response and slow down as it approaches0. Waiting for a429to start pacing wastes a round trip on every worker at once. - Honor
Retry-After. It is the wait the limiter actually needs; guessing usually makes things worse. If it is absent, fall back toRateLimit-Reset. Raising concurrency after a429never helps — the limit is per key.
Asking for more
Limits are set per key, so a higher limit is a configuration change rather than a plan upgrade. Email support@unifapi.com with your workspace ID, the Skill or operation you are running, and a rough QPS estimate.