Guides

Rate limits

Each organization gets 300 reads and 60 writes a minute, with at most 10 requests in flight. Read the RateLimit headers and back off on 429.

Limits apply per organization, not per key. Every key and OAuth token in your organization shares them.

BucketLimitCounts
Read300 requests per minuteOperations that only read
Write60 requests per minuteEverything else
Concurrency10 requests in flightAll requests
Timeout30 seconds per requestSlower requests answer 504 timeout

The two buckets are separate. Running out of writes doesn't stop your reads.

Headers

Every response that counted against a bucket carries:

HeaderWhat
RateLimit-LimitThe bucket's limit per minute (300 or 60)
RateLimit-RemainingRequests left in the current window
RateLimit-ResetSeconds until the window resets
Retry-AfterOn a 429 only: seconds to wait
HTTP/1.1 200 OK
Koast-Request-Id: 49654b35-c27f-49c5-9f54-a4f64f7d11d2
RateLimit-Limit: 300
RateLimit-Remaining: 299
RateLimit-Reset: 47

When you hit the limit

Past the limit, calls answer 429 rate_limited until the window resets:

HTTP/1.1 429 Too Many Requests
Koast-Request-Id: a142322f-4394-494b-b6a3-78731e3561b8
RateLimit-Limit: 60
RateLimit-Remaining: 0
RateLimit-Reset: 21
Retry-After: 21
{
  "error": {
    "code": "rate_limited",
    "message": "Too many write requests: this organization's allowance is 60 per minute. Retry in 21 seconds.",
    "details": null,
    "requestId": "a142322f-4394-494b-b6a3-78731e3561b8"
  }
}

With more than 10 requests in flight, the 11th answers 429 rate_limited with Retry-After: 1 and the message Too many concurrent requests for this organization. Retry shortly.

Handling 429

  • Wait the number of seconds in Retry-After, then retry the same request.
  • For a POST, retry with the same Idempotency-Key. See Idempotency.
  • Watch RateLimit-Remaining and slow down before it reaches 0, especially when several jobs share one organization.
  • Keep at most a few requests in parallel. Ten is the hard ceiling for the whole organization.
  • Read metrics less often: a metrics read can be served from a short cache anyway. See Metrics freshness.

The koast CLI does this for you: on a 429 it waits as the API asks, never more than 60 seconds, and retries up to 3 times.

async function call(url, init = {}, attempt = 0) {
  const res = await fetch(url, {
    ...init,
    headers: { ...init.headers, Authorization: `Bearer ${process.env.KOAST_API_KEY}` },
  });
  if (res.status === 429 && attempt < 3) {
    const wait = Number(res.headers.get('Retry-After') ?? 1);
    await new Promise((resolve) => setTimeout(resolve, wait * 1000));
    return call(url, init, attempt + 1);
  }
  return res;
}

Timeouts

A request that takes longer than 30 seconds answers 504 timeout with the message The request did not finish within 30 seconds. It may still complete. The work isn't cancelled. Retry a POST with the same Idempotency-Key to get its result instead of running it twice.

On this page