Authentication
Sume uses workspace-scoped API keys to authenticate Developer API requests. You create these keys in the API Keys dashboard.
Send an API key
Use a server-side environment variable:
Then send either Bearer auth:
or the API key header:
The current API accepts both forms. In each integration, use one form
consistently. The Sume CLI uses x-api-key by default. If you configure it, the
CLI can use Bearer mode.
Send exactly one. If a request has both Authorization: Bearer and
x-api-key, the API rejects it with 401 unauthorized and the message
Send only one API key credential. Neither header wins. The second header does
not silently hide the first.
This problem occurs with gateways and fetch wrappers that add their own
Authorization header to a client that already sends x-api-key. Remove one
of the two headers. Do not expect a precedence rule, because that rule does not
exist.
Scope
Sume gets the workspace, owner, and API key metadata from the key. Do not put
workspace_id, owner_user_id, or user_id in public API request bodies.
Responses show key metadata such as id, name, prefix, scopes, and last-used time. Responses do not show the full secret.
When you create a key, Sume fixes its scopes. You cannot add scopes later. A key that you created before a scope existed does not have that scope.
This rule is important for actions:read and actions:write. Calls to
Scheduled via the Actions API must have these
scopes. Sume mints them only on keys that you created after the Actions
API-call trigger shipped. An older key returns 403 insufficient_scope on each
Action run request. Create a new key, then rotate to it.
The same problem applies to formats:read / formats:write. If a pre-Formats
key calls a Format, the result is 403 insufficient_scope, not
404 format_not_found. There is no API to add scopes to a key that already
exists.
Server-side proxy pattern
Browser and mobile clients must call your backend. Your backend must attach the Sume API key.
Validate the user input before you forward requests to Sume. Also enforce your own authorization before you forward the requests.
Rotation
Create a replacement key. Deploy the new key to your server. Use GET /v1/me
to verify the new key. Then revoke the old key from the dashboard. If a key
shows in logs or chat history, rotate it.
Safety rules
- Keep API keys on trusted servers, CI secret stores, or local developer machines.
- Do not put API keys in frontend JavaScript, mobile apps, support tickets, or screenshots.
- Signed upload and download URLs are temporary secrets. Protect them.
- If a key is exposed, rotate keys from the dashboard.
- Give agents read-only commands first. Make explicit confirmation mandatory before write or paid generation commands.
Rate limits
Each API key gets a request budget per minute for all of /v1. The
subscription plan of the workspace that owns the key sets this budget. Reads
and writes have separate budgets. Thus, a tight status-poll loop cannot cause
a 429 on your own submits.
| Plan | Writes per minute | Reads per minute |
|---|---|---|
| Free | 120 | 4800 |
| Pro | 300 | 12000 |
| Startup | 600 | 24000 |
| Scale | 1200 | 48000 |
| Enterprise | Contact sales | Contact sales |
A read is any GET or HEAD, for example a poll of status_url,
events_url, or result_url, or a list of Formats or runs. The two POSTs that
submit nothing are also reads:
/v1/generation/admission-preview and the MCP endpoint itself.
All other requests are writes: run creation, cancellation, and uploads. The plan number is the write number. Reads get forty times that number in their own bucket.
We sized that multiple for agents, not for a person who monitors one run. An agent harvest can hold twenty-odd jobs open and poll each of them. That is thousands of reads a minute for work that costs nothing. Thus, reads are deliberately cheap. The write budget is the tier that a plan actually buys, and we do not change it.
Enterprise is not self-serve. Until Sume provisions a contracted number, an Enterprise key uses the Scale row above.
An MCP tool call spends the write budget one time, for the run that it creates.
It does not spend the write budget for the JSON-RPC request that carried it. A
jobs_status poll over MCP does not spend any write budget.
Do not count requests yourself. Read ratelimit-remaining. On retry-after,
wait before you send more requests. The headers describe the budget that the
current request spent from. A 429 names that budget in error.details.scope
(read or write).
Request rate is not the same as generation capacity. The concurrency limit of
your plan controls the number of generations that run at the same time. The
generation_limits object reports this limit. If you increase your request
rate, the concurrency limit does not increase.
Each response shows the current state:
| Header | Meaning |
|---|---|
ratelimit-limit | Requests permitted in the current window. |
ratelimit-remaining | Remaining requests in the current window. |
ratelimit-reset | Seconds until the window resets. |
retry-after | Seconds to wait, sent on 429. |
Sume limits unauthenticated requests per client IP at the Free rate. Their read bucket stays at four times the write rate, not forty. The agent-sized read budget is for callers who own the jobs that they poll. A larger anonymous bucket only makes the abuse surface larger.
The read multiple is a deployment configuration value
(SUME_COM_API_RATE_LIMIT_READ_MULTIPLIER). Thus, a self-hosted or preview
deployment can be different. ratelimit-limit on the response is always the
authority for the deployment that you send requests to. The table above shows
the shipped default.
Common failures
| Status | Common cause | Next step |
|---|---|---|
401 | Missing, malformed, or revoked key. | Examine the header. If necessary, create a new key. |
403 | The key is valid, but it is not permitted to access the requested surface. | Examine the workspace membership and the key scope. |
429 | Your requests went above the requests-per-minute budget of your plan. | Wait for the retry-after time, then try again. |