Generation admission
Sume generation APIs use queue-first admission for paid work that runs for a long
time. A submit request creates a durable job when three conditions are true. The
request is valid, Sume can reserve the balance, and the workspace still has capacity
for accepted jobs. The job can start immediately, or it can wait in queued until a
workspace concurrency slot opens.
Concurrency is a dispatch limit, not a submit limit. If your workspace is
already at its generation concurrency limit, Sume can still accept more jobs as
queued. This is true while queue capacity remains. Later, workers move queued jobs
to processing under the concurrency guard of each workspace.
Limits at a glance
Sume separates four controls that are easy to confuse:
| Control | Applies to | What happens when full |
|---|---|---|
| Generation concurrency | Paid generation jobs with status processing. | Sume can still accept new valid jobs as queued if queue capacity remains. |
| Queue capacity | Paid generation jobs that Sume accepted and that did not start yet. | New paid generation submissions fail with 429 queue_full. |
| Submit rate limits | Request volume for public API submit endpoints. | Requests fail with 429 rate_limited. Retry with backoff and an idempotency key. |
| Balance and reservation | Spendable USD balance for the authenticated workspace. | Generation submit fails with 402 insufficient_credits before provider work starts. |
Read/status/list endpoints can also have rate limits. Treat those limits as poll backpressure, not generation concurrency.
Processing concurrency
Generation concurrency is plan-only. Prepaid top-ups do not increase the
processing concurrency limit. Admin overrides can increase the effective
concurrency_limit (limit_source: admin_override). The default queue capacity is
max(3, concurrency_limit × 5).
| Plan | Processing concurrency | Queue capacity (default) | Accepted job capacity |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
| Enterprise | 20 | 100 | 120 |
The dashboard Concurrency tab is the source of truth for the configured
processing cap of the workspace. The API shows this cap as
generation_limits.concurrency_limit. Org workspaces have a floor of 10. Enterprise
has a default of 20 and uses admin overrides for higher contract limits. Always
prefer the effective field to this static table.
accepted job capacity is
concurrency_limit + queued_jobs_limit. This value is the maximum number of paid
generation jobs that can be processing or queued for the workspace at the same time.
Queue-first behavior
If your workspace has concurrency_limit: 1, you can submit several valid jobs
at the same time. Sume can return all of them as queued while balance and queue
capacity are available. Only one generation job of the same workspace can move to
processing at a time.
This is the intended behavior:
Do not treat queued as a failure. Store the job_id. Poll the status with backoff.
Fetch the result only when the job reports result_ready: true or
status: completed.
Immediate rejection
Sume rejects a request immediately only when it cannot safely accept the request.
| Status | Code | Why it happens | Client behavior |
|---|---|---|---|
400 | invalid_request | The request body, model id shape, mode, webhook options, or headers are not valid. | Correct the request before you retry. |
401 | unauthorized | The API key is missing, malformed, revoked, or not valid. | Correct the authentication. |
402 | insufficient_credits | Sume cannot reserve the estimated generation cost from the workspace balance. | Upgrade the plan / wait for included Gen$, or submit a less expensive request. Do not invent prepaid top-ups. |
404 | model_not_found or not_found | The public model or resource does not exist in this workspace. | Use /v1/catalog, or make sure that the ids are correct. |
409 | idempotency_conflict | A client used the same idempotency key again for a different operation or payload. | Use a key again only for an exact retry. |
429 | queue_full | The workspace has no remaining accepted generation capacity. | Wait for jobs to finish, or cancel queued jobs. Then retry with the same idempotency key. |
429 | rate_limited | API request volume exceeded an abuse-protection limit. | Use retry-after for the backoff when it is present. |
503 | provider_capacity_exceeded or runtime configuration errors | Sume cannot start or dispatch generation work safely. | Retry later with the same idempotency key, unless the error tells you not to retry. |
Full concurrency alone is not an error. It becomes a submit error only when the queue is also full.
generation_limits
Generation submit responses include generation_limits when Sume can compute
the workspace admission snapshot.
The fields have these meanings:
| Field | Meaning |
|---|---|
plan_id | Subscription plan that sets the default concurrency map. |
limit_source | plan or admin_override for the effective concurrency. |
plan_concurrency_limit | Plan-default processing concurrency (when an override applies, do not use it to calculate the wave size). |
concurrency_limit | Effective maximum number of paid generation jobs in the same workspace that can be processing. |
queued_jobs_limit | More paid generation jobs in the same workspace that can wait in queued. |
accepted_generation_jobs_limit | concurrency_limit + queued_jobs_limit. |
active_generation_jobs | Current generation jobs in the same workspace with status processing. |
queued_generation_jobs | Current generation jobs in the same workspace with status queued. |
queue_capacity_remaining | The remaining queued-job budget plus the idle processing seats, before queue_full. |
wave_size_hint | Submission-wave hint only: max(1, floor(queue_capacity_remaining * 0.75)). Not a concurrency limit, override, or processing width. Never use it to calculate the size of in-flight work. |
The counts are a snapshot. They can change immediately after the response, when workers claim jobs or other clients submit work.
Sizing in-flight work
Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs)
as the budget for new in-flight work. Limit that budget to queue_capacity_remaining.
Count each newly submitted job against that budget until the next live snapshot.
At zero headroom, wait and refresh the preview before you submit more. If the counts
are not available, refresh before you select a width.
This pace control on the client keeps open work in the processing cap. The API can still accept queued work under its separate queue-first admission policy.
For example, concurrency_limit: 100, queued_jobs_limit: 500, and no active
or queued jobs give queue_capacity_remaining: 600 and wave_size_hint: 450.
The workspace is set to 100, with a maximum of 100 new in-flight jobs in this
snapshot. 450 is only a submission-wave hint, and it includes queue slots. With
30 processing jobs and 10 queued jobs, the new in-flight budget is 60.
Never show the hint as concurrency. Do not use the pre-override
plan_concurrency_limit / purchased_concurrency_limit fields in its place. A full
queue still means that you must wait, although the minimum value of the hint is 1.
Before bulk submissions
For launch integrations, use GET /v1/balance and the generation_limits
from generation submit responses to make conservative queue decisions.
Later, the live OpenAPI can show a read-only admission preview endpoint for
your environment. If it does, treat that endpoint only as an optional preflight. It
must not create a job, reserve credits, capture usage, refund usage, or call
generation providers.
Submit and poll pattern
For production integrations, prefer async submit with an idempotency key.
Then poll the status:
Fetch the result after completion:
Recommended client behavior:
- Treat
queuedandprocessingas normal non-terminal states. - Use exponential backoff for polls. Do not use tight loops across many jobs.
- Continue to poll until
terminal: true, or until your own application deadline. - Use
Idempotency-Keyfor each paid submit that a client can retry. - Do not submit a paid request again only because your local worker timed out.
- Store
status_url,result_url,events_url, andcancel_urlwhen they are present. - Examine
generation_limits. When queue capacity is low, do not add more work.
The Sume CLI uses the same model:
Queue-full handling
queue_full means that the workspace used all of its accepted generation capacity:
The error details can include a generation_limits snapshot and job metadata
for the failed admission attempt. When applicable, Sume releases or refunds the
reservation for the failed admission.
When you receive queue_full:
- Do not add more generation work for that workspace.
- Poll the current jobs until at least one job reaches a terminal state.
- Cancel the queued jobs that you no longer need.
- Retry with the same idempotency key after capacity opens.
- Use
retry-afterwhen it is present.
Cancellation and billing
Paid generation uses public Sume USD estimates. At submit time, Sume reserves the estimated amount when it accepts the request. A successful completion captures the reserved usage. Where applicable, failed jobs and failed queue admission release or refund the reservation.
Cancellation succeeds only before generation work starts:
Cancel the queued jobs that you no longer need, before they start to process. After
generation starts, cancel returns 409 job_generation_already_started
(details.cancelable: false). The job then completes or fails normally.
A cancel of a job that is already canceled is idempotent.
Edge cases and current boundaries
- Sume currently shows queue counts and the remaining accepted capacity. It does not show a precise queue position or ETA for each job.
syncandsubscribemodes can wait for a maximum of 30 seconds. When the wait budget ends, continue to poll the job id.- Queue expiration and an explicit client-supplied fail-fast queue length are not currently public API options. If they are not in the live OpenAPI schema, treat them as future contract additions.
- Public API responses are provider-neutral. They do not show hidden provider names, raw provider task ids, raw provider URLs, or internal workflow names. They also do not show storage object keys, API keys, or private workspace/user metadata.