Generation admission

Sume generation APIs use queue-first admission for paid work that runs for a long time. A submit request creates a durable job when three conditions are true. The request is valid, Sume can reserve the balance, and the workspace still has capacity for accepted jobs. The job can start immediately, or it can wait in queued until a workspace concurrency slot opens.

Concurrency is a dispatch limit, not a submit limit. If your workspace is already at its generation concurrency limit, Sume can still accept more jobs as queued. This is true while queue capacity remains. Later, workers move queued jobs to processing under the concurrency guard of each workspace.

Limits at a glance

Sume separates four controls that are easy to confuse:

ControlApplies toWhat happens when full
Generation concurrencyPaid generation jobs with status processing.Sume can still accept new valid jobs as queued if queue capacity remains.
Queue capacityPaid generation jobs that Sume accepted and that did not start yet.New paid generation submissions fail with 429 queue_full.
Submit rate limitsRequest volume for public API submit endpoints.Requests fail with 429 rate_limited. Retry with backoff and an idempotency key.
Balance and reservationSpendable USD balance for the authenticated workspace.Generation submit fails with 402 insufficient_credits before provider work starts.

Read/status/list endpoints can also have rate limits. Treat those limits as poll backpressure, not generation concurrency.

Processing concurrency

Generation concurrency is plan-only. Prepaid top-ups do not increase the processing concurrency limit. Admin overrides can increase the effective concurrency_limit (limit_source: admin_override). The default queue capacity is max(3, concurrency_limit × 5).

PlanProcessing concurrencyQueue capacity (default)Accepted job capacity
Free156
Pro42024
Startup84048
Scale20100120
Enterprise20100120

The dashboard Concurrency tab is the source of truth for the configured processing cap of the workspace. The API shows this cap as generation_limits.concurrency_limit. Org workspaces have a floor of 10. Enterprise has a default of 20 and uses admin overrides for higher contract limits. Always prefer the effective field to this static table.

accepted job capacity is concurrency_limit + queued_jobs_limit. This value is the maximum number of paid generation jobs that can be processing or queued for the workspace at the same time.

Queue-first behavior

If your workspace has concurrency_limit: 1, you can submit several valid jobs at the same time. Sume can return all of them as queued while balance and queue capacity are available. Only one generation job of the same workspace can move to processing at a time.

This is the intended behavior:

Do not treat queued as a failure. Store the job_id. Poll the status with backoff. Fetch the result only when the job reports result_ready: true or status: completed.

Immediate rejection

Sume rejects a request immediately only when it cannot safely accept the request.

StatusCodeWhy it happensClient behavior
400invalid_requestThe request body, model id shape, mode, webhook options, or headers are not valid.Correct the request before you retry.
401unauthorizedThe API key is missing, malformed, revoked, or not valid.Correct the authentication.
402insufficient_creditsSume cannot reserve the estimated generation cost from the workspace balance.Upgrade the plan / wait for included Gen$, or submit a less expensive request. Do not invent prepaid top-ups.
404model_not_found or not_foundThe public model or resource does not exist in this workspace.Use /v1/catalog, or make sure that the ids are correct.
409idempotency_conflictA client used the same idempotency key again for a different operation or payload.Use a key again only for an exact retry.
429queue_fullThe workspace has no remaining accepted generation capacity.Wait for jobs to finish, or cancel queued jobs. Then retry with the same idempotency key.
429rate_limitedAPI request volume exceeded an abuse-protection limit.Use retry-after for the backoff when it is present.
503provider_capacity_exceeded or runtime configuration errorsSume cannot start or dispatch generation work safely.Retry later with the same idempotency key, unless the error tells you not to retry.

Full concurrency alone is not an error. It becomes a submit error only when the queue is also full.

generation_limits

Generation submit responses include generation_limits when Sume can compute the workspace admission snapshot.

The fields have these meanings:

FieldMeaning
plan_idSubscription plan that sets the default concurrency map.
limit_sourceplan or admin_override for the effective concurrency.
plan_concurrency_limitPlan-default processing concurrency (when an override applies, do not use it to calculate the wave size).
concurrency_limitEffective maximum number of paid generation jobs in the same workspace that can be processing.
queued_jobs_limitMore paid generation jobs in the same workspace that can wait in queued.
accepted_generation_jobs_limitconcurrency_limit + queued_jobs_limit.
active_generation_jobsCurrent generation jobs in the same workspace with status processing.
queued_generation_jobsCurrent generation jobs in the same workspace with status queued.
queue_capacity_remainingThe remaining queued-job budget plus the idle processing seats, before queue_full.
wave_size_hintSubmission-wave hint only: max(1, floor(queue_capacity_remaining * 0.75)). Not a concurrency limit, override, or processing width. Never use it to calculate the size of in-flight work.

The counts are a snapshot. They can change immediately after the response, when workers claim jobs or other clients submit work.

Sizing in-flight work

Use max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs) as the budget for new in-flight work. Limit that budget to queue_capacity_remaining. Count each newly submitted job against that budget until the next live snapshot. At zero headroom, wait and refresh the preview before you submit more. If the counts are not available, refresh before you select a width.

This pace control on the client keeps open work in the processing cap. The API can still accept queued work under its separate queue-first admission policy.

For example, concurrency_limit: 100, queued_jobs_limit: 500, and no active or queued jobs give queue_capacity_remaining: 600 and wave_size_hint: 450. The workspace is set to 100, with a maximum of 100 new in-flight jobs in this snapshot. 450 is only a submission-wave hint, and it includes queue slots. With 30 processing jobs and 10 queued jobs, the new in-flight budget is 60.

Never show the hint as concurrency. Do not use the pre-override plan_concurrency_limit / purchased_concurrency_limit fields in its place. A full queue still means that you must wait, although the minimum value of the hint is 1.

Before bulk submissions

For launch integrations, use GET /v1/balance and the generation_limits from generation submit responses to make conservative queue decisions. Later, the live OpenAPI can show a read-only admission preview endpoint for your environment. If it does, treat that endpoint only as an optional preflight. It must not create a job, reserve credits, capture usage, refund usage, or call generation providers.

Submit and poll pattern

For production integrations, prefer async submit with an idempotency key.

Then poll the status:

Fetch the result after completion:

Recommended client behavior:

  • Treat queued and processing as normal non-terminal states.
  • Use exponential backoff for polls. Do not use tight loops across many jobs.
  • Continue to poll until terminal: true, or until your own application deadline.
  • Use Idempotency-Key for each paid submit that a client can retry.
  • Do not submit a paid request again only because your local worker timed out.
  • Store status_url, result_url, events_url, and cancel_url when they are present.
  • Examine generation_limits. When queue capacity is low, do not add more work.

The Sume CLI uses the same model:

Queue-full handling

queue_full means that the workspace used all of its accepted generation capacity:

The error details can include a generation_limits snapshot and job metadata for the failed admission attempt. When applicable, Sume releases or refunds the reservation for the failed admission.

When you receive queue_full:

  • Do not add more generation work for that workspace.
  • Poll the current jobs until at least one job reaches a terminal state.
  • Cancel the queued jobs that you no longer need.
  • Retry with the same idempotency key after capacity opens.
  • Use retry-after when it is present.

Cancellation and billing

Paid generation uses public Sume USD estimates. At submit time, Sume reserves the estimated amount when it accepts the request. A successful completion captures the reserved usage. Where applicable, failed jobs and failed queue admission release or refund the reservation.

Cancellation succeeds only before generation work starts:

Cancel the queued jobs that you no longer need, before they start to process. After generation starts, cancel returns 409 job_generation_already_started (details.cancelable: false). The job then completes or fails normally. A cancel of a job that is already canceled is idempotent.

Edge cases and current boundaries

  • Sume currently shows queue counts and the remaining accepted capacity. It does not show a precise queue position or ETA for each job.
  • sync and subscribe modes can wait for a maximum of 30 seconds. When the wait budget ends, continue to poll the job id.
  • Queue expiration and an explicit client-supplied fail-fast queue length are not currently public API options. If they are not in the live OpenAPI schema, treat them as future contract additions.
  • Public API responses are provider-neutral. They do not show hidden provider names, raw provider task ids, raw provider URLs, or internal workflow names. They also do not show storage object keys, API keys, or private workspace/user metadata.