Jobs and results
Sume generation endpoints create durable jobs. Store the job id from submit responses, so that your integration can recover work after process restarts.
For paid generation, queued is a normal accepted state. Workspace concurrency
limits apply when workers move jobs into processing, not when the API accepts
valid jobs. Refer to Generation admission for
queue capacity, tier limits, and queue-full errors.
Statuses
| Status | Meaning | Terminal |
|---|---|---|
queued | Sume accepted the request, and the job waits to run. | No |
processing | The job runs, or Sume finalizes it. | No |
completed | The result is ready. | Yes |
failed | The job reached a terminal failure with a public error. | Yes |
canceled | A cancellation request arrived, and the job is terminal. | Yes |
Who can read a job
A job belongs to its workspace and to the member whose key or Agent turn created it. Sume records its usage against that member. Only that member can cancel it.
The reads (GET /v1/jobs/:id, /status, /result, /events, and
GET /v1/jobs) obey one rule:
- An API key reads the jobs that its own member created in the workspace of the key.
- A Studio Agent turn reads each job in the thread that it runs on, whoever created the job. This includes the jobs that an API-fired Format run made in that thread. Thus, a teammate who continues the thread can read the results of the run. An example is the voice and model of each narration job. An interactive turn can also read the jobs of its own member in sibling threads. An unattended run stays on its thread.
- All other reads get
404 not_found: jobs in other workspaces, the jobs of other members in other threads, and, for an API key, any job that its member did not create. Athread_idfilter makes a list smaller. It never increases what a key can read.
When a turn reads a job that another member created, webhook_delivery.url and
webhook_delivery.last_error are null. The turn sees the delivery state, but
not the callback endpoint of the owner.
Poll status
Use exponential backoff. Stop the polls on completed, failed, or
canceled. Do not submit the original paid request again only because a local
process timed out.
jobs_wait
Remote MCP jobs_wait accepts one of these:
job_id— single-job wait. The response shape does not change (object: "job_wait").job_ids— 1–20 ids with optionalwait_for: "all" | "any"(defaultall). The response isobject: "job_wait_batch", with a status snapshot for each requested id.
After parallel fan-outs, prefer one batch wait to N single waits. Each remote HTTP
POST /mcp caller, and also Studio Agent threads, holds a maximum of 55s for each
call (omitted default 50). On wait_slice_expired, retry jobs_wait with the same
ids. Never submit the paid create again. There is no longer a thread-header 600s
server-side hold. An HTTP request that stays open for that long dies at the edge
(502 / Transport send error) before it can answer.
wait_for: "any" still reports each id. The remaining jobs continue and still bill.
Unknown or foreign-workspace ids cause the full call to fail.
If you pass include_results: true, each completed id comes back with its
jobs_result answer in results[]. These are the same entries that a batch
jobs_result returns. Thus, a wave needs no separate result read.
results_omitted.job_ids names the results that do not fit in one answer. Read
those results with one batch jobs_result.
outcome: "operator_stopped" means that Sume operations stopped at least one id
(error_code ops_*, public_reason job_stopped_by_operations). Those jobs
are terminal and produce no output. Sume refunded their holds. The wait answers
immediately when it sees one stopped id. operator_stopped.pending_job_ids names
the ids that still run. If you issue the wait again on the stopped ids, the answer
cannot change.
Slices are bounded, and the server enforces it
On remote MCP, the default of timeout_seconds is 50, and its cap is 55.
(The API accepts values up to 600 and clamps them.) A wait returns at the moment
that its job is terminal. An image that still runs at 50s is usually stuck, not
slow.
A wait is one HTTP request that stays open for the full slice, and no data moves
during that time. Each edge closes such a request at some time. The caller then
gets no tool result, while the job continues to run and to bill. The API clamps a
larger timeout_seconds and does not reject it. The response tells you about the
clamp in wait_slice_clamped.
To wait for a ten-minute render, do the wait again. Do not ask for a longer
wait. A 524 (or 522 / 523 / 525) on jobs_wait is a transport failure,
never a job outcome. Issue jobs_wait again on the same ids, or read jobs_status
one time. Do not submit the paid create again. Do not report the job as blocked.
Fetch a result
Results are available only after completion. If the job is not complete yet, the API returns a conflict response. It does not return an empty result.
Reading a whole wave (MCP)
Over MCP, jobs_result also takes job_ids, with the same 1–20 ceiling as
jobs_wait. Thus, if you waited on a wave in one call, you can read it back in
one call, not in N calls:
The response is a job_result_batch: results[] in request order, with one entry
for each id. Each entry has ok and either value or a typed error.
Partial success is normal and intentional. An id that still runs comes back as
job_not_completed, and each finished id still returns its result.
partial_failure.failed_job_ids names exactly the ids that need a second read. Read
ok for each entry. A failure on one id tells you nothing about the other ids.
Read events
Events give a public timeline that you can use to debug and to recover work:
job.createdjob.queuedjob.startedgeneration.submittedjob.completedjob.failedjob.canceledwebhook.delivery
Public events do not show raw provider task ids or raw provider URLs.
Request cancellation
Cancellation succeeds only before generation work starts. After generation
starts, the API returns 409 job_generation_already_started with
details.cancelable: false, and the job runs to completion. A cancel of a job
that is already canceled is idempotent. It returns the same canceled job.
Communication modes
Each submit endpoint accepts a mode. The mode decides how you learn the
outcome. It never changes whether Sume creates a job, what the job costs, or how
long the job takes to run.
| Mode | HTTP returns | Job id in the first response | Server blocks | What the client does next |
|---|---|---|---|---|
async (default) | 202 with the job envelope and poll URLs | Yes | No | Poll status_url until terminal is true. Then GET result_url when result_ready is true. |
sync | The same envelope, after a wait of up to wait_timeout_seconds (max 30) for a terminal transition | Yes | Yes, at most 30s. Less when waiter capacity is not available. | Terminal? Read the job from the response. Not terminal? Poll. Do not resubmit. |
subscribe | The same as sync: the same bounded wait | Yes | Same as sync | Same as sync. For a fal-style long wait, use the client subscribe recipe below with async. |
webhook | 202 with the job envelope and poll URLs. Sume stores the callback. | Yes | No | Wait for the terminal callback, and verify its signature. Continue to poll as a backup. |
If you omit mode, you get async. If you send webhook_url (or its alias
callback_url) without a mode, you get webhook.
Sume accepts a submit at the moment that it has a durable job id. Thus, every mode
returns the job id in its first response. A 2xx means that the job exists and
paid work is in flight. It does not mean that the job finished. To know the
difference, read terminal and result_ready from the envelope.
and
They run the same bounded waiter and return the same envelope. Sume keeps
subscribe because clients ported from other queue APIs use it. It is not a
long-lived subscription, an event stream, or a longer wait. There is no SSE or
WebSocket transport on the Developer API today. GET /v1/jobs/:id/events is a
pull snapshot, not a stream.
"Subscribe" means three different things
The word occurs on three unrelated surfaces with three different waits. None of them is a push stream. Read the table before you set a timeout.
| Where you see it | What it is | How long it waits |
|---|---|---|
Job mode: "subscribe" (this page) | An alias of sync. One bounded HTTP wait on the submit call. | At most wait_timeout_seconds, with a cap of 30s. |
SDK subscribeFormatRun() (TypeScript SDK) | Create the run, then poll it client-side to a terminal receipt. | Minutes. The SDK uses its own timeout, not an HTTP hold. |
Format / Action / Agent communication.mode | Delivery selection for a run, not a job. The values are async and webhook. subscribe is not one of them. | Nothing blocks. |
Two results are important:
- If you send
mode: "subscribe", you do not get progress events. You get the same 30-second wait thatsyncgets. For progress, submitasyncand readGET /v1/jobs/:id/events, or use a webhook. communication.modehas nosubscribevalue, and its two values operate identically. Only a suppliedwebhook_urlmakes delivery active.
For new integrations, we recommend async (poll, or read events) or
webhook (Sume tells you). sync and subscribe stay supported, and Sume
will not remove them. But they are the wrong tool for work that can last longer
than 30 seconds. Most video work is in that group.
30 seconds is a wait budget, not a job duration
The API clamps wait_timeout_seconds to 0..30. It sets a limit on how long the
HTTP request blocks, not on how long the job can take. Image jobs often
finish in that time. Video, avatar-video, and face-swap jobs usually do not.
When the budget ends, or when the API process has no waiter capacity left and skips the wait fully:
- The response is still
2xxand still carries the job id. The end of the wait budget is not an admission failure. - The envelope carries
status_url,result_url,events_url,cancel_url, and asyncobject.sync.timed_outis true when the wait returned before a terminal state.sync.capacity_exhaustedis true when Sume skipped the wait because the waiter budget of the process was full. - You must continue with
GET status_url. Whennext_poll_after_secondsis present, obey it. If it is not present, do a backoff. - You must not submit a new paid job for the same intent. You can retry the submit itself. Use the same
Idempotency-Keyagain, so that the retry returns the original job and does not bill a second job.
sync is null on async and webhook responses.
Client subscribe: poll a job to a terminal state
This is the official equivalent of a client-side subscribe(). It is the correct
answer for work that can last longer than 30 seconds. Submit with async, poll,
then read the result. The wait lives in your client. Thus, its timeout can be
minutes, and no HTTP request stays open.
Poll on the booleans (terminal, result_ready) or on sume_status. The
status endpoint also returns a queue-shaped status field
(IN_QUEUE / IN_PROGRESS / COMPLETED / FAILED / CANCELED) for clients
ported from other queue APIs. This field maps one-to-one onto sume_status. The
two fields always agree, but do not mix them.
GET /v1/jobs/:id/result is only for completed jobs. For other jobs, it answers
409 job_not_completed. Thus, read the failure from the job record.
Submit (step 1):
The response carries request_id (the job id), status_url, result_url, and
next_poll_after_seconds.
Poll (step 2), until "terminal": true:
Fetch (step 3), when "result_ready": true:
Any HTTP client can run this loop. In Kotlin with OkHttp or Ktor, the shape is:
A client-side timeout does not cancel the job. The job continues to run and still
bills. You only stopped the wait. Store the job id. Get the job again from
status_url, or cancel it explicitly.
In TypeScript, waitForJob from @sume-com/sdk is this loop.
Job webhooks are terminal-only
mode: "webhook" delivers exactly three events to a public HTTPS webhook_url:
job.completed, job.failed, and job.canceled. There are no progress or
partial webhooks. Sume signs the raw body with HMAC SHA-256 over
{timestamp}.{raw_body}. It sends x-sume-webhook-timestamp and
x-sume-webhook-signature: sume-v1=…. Refer to Webhooks for
the payload and a verifier.
A webhook is a delivery optimization, not your only recovery path. Keep the
status_url polls available for missed or retried deliveries.
Action, Format, and Agent Completion runs are a different surface with their
own *.run.terminal events. Refer to Run webhooks.
Idempotency
When you retry after client-side timeouts or network failures, send
Idempotency-Key on submit requests.
Use the same key again only for the same operation and payload.
Result shape
Completed jobs can include artifacts:
Use Sume media URLs from the result. Raw provider URLs are not public API outputs.
A completed text-to-speech job (text_to_speech) also records how Sume made its
audio: model_id (the engine), voice ({ "mode": "id", "id": "…" }),
language, output_format, and the synthesis settings, generation_config and
speed. Each setting is null when the request did not send it. Read these values
from the job to make the next line sound the same.