Waiting for runs and jobs
Each Sume run is asynchronous. POST .../runs gives back a receipt with a status_url,
and you learn the result later. A bulk queue is a server-side list of those runs. Poll
GET /v1/format-run-queues/{queue_id} for counts. The helpers on this page still wait on
one run id. Refer to Bulk runs.
In @sume-com/sdk@0.2.0, the Format path for partners is subscribeFormatRun
(create + wait). When you already have a run id, use waitForRun. For generation
jobs, use waitForJob. Generation jobs are a separate surface. When you can skip the
wait fully, prefer run webhooks.
There is no SSE event stream today. Thus, "subscribe" means create, then poll.
onStatus shows the result of status polls, not a live log feed. It also gives a snapshot
with more data, for example next_action, timestamps, and cancelable. On a Format run, you can also
poll the phase timeline. Refer to Watching phases while you wait.
subscribeFormatRun
| Option | Default | Notes |
|---|---|---|
path | — | Vanity { handle, slug } or { format_id }. |
body | — | The same body as createFormatRun* (input, caps, schema, attachments, …). |
idempotencyKey | auto UUID | The helper sends it as Idempotency-Key. A replay of a finished run returns immediately. To send no key, pass null. |
timeout | 20 minutes | Longer than the 10 minutes of waitForRun. Video Formats usually run 10–20. |
pollInterval | 2 seconds | The time between status reads. |
signal | — | Aborts the wait and the in-flight request. |
timeline | false | Also reads the phase timeline on each poll and gives it to onStatus as snapshot.timeline. |
onStatus | — | (status, snapshot) on each status read, and also on the terminal read. |
onCreated | — | The helper calls it one time with the accepted run, before the polls start. |
It resolves for any terminal status. It throws SumeRunRequestError when the API
refuses the create call. An example is 403 workspace_key_required when a personal key
calls a team Format. In that condition, there is no run to wait for. It throws the same
error when a status read fails and the failure is not transient. It throws
SumeRunTimeoutError when timeout occurs first.
Runs and results documents each field of the receipt.
waitForRun
Use this helper for Action / Agent Completion runs. Also use it when you created the Format run yourself and need only the poll loop.
| Option | Default | Notes |
|---|---|---|
family | — | Required. "format", "action", or "agent". |
client | module default | The client from createSumeClient(). Pass it. The module default has no base URL or key. |
timeout | 10 minutes | If the wait is longer, the helper throws SumeRunTimeoutError. |
pollInterval | 2 seconds | The time between status reads. |
signal | — | Aborts the wait and the in-flight request. Rejects with the reason of the signal. |
timeline | false | Format runs only. Also reads the phase timeline on each poll. |
onStatus | — | (status, snapshot) on each status read, and also on the terminal read. |
family is required. The helper cannot infer it. A run id does not show its surface.
The three families live behind three different URL prefixes
(/v1/format-runs/…, /v1/action-runs/…, /v1/agent-runs/…). This option also sets the
type of the return value: family: "format" resolves as PublicFormatRun.
The helper checks the deadline before it sleeps, not after. If a caller asks for a 5-second timeout, the caller gets the timeout in 5 seconds. The caller does not wait 5 seconds plus one full poll interval.
Watching phases while you wait
status tells you that a Format run is processing. It does not tell you what the run
does. For a fifteen-minute video run, that is most of the information that you want. The
phase timeline gives this information: preparing, running, finalizing, each with
a timestamp and a status.
Pass timeline: true. The timeline then comes on the onStatus snapshot:
For a run that you do not wait on, you can also read it directly:
Know these three things:
- It is a phase timeline, not a log stream. Sume never publishes agent output, tool calls, or sandbox internals. Refer to Watch a run progress.
timeline: truedoubles the request rate of the wait. Reads have their own rate-limit budget. Thus, this option cannot cause a 429 on your run creates. But it still sends twice the requests for the same run. Thus, it is off by default.- A timeline read that fails does not end the wait. The run still executes and still spends. The helper reports the failure through
onTransientErrorand keeps the previous timeline.
This option is for Format runs only. The helper ignores timeline: true for
family: "action" and "agent". The receipts of those runs report events_url: null,
because they have no events route.
waitForJob
Runs and jobs are different surfaces. POST /v1/image-1.0/generate,
/v1/video-1.0/generate, the Avatar routes, and the /v1/models/sume/…/runs
aliases all create jobs at /v1/jobs/:id, not runs. Thus, they need
waitForJob, not waitForRun. You cannot use a job id as a run id, or a run id as a
job id.
| Option | Default | Notes |
|---|---|---|
client | module default | The client from createSumeClient(). Pass it — the module default has no base URL or key. |
timeout | 20 minutes | Longer than the 10 minutes of waitForRun. Video and avatar-video jobs usually run for minutes. If the wait is longer, the helper throws SumeJobTimeoutError. |
pollInterval | 2 seconds | A floor. When next_poll_after_seconds in the status payload asks for a longer time, that value wins. |
signal | — | Aborts the wait and the in-flight request. |
onStatus | — | (status, snapshot) on every status read, including the terminal one. |
Submit with mode: "async" (or omit mode). The server-side sync and
subscribe modes are the same bounded wait, with a cap of 30 seconds. That cap is a
budget for the HTTP request, not for the job. Refer to
Communication modes. waitForJob is the
client-side wait, and it can wait longer than that cap.
It resolves with the job record from /v1/jobs/:id, not from
/v1/jobs/:id/result. For failed and canceled jobs, /result answers
409 job_not_completed, and then the helper has nothing to give back. Read status,
result, and error from the record. The errors are the same as the errors of the run
helpers: SumeJobTimeoutError and SumeJobRequestError. Both errors carry jobId.
A timeout does not cancel the job. The job continues to run, and it still bills. Store
the job id. Read the job again later with getApiJob, or cancel it with
cancelApiJob.
Terminal is not the same as successful
Both run helpers resolve for any terminal status: completed, failed, canceled,
skipped. A failed run is a result that you asked for, not an exception. Thus, read
status and error from the receipt, the same as a webhook handler does:
Give skipped its own branch. It means that a run was already in flight and you
passed on_active_run: "skip". (Format runs allow concurrency by default.)
Refer to Runs and results for the full lifecycle.
Errors they throw
The generated operations resolve with { data, error }. These helpers are different:
they throw, because a poll loop has no place to put a non-result.
| Error | When |
|---|---|
SumeRunTimeoutError | timeout occurred first. Carries runId and lastStatus. |
SumeRunRequestError | The API refused the create, or a read failed and the failure was not transient. Carries runId (or "(not created)") and all the fields of SumeApiError. |
the signal's reason | You aborted. |
SumeRunRequestError extends SumeApiError. Thus, the envelope is available as typed
fields. You do not have to get it out of body:
SumeApiError subclasses: SumeAuthenticationError (401), SumeInsufficientCreditsError
(402), SumePermissionError (403), SumeNotFoundError (404), SumeConflictError (409),
SumeRateLimitError (429), SumeServerError (5xx). Each one carries code, requestId,
retryable, retryAfterSeconds, nextAction, details, and the raw body. The run
helpers always throw SumeRunRequestError itself, never one of these subclasses. Thus,
branch on its status or code.
Transient failures do not end the wait
A 429 or a 5xx on a status read means that the read failed, not the run. The run
still executes and still spends. Thus, the helpers do a backoff and poll again, and they do
not throw. If you lose the handle to a live run, the cost is much higher than one more
second of wait.
Two layers do this. Both are on by default:
| Layer | Default | Option |
|---|---|---|
createSumeClient retries 408/429/5xx and transport failures | 2 retries, exponential backoff + jitter, honors retry-after | maxRetries, timeout |
waitForRun accepts consecutive transient read failures | 6 | maxTransientFailures, onTransientError |
The client retries a POST only when it carries an Idempotency-Key. Without a key, a
replay starts and bills a second run. subscribeFormatRun generates a key for you, unless
you pass your own key (or null).
The helpers add jitter to polls. Reads have their own rate-limit budget. But without jitter, several clients that start together stay in phase and hit that ceiling as a group.
A timeout does not cancel the run. The run continues. You only stopped the wait.
Store the run id. Get the run again later with getFormatRun, or cancel it explicitly
with cancelFormatRun.
Prefer a webhook where you can
A poll loop needs one timer and one open request for each run in flight, and runs usually
take minutes. Run webhooks deliver the same receipt without these
costs. verifyWebhook is the half on the receiver side.
Polls are the correct tool in three conditions. You are in a job that can block without a
problem. You make a prototype. Or webhook delivery is still off for your environment.
Build the receiver now. Keep subscribeFormatRun / waitForRun as the fallback.
Next
- Verifying webhooks — the receiver check for the push path
- Runs and results — each field of the receipt
- Bulk runs — queue many runs, then poll
GET /v1/format-run-queues/{id} - Embed a Format in your product — the whole partner integration