Waiting for runs and jobs

Each Sume run is asynchronous. POST .../runs gives back a receipt with a status_url, and you learn the result later. A bulk queue is a server-side list of those runs. Poll GET /v1/format-run-queues/{queue_id} for counts. The helpers on this page still wait on one run id. Refer to Bulk runs.

In @sume-com/sdk@0.2.0, the Format path for partners is subscribeFormatRun (create + wait). When you already have a run id, use waitForRun. For generation jobs, use waitForJob. Generation jobs are a separate surface. When you can skip the wait fully, prefer run webhooks.

There is no SSE event stream today. Thus, "subscribe" means create, then poll. onStatus shows the result of status polls, not a live log feed. It also gives a snapshot with more data, for example next_action, timestamps, and cancelable. On a Format run, you can also poll the phase timeline. Refer to Watching phases while you wait.

subscribeFormatRun

OptionDefaultNotes
path—Vanity { handle, slug } or { format_id }.
body—The same body as createFormatRun* (input, caps, schema, attachments, …).
idempotencyKeyauto UUIDThe helper sends it as Idempotency-Key. A replay of a finished run returns immediately. To send no key, pass null.
timeout20 minutesLonger than the 10 minutes of waitForRun. Video Formats usually run 10–20.
pollInterval2 secondsThe time between status reads.
signal—Aborts the wait and the in-flight request.
timelinefalseAlso reads the phase timeline on each poll and gives it to onStatus as snapshot.timeline.
onStatus—(status, snapshot) on each status read, and also on the terminal read.
onCreated—The helper calls it one time with the accepted run, before the polls start.

It resolves for any terminal status. It throws SumeRunRequestError when the API refuses the create call. An example is 403 workspace_key_required when a personal key calls a team Format. In that condition, there is no run to wait for. It throws the same error when a status read fails and the failure is not transient. It throws SumeRunTimeoutError when timeout occurs first.

Runs and results documents each field of the receipt.

waitForRun

Use this helper for Action / Agent Completion runs. Also use it when you created the Format run yourself and need only the poll loop.

OptionDefaultNotes
family—Required. "format", "action", or "agent".
clientmodule defaultThe client from createSumeClient(). Pass it. The module default has no base URL or key.
timeout10 minutesIf the wait is longer, the helper throws SumeRunTimeoutError.
pollInterval2 secondsThe time between status reads.
signal—Aborts the wait and the in-flight request. Rejects with the reason of the signal.
timelinefalseFormat runs only. Also reads the phase timeline on each poll.
onStatus—(status, snapshot) on each status read, and also on the terminal read.

family is required. The helper cannot infer it. A run id does not show its surface. The three families live behind three different URL prefixes (/v1/format-runs/…, /v1/action-runs/…, /v1/agent-runs/…). This option also sets the type of the return value: family: "format" resolves as PublicFormatRun.

The helper checks the deadline before it sleeps, not after. If a caller asks for a 5-second timeout, the caller gets the timeout in 5 seconds. The caller does not wait 5 seconds plus one full poll interval.

Watching phases while you wait

status tells you that a Format run is processing. It does not tell you what the run does. For a fifteen-minute video run, that is most of the information that you want. The phase timeline gives this information: preparing, running, finalizing, each with a timestamp and a status.

Pass timeline: true. The timeline then comes on the onStatus snapshot:

For a run that you do not wait on, you can also read it directly:

Know these three things:

  • It is a phase timeline, not a log stream. Sume never publishes agent output, tool calls, or sandbox internals. Refer to Watch a run progress.
  • timeline: true doubles the request rate of the wait. Reads have their own rate-limit budget. Thus, this option cannot cause a 429 on your run creates. But it still sends twice the requests for the same run. Thus, it is off by default.
  • A timeline read that fails does not end the wait. The run still executes and still spends. The helper reports the failure through onTransientError and keeps the previous timeline.

This option is for Format runs only. The helper ignores timeline: true for family: "action" and "agent". The receipts of those runs report events_url: null, because they have no events route.

waitForJob

Runs and jobs are different surfaces. POST /v1/image-1.0/generate, /v1/video-1.0/generate, the Avatar routes, and the /v1/models/sume/…/runs aliases all create jobs at /v1/jobs/:id, not runs. Thus, they need waitForJob, not waitForRun. You cannot use a job id as a run id, or a run id as a job id.

OptionDefaultNotes
clientmodule defaultThe client from createSumeClient(). Pass it — the module default has no base URL or key.
timeout20 minutesLonger than the 10 minutes of waitForRun. Video and avatar-video jobs usually run for minutes. If the wait is longer, the helper throws SumeJobTimeoutError.
pollInterval2 secondsA floor. When next_poll_after_seconds in the status payload asks for a longer time, that value wins.
signal—Aborts the wait and the in-flight request.
onStatus—(status, snapshot) on every status read, including the terminal one.

Submit with mode: "async" (or omit mode). The server-side sync and subscribe modes are the same bounded wait, with a cap of 30 seconds. That cap is a budget for the HTTP request, not for the job. Refer to Communication modes. waitForJob is the client-side wait, and it can wait longer than that cap.

It resolves with the job record from /v1/jobs/:id, not from /v1/jobs/:id/result. For failed and canceled jobs, /result answers 409 job_not_completed, and then the helper has nothing to give back. Read status, result, and error from the record. The errors are the same as the errors of the run helpers: SumeJobTimeoutError and SumeJobRequestError. Both errors carry jobId.

A timeout does not cancel the job. The job continues to run, and it still bills. Store the job id. Read the job again later with getApiJob, or cancel it with cancelApiJob.

Terminal is not the same as successful

Both run helpers resolve for any terminal status: completed, failed, canceled, skipped. A failed run is a result that you asked for, not an exception. Thus, read status and error from the receipt, the same as a webhook handler does:

Give skipped its own branch. It means that a run was already in flight and you passed on_active_run: "skip". (Format runs allow concurrency by default.) Refer to Runs and results for the full lifecycle.

Errors they throw

The generated operations resolve with { data, error }. These helpers are different: they throw, because a poll loop has no place to put a non-result.

ErrorWhen
SumeRunTimeoutErrortimeout occurred first. Carries runId and lastStatus.
SumeRunRequestErrorThe API refused the create, or a read failed and the failure was not transient. Carries runId (or "(not created)") and all the fields of SumeApiError.
the signal's reasonYou aborted.

SumeRunRequestError extends SumeApiError. Thus, the envelope is available as typed fields. You do not have to get it out of body:

SumeApiError subclasses: SumeAuthenticationError (401), SumeInsufficientCreditsError (402), SumePermissionError (403), SumeNotFoundError (404), SumeConflictError (409), SumeRateLimitError (429), SumeServerError (5xx). Each one carries code, requestId, retryable, retryAfterSeconds, nextAction, details, and the raw body. The run helpers always throw SumeRunRequestError itself, never one of these subclasses. Thus, branch on its status or code.

Transient failures do not end the wait

A 429 or a 5xx on a status read means that the read failed, not the run. The run still executes and still spends. Thus, the helpers do a backoff and poll again, and they do not throw. If you lose the handle to a live run, the cost is much higher than one more second of wait.

Two layers do this. Both are on by default:

LayerDefaultOption
createSumeClient retries 408/429/5xx and transport failures2 retries, exponential backoff + jitter, honors retry-aftermaxRetries, timeout
waitForRun accepts consecutive transient read failures6maxTransientFailures, onTransientError

The client retries a POST only when it carries an Idempotency-Key. Without a key, a replay starts and bills a second run. subscribeFormatRun generates a key for you, unless you pass your own key (or null).

The helpers add jitter to polls. Reads have their own rate-limit budget. But without jitter, several clients that start together stay in phase and hit that ceiling as a group.

A timeout does not cancel the run. The run continues. You only stopped the wait. Store the run id. Get the run again later with getFormatRun, or cancel it explicitly with cancelFormatRun.

Prefer a webhook where you can

A poll loop needs one timer and one open request for each run in flight, and runs usually take minutes. Run webhooks deliver the same receipt without these costs. verifyWebhook is the half on the receiver side.

Polls are the correct tool in three conditions. You are in a job that can block without a problem. You make a prototype. Or webhook delivery is still off for your environment. Build the receiver now. Keep subscribeFormatRun / waitForRun as the fallback.

Next