Create a run

One POST starts a run. Production integrations send this shape: your data in input, a schema for the result, a per-run spend cap, and a webhook. With the webhook, you do not have to poll.

202 Accepted:

Two status codes mean success. 202 is a fresh run. 200 is an idempotent replay (the same Idempotency-Key with the same body), and returns the original run with idempotency_hit: true. Both carry the full receipt. Store data.id and use the URLs on it. Runs and results covers the next steps.

Build your own call

Fill in the fields to rewrite the cURL, TypeScript, JavaScript and Python snippets. Operations in the TypeScript SDK resolve with { data, error, response }. They do not throw. Thus, examine error before you read data.

Create a Format run

POST /v1/formats/{handle}/{slug}/runs

Required

Address a Format

ShapeExampleUse it when
{handle}/{slug}POST /v1/formats/acme/live-commerce/runsEvery new integration. It is the address that the Format detail page shows, and the shape that production callers use.
{format_id}POST /v1/formats/skl_…/runsA stored URL that must stay valid after a handle or slug rename. Permanent, with the same behavior.
sume/{slug}POST /v1/formats/sume/sume-product-commercial/runsThe Formats by Sume catalog. Any key with the scopes can call it. The run belongs to that key.

The two shapes resolve to the same Format and run the same pipeline: same body, headers, idempotency, caps and receipt. The receipt's format.id is always the opaque skl_… id, for each shape that you use. A renamed handle continues to resolve for 90 days.

A team Format is at the team's handle. Any key created in that workspace with the correct scopes can POST it. Refer to Team Formats need a team key. An unknown handle, an unknown slug, and a handle that you cannot see all answer the same 404 format_not_found.

GET /v1/formats/{handle}/{slug} and GET /v1/formats/{handle}/{slug}/runs take the same address.

Request body

Every field is optional on its own, but the body must name at least one of instruction, input, previous_run_id or attachments. {} or {"input": {}} is 400 invalid_request, not a run on the Format's default. Unknown top-level fields are 400 unknown_parameter, with a suggestion when the name is close (webook_url → webhook_url).

FieldNotes
instructionThe task in your words, up to 8000 characters. If you omit it, the run uses the Format's own default instruction. Sume puts it after the Format body, so it wins where the two do not agree. Refer to What is accepted and what is carried.
inputA JSON object of caller data: at most 64 top-level keys and 2 MiB (2097152 UTF-8 bytes, compact). You choose the shape. The Format reads the keys that it knows. Media URLs in it share the run's attachment budget. Refer to input caller data.
attachmentsUp to 30 images that the agent can see: { "type": "input_image", "image_url": … } or { …, "asset_id": … }. Refer to Attachments.
output_schemaIf you bind a JSON Schema, output comes back in that shape. response_format is the OpenAI-shaped alias. If you send both, you get 400 invalid_request. For rules and failure modes, refer to Structured output.
primary_output_keyThe key in output that gives the URL for primary_output_url. Up to 64 characters.
generation_spend_cap_usdThe generation ceiling of this run, up to the platform maximum of $500. If you omit it, the run inherits the Format's cap. Sume accepts a number above the Format's cap and does not clamp it. With null, the run uses $500. Sume rejects 0. Refer to Spend caps.
communication.webhook_urlPublic HTTPS URL that receives one signed format.run.terminal POST when the run completes or fails. callback_url is an accepted alias. Sume normalizes top-level webhook_url / callback_url / mode into communication. Sume delivers on both hosts. Refer to Webhook.
communication.modeasync (default) or webhook. Descriptive only. The URL arms delivery, not this field.
previous_run_idContinue an earlier run of this Format as one more turn of the same conversation, not as a fresh start. Refer to Continue a run.
on_active_runWhat to do when a run of this Format is already in flight. Default allow (runs concurrently, and workspace generation concurrency still applies). skip records a skipped run. reject answers 409 format_run_in_progress. Scheduled Actions default to skip, so do not copy their bodies here.
modelAgents catalog id for the LLM that orchestrates the run. If you omit it, the run uses the gpt-6-sol default. A request for retired gpt-5.6-sol runs on gpt-6-sol. This field selects only the orchestrator. The Format's tools select the image, video and audio models. An id outside the catalog gets 400 invalid_request. The receipt shows the id that ran.
idempotency_keyBody form of the Idempotency-Key header. If you send both, the header wins.

Headers: Authorization: Bearer $SUME_API_KEY or x-api-key: $SUME_API_KEY (one, never both), Content-Type: application/json, and Idempotency-Key on every create. The maximum request body is 4 MiB (413 payload_too_large).

caller data

input is the JSON object that your service gives to the run. It is not a wire schema, and Sume publishes no field list for it. You choose the shape, and the Format's recipe reads the keys that it recognizes. Two integrations that call the same Format can send fully different objects, and both are correct. The example bodies on these pages are one integrator's convenient shape, not a contract.

The API checks only these items:

CheckRuleOn failure
TypeA JSON object. The API refuses arrays, strings and numbers. null and omission both mean no input.400
Property countAt most 64 top-level keys. The API does not count nested keys, so groups of keys are free.400
SizeAt most 2097152 UTF-8 bytes (2 MiB) on the compact serialization.400
Media referencesHTTPS URLs to image, video or audio files, at any depth, share the run's attachment budget: 30 in total, at most 30 images, 10 videos, 10 audio. Refer to Media referenced from input.400 invalid_attachment

Do not confuse the two JSON fields. input is loose data that goes in, and output_schema is a strict contract for the data that comes out. Sume writes input whole to a file in the run's workspace. Sume tells the agent that this file is caller-supplied data, not instructions. Thus, the file is the correct location for scraped product copy, a customer's message or a supplier's field. Do not concatenate that data into instruction.

The file is a trust boundary, not a sandbox. Runs are spend-capped, so the cap sets a limit on the blast radius of a hostile payload. But do not pass raw untrusted text through on purpose.

input does not reach the structured output. Sume makes output from what the run made and said. Thus, a value that you sent, for example an order id or a SKU, cannot come back in the output unless the run repeats it. Keep your identifiers on your side, keyed by data.id or by your Idempotency-Key. Refer to Where your object comes from.

What is accepted and what is carried

FieldAcceptedCarried to the run
instruction8000 charactersThe first ~4000 characters, as prompt text. Keep it well below that limit. Put data in input.
input2 MiBAll of it, as a file that the agent reads. Sume never truncates it.
The Format bodyNo cap other than 100 MiB per package fileAll of it, attached as files.

An empty input ({}) adds no file and no block. The result is byte-identical to a request without the field. Then a Format that says "read product_url from the input" has nothing to read.

Idempotency

Send Idempotency-Key on every create. Derive it from the item that the run makes: your order id and a version that you bump only when you want a re-run. Do not derive it from the time of the request. If you use a new uuidgen per request, the header has no effect.

ReplayResult
Same key, same body200 with the original receipt and idempotency_hit: true. Sume does not start a second run or make a second charge.
Same key, different body (a different instruction or attachment list is also a different body)409 idempotency_conflict. Nothing runs.
Same key, two requests at the same momentOne request wins. The other gets 409 idempotency_key_in_use, which is retryable. Wait approximately one second, then send again to get the original run.
Same key after a create that failed (402, 503, …)Sume released the key. Correct the cause and retry with the same key.

The scope of a key is one Format. If you send the same key to two Formats, you start two runs. A key is up to 255 characters.

Spend caps

Every Format has a generation spend cap. A run can never spend more than its own effective cap. Read the Format's cap from generation_spend_cap_usd_micros on GET /v1/formats/…. A Format that never named a cap reports the platform default of $400.

generation_spend_cap_usd on the request names this run's own ceiling:

You sendThe run's cap
NothingThe Format's cap.
A number up to 500That number. Sume accepts a number above the Format's own cap and does not clamp it.
nullThe platform maximum, $500. It lifts the ceiling, but it does not remove it.
0, or above 500400. A run that cannot spend cannot deliver.

Every receipt gives the effective cap as usage.generation_spend_cap_usd_micros. It gives the actual spend of the run against the cap as usage.billable_amount_usd_micros. Production live-commerce integrations run with caps of approximately $120. A single-scene retry on the same thread needs a fraction of that. Sume meters the spend against the cap at the rates on the API pricing page. Errors and spend covers what happens at the wallet.

Keys and scopes

Scopes

ScopeNeeded for
formats:readList and read Formats, read and list runs, read queues.
formats:writeCreate a run, create a bulk queue, cancel a run, redeliver a webhook.

Sume fixes the scopes when it mints a key. Keys created before the release of the Formats API do not carry these scopes. A key without one of them fails every Format request with 403 insufficient_scope, never a 404. Create a new key at API keys and rotate to it.

Service-account keys cannot create Format runs. They fail with 403 insufficient_scope and details.reason of service_account_format_runs_unsupported.

Team Formats need a team key

To invoke a Format that a team workspace owns, use an API key created in that workspace. Membership is not sufficient. If a team member uses a personal key, the API refuses it with 403 workspace_key_required, and details.workspace_id names the workspace that the key must come from.

The rule agrees with the flow of money. A team Format's runs bill the team wallet, count against the team's generation concurrency, and read their media back through the team workspace. A personal key can split those items. In the past, a personal key produced runs that made a real video and then reported output_schema_unsatisfied with nothing harvested. Create the key from the team's dashboard. Personal keys stay correct for personal Formats.

Reads work the same way, by key. A team key lists the Formats of that workspace for every member, and never your personal Formats. A team handle that you are not a member of gets 404. You cannot tell this answer apart from a handle that does not exist. Thus, a 403 workspace_key_required always means "right team, wrong key".

Running a Format another workspace shared with you

An owner can share a team Format with a different workspace, but never with a user. The method is the same as when a GitHub repository adds an outside collaborator. The owner workspace adds your team handle on the Format's Access tab (live immediately). Or, the owner invites your workspace with POST /v1/formats/{handle}/{slug}/grants. Then your admin accepts with POST /v1/format-grants/{grant_id}/accept and a key created in your workspace. After that, you call the Format at the owner's address with your own team key:

The request has no workspace field. The key is the actor. The run, its spend, its concurrency slot and its media belong to your workspace, not to the owner. GET …/runs with your key lists only your runs. The owner still decides if the Format accepts any API calls. Its Status and API call trigger apply to every caller.

The owner can also remove your access. After that, the address is a 404 for you again. The API refuses a personal key in the same way as on any team Format. Membership of the owner workspace does not replace a grant. If you have seats in both workspaces, the key that you use decides. Each workspace keeps its own run history and its own bill, and a key from one workspace never reads the runs of the other.

Runs over the API are unattended

A Format written for chat can pause and wait for a person. In the Agents UI, "approve these stills before I make the video" is a quality gate by design. Over the API, nobody is there. Thus, Sume tells the run that those approvals are already granted. The run then continues to the paid step in its spend cap.

A run that really cannot finish comes back failed, never a half-finished completed:

One qualification: a completed run always did real work and always fills artifacts[]. But it can still carry output_error when the projection did not match your output_schema. Examine output_error before you read output. Refer to When output cannot be produced.

401 vs 403 vs 404

The HTTP class is the first branch. error.code is the second. A missing key gets 401 unauthorized. A known key without formats:read / formats:write gets 403 insufficient_scope, never 404 format_not_found. Scopes cannot be patched onto an existing key. Mint a new key and rotate.

HTTPerror.codeWhen
401unauthorizedNo key, malformed key, two credentials at the same time, revoked, or unknown. next_action is authenticate.
403insufficient_scopeValid key without formats:read / formats:write. details.required_scope names the scope. Also a service-account key on a run create or package write, with details.reason set. next_action is authenticate.
403workspace_key_requiredYou are a member of the team workspace, but you used a personal key. details.workspace_id names the workspace to mint a key in. For a team key from a different workspace, the grant decides. The run starts when that workspace has an accepted grant. Otherwise, the answer is a 404.
404format_not_foundUnknown, archived, outside this key's workspace, a team handle that you are not a member of, or a shared Format with a grant that is still pending or that the owner removed. A member's team key on the correct handle never gets this 404.
404format_run_not_foundUnknown run id, or a run that a different owner has.
404format_run_queue_not_foundUnknown bulk queue, or the queue of a different owner.
404previous_run_not_foundprevious_run_id is unknown or not yours.
404format_content_not_foundThe Format exists for this key but that package path does not.

Status classes obey RFC 9110. insufficient_scope is the RFC 6750 vocabulary. The 404 on a Format that someone else owns is a tenancy hide by design. It does not show that the Format exists at a different location (design record: #2393).

Errors

Errors and spend gives all the answers of the create call in one table. These are the errors that you will see first:

Statuserror.codeWhat to do
400invalid_requestThe body named none of instruction / input / previous_run_id / attachments, sent both output_schema and response_format, or failed a size cap. Read message.
400output_schema_invalidYour schema is outside the supported subset. details.violations[] names every problem.
401unauthorizedCorrect the header, not the body.
402insufficient_creditsThe workspace cannot fund the run. next_action is add_funds.
403insufficient_scope, workspace_key_requiredMint the correct key. Refer to the text above.
404format_not_foundExamine the address and the key that you use.
409format_inactive, format_api_trigger_disabledThe owner workspace sets these on the Format page's API tab. They apply to every caller, and also to the workspaces that the owner shared the Format with.
409idempotency_conflict, idempotency_key_in_useUnstable key derivation, or a concurrent duplicate. Refer to Idempotency.
429rate_limitedWait retry-after seconds.
503studio_agent_upstream_unavailableA Sume-side outage. Retry with the same Idempotency-Key.

error.code is a lowercase token that you can switch on. message is for humans and can change.

Rate limits

Every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and a 429 adds retry-after. Reads and writes have separate budgets. The read budget is forty times the write budget, so a poll loop cannot starve your own creates. A 429 names the budget that it came from in error.details.scope. For the full table, refer to Errors and spend.

Next