Calling a Format

Start a Format run from your own service. This page covers the invoke contract only; see Runs and results for the receipt, Structured output for schemas, and Embed a Format in your product for the whole partner integration — key custody, idempotency, webhooks, and artifact handling end to end.

Exact request and response schemas come from live OpenAPI (https://api.sume.com/reference/json). The tables here are a readable summary, not a second schema.

Prerequisites

Every run request needs all three:

  1. The Format's status is active.
  2. The Format's api_trigger_enabled is true.
  3. Your API key carries formats:read and formats:write.
  4. For a Format owned by a team workspace, your key was issued in that workspace.

A Format you have never run reports inactive / false on the first two, and still runs. Both fields are read off a hidden runner that is provisioned lazily, so GET /v1/formats/{id} shows the pre-provisioning state until the first POST .../runs creates it. Those two gates reject a run only once a runner exists and has been turned off. Do not block your integration on polling them true — call the Format.

Scopes

ScopeNeeded for
formats:readList and read Formats, read and list runs.
formats:writeCreate a run, create a bulk-run queue, cancel a run.

Keys created before the Format API-call trigger shipped do not carry these scopes. An older key fails every run request with 403 insufficient_scope, and scopes cannot be added to an existing key. Create a new key at API Keys and rotate to it — see Authentication.

Service-account keys cannot create Format runs. They fail with 403 insufficient_scope and details.reason of service_account_format_runs_unsupported.

Team Formats need a team key

A Format owned by a team workspace is invoked with an API key issued in that workspace. Membership is not enough: a personal key held by a team member is refused with 403 workspace_key_required, and details.workspace_id names the workspace the key has to come from.

The rule follows the money. A team Format's runs are billed to the team wallet, counted against the team's generation concurrency, and read back through the team's workspace — including the harvest that turns generated media into structured output. A personal key would split those: the spend would land on your personal wallet while the run belonged to the team, and the run's own media would be filed somewhere its Format could not see it. Runs like that used to produce a real video and then report output_schema_unsatisfied with zero videos harvested.

Create the key from the team's dashboard rather than your own. Personal keys stay correct for your own personal Formats, which are unaffected.

A team handle you are not a member of is 404, indistinguishable from a handle that does not exist — so a 403 here always means "right team, wrong key".

Reading a Format is keyed the same way as calling one. GET /v1/formats, GET /v1/formats/{format_id}, and GET /v1/formats/{handle}/{slug} return what your key's workspace owns, plus the first-party catalog: a personal key lists your personal Formats and never a team's, and a team key lists that workspace's Formats — for every member, not just whoever created them — and never your personal ones. Any key issued in that workspace with formats:read / formats:write may use the team vanity URL {handle}/{slug}; 404 is not "you are not the author." A Format outside your key's workspace is 404 format_not_found, the same answer as an id that does not exist. If a Format you expect is missing from the list, you are holding the other key.

On a team Format, status and api_trigger_enabled are workspace facts. A teammate's GET does not report inactive / trigger off just because that user has never invoked the Format themselves.

Invoke

Call a Format your key's workspace owns by handle and slug — the same address the Agents Format detail page shows. For a team Format that is the team handle, and any workspace key of that org with the right scopes may POST it. Fill in the required fields below to rewrite the cURL, JavaScript, TypeScript SDK, and Python snippets.

Create a Format run

POST /v1/formats/{handle}/{slug}/runs

Required

Operations resolve with { data, error, response } rather than throwing — check error before reading data. The fields below are the contract either way.

An accepted run returns 202 with a receipt:

format.id in the receipt is always the opaque skl_… id (internal SoT for runs and billing). You do not need it for day-to-day calling.

GET /v1/formats/{handle}/{slug} and GET /v1/formats/{handle}/{slug}/runs work the same way.

Bulk run queue

Leave a list of runs overnight without driving fan-out from your laptop. That contract lives on Bulk runs: POST …/bulk-runs (opaque or vanity), concurrency 1–16, up to 100 items, then poll GET /v1/format-run-queues/{queue_id}. Each item is still one Format run; the queue is not a second execution engine.

Opaque URLs (compat)

The opaque path remains valid forever for clients that already store it:

Request body, headers, scopes, idempotency, spend caps and the run receipt are identical — vanity resolves to the same Format and runs the same pipeline. Prefer {handle}/{slug} in new integrations and docs; persist invoke_url only when a rename must not break a stored URL.

PublicFormat exposes both:

FieldMeaning
handleYour current handle, when one is resolvable. sume for Formats by Sume.
slugThe Format's URL segment.
vanity_invoke_urlThe {handle}/{slug} path, or null when either half is unknown. Prefer this.
invoke_urlOpaque path. Always present, always permanent.

Formats by Sume live at the sume handle. Nobody owns them, so no account handle addresses them — chase/sume-product-promo returns 404 format_not_found. The reserved sume namespace is their address, it is the same for every caller, and you call it directly:

The run, its media and its spend belong to the key that called — the catalog Format itself stays unowned and shared. Forking is still there if you want to edit one, but it is no longer a prerequisite for calling one.

An unknown handle, an unknown slug, and a handle you do not own all return the same 404 format_not_found.

Team-owned Formats use the same vanity and opaque invoke paths. Authenticate with a key issued in that team workspace — see Team Formats need a team key.

Request body

FieldNotes
instructionThe task, in your words. Up to 8000 characters. Optional — omit it to run the Format's own default instruction. Composed after the Format body; see instruction composition.
inputJSON object of caller data, written to a file in the run's workspace and treated as data, never as instructions. Capped at 64 top-level properties and 2097152 UTF-8 bytes (2 MiB). Media URLs inside it share the run's attachment budget — see Media referenced from input. You pick the shape — see input caller data.
attachmentsUp to 30 images the agent can see, as { "type": "input_image", "image_url": … } or { …, "asset_id": … }. See Attachments.
on_active_runFormat default is allow (omit to run concurrently; workspace generation concurrency still applies). Pass skip to record a skipped run when one is already in flight, or reject for 409 format_run_in_progress. Do not copy Action examples here — Actions default to skip.
generation_spend_cap_usdPer-run cap. Clamped down to the Format's own cap — it can lower, never raise. See Spend caps.
output_schema / response_formatBind a JSON Schema so the run comes back as typed JSON. Sending both is 400 invalid_request. Full rules: Structured output.
primary_output_keyKey in output whose URL becomes primary_output_url. Up to 64 characters.
communication.webhook_urlPublic HTTPS delivery target for the terminal receipt. See Run webhooks — accepted, stored, and delivered on api.dev.sume.com and api.sume.com.
previous_run_idContinue a completed run of this Format as a second agent turn instead of starting fresh. See Continue a run.

Every field is optional individually, but the body must name at least one of instruction, input, previous_run_id, or attachments. A body that names none of them — {}, or {"input": {}} — is 400 invalid_request, not a run on the Format default: an empty POST is the cheapest mistake to make and used to be a billable one.

caller data

input is the JSON object your service hands to a run. It is not a fixed wire schema, and Sume publishes no field list for it: you choose the shape, and the Format's recipe reads the keys it recognises. Two integrations calling the same Format may send completely different objects and both are correct.

It is the exact opposite of output_schema, which is a strict contract. Keep the two apart in your head:

inputoutput_schema
DirectionYou → the runThe run → you
ShapeAny JSON object that suits your backendJSON Schema, inside a supported subset
Validated againstObject type, key count, byte size. Nothing else.Every rule in the subset
A shape Sume does not expectRuns anyway; unknown keys are simply more data400 output_schema_invalid — nothing runs, nothing is charged
Where it landsA file in the run's workspace (/workspace/inputs/sume-action-input.json)The post-run projection

What the API validates

Exactly four things, and nothing else:

CheckRuleOn failure
TypeA JSON object. An array, string, or number is refused. null and omission both mean {}.400
Property countAt most 64 top-level keys. Keys nested inside them are not counted.400
SizeAt most 2097152 UTF-8 bytes (2 MiB), measured on the compact serialization — no indentation.400
Media referencesHTTPS URLs to image, video, or audio files — anywhere in the object, at any depth — share the run's attachment budget: 30 in total, of which 30 images, 10 videos, 10 audio. See Media referenced from input.400

There is no reserved key, no required key, no type rule on values, and no naming convention. {"a": 1} and a forty-key nested order payload are equally valid. Do not treat the example input objects in these docs as a wire contract you must match field-for-field — they are one integrator's convenient shape, not a schema.

Because only top-level keys count against the 64, nesting is free. A payload with more than 64 leaves is fine as long as you group it:

What the run actually receives

input is serialized with two-space indentation and written to a file at a locked path in the run's workspace, /workspace/inputs/sume-action-input.json — whole, at every size. The prompt never carries the JSON itself; it carries a bounded pointer block that names the file, its size, and its top-level keys, and tells the agent to read it before acting:

This is the same attach-always, inline-never doctrine the Format body follows: the payload lives on disk where the agent's Read tool can open it, and the turn stays readable.

The full composition order on a Format run is:

The Format is the how and comes first. Your instruction comes after it, so where the two disagree the model follows what you asked for. Your input comes last, as data the instruction can refer to. See instruction composition.

The recipe itself reaches the agent as a file, not as prompt text: the whole package is written to /workspace/skills/{slug}/ before the turn, and [Format attached: …] names that path. This holds for every Format at every size — see How big SKILL.md should be.

The conversation the run happens in records this turn as it was sent. Open the run's thread_id in Agents and the first message is the text above, attachments and all — which is what to read first when a run did something you did not expect.

Two behaviours fall out of this that are easy to trip over:

  • An empty input adds no block and no file at all. {} — and any object whose keys were all dropped — produces a prompt with no [Sume action input] section, byte-identical to a call that omitted the field. A Format body that says read product_url from the input has nothing to read.
  • There is no separate JSON mode. Prose in instruction, structured data in input, or both, are all normal. The recipe glances at the keys it knows; everything else is context.

is data, never instructions

The pointer block's wording exists to hold a trust boundary: what arrives through sume-action-input.json is caller-supplied data, not a command to the agent. Content you did not write — a scraped product description, a customer's message, a supplier's field — belongs in input rather than concatenated into instruction, precisely so it inherits that framing.

This is a boundary, not a sandbox. Treat it the way you would treat a prompt-injection control in your own product: it is meaningfully better than splicing untrusted text into your instruction, and it is not a licence to pass a hostile payload straight through. Runs are spend-capped, so the blast radius of a bad input is bounded by the Format's cap.

input

input is never truncated. The whole object, up to the 2 MiB accept cap, is written to /workspace/inputs/sume-action-input.json and reaches the run on disk. Only the pointer block sits in the prompt, and it is small and bounded by construction — path, byte count, and the first 32 top-level key names. An earlier iteration inlined the JSON into the prompt and cut it at roughly 3860 characters; that behaviour is gone, and a live-commerce script or a long product list arrives whole.

instruction is different: it is prompt text, carried as its own [Format run instruction] block, and a block is truncated at 4000 characters, keeping the beginning.

FieldAcceptedCarried
instruction8000 charactersthe first ~4000
input2097152 UTF-8 bytes, compactall of it, as a file the agent reads

Keep instruction well inside 4000 characters, and put data in input rather than prose — the file has no such budget.

The Format body escapes this entirely, because it is never in the turn to be truncated. The turn carries the [Format attached: …] pointer and the agent opens the file, so a body of any size arrives whole rather than as a fragment cut mid-sentence — there is no cap on it here at all. If you author Formats, see How big SKILL.md should be.

does not reach the structured output

Worth knowing before you design either half. output is not produced by the agent: it is projected afterwards from the run's generated media and its closing text, and your input is not among the projection's inputs. A value you sent — an order id, a SKU, your own locale — cannot be echoed into output unless the run itself repeats it in the text it finishes with.

So do not use output_schema to route your own identifiers back to yourself. Keep them on your side, keyed by run.id or by your Idempotency-Key, and let output carry only what the run made. Full mechanics: Parsing happens after the run.

Idempotency

Send Idempotency-Key on every call. A replay with the same body returns 200 and the original run; a replay with a different body — including a different instruction — returns 409. A 200 replay sets idempotency_hit: true on the receipt, which is how you tell it from a fresh 202.

Two properties worth knowing before you pick a key:

  • Keys are scoped to one Format. The same key sent to two different Formats starts two runs and never conflicts. If you are deriving keys from an order id, include the Format in the key only when you want per-Format independence.
  • Concurrency. Two identical requests sent at the same moment are not both accepted: one wins and the other returns 409 idempotency_key_in_use, which is retryable — wait about a second and retry to receive the original run. That is a different code from idempotency_conflict, which means the payloads genuinely differ and retrying will not help.
  • A create that failed releases the key. If the call came back non-2xx — 402 with an empty wallet, say — no run was started and the key is free. Fix the cause and retry with the same key.

Spend caps

Every Format carries a generation spend cap, and a run can never spend past its own effective cap. The Format's cap is what a run inherits when it names none.

Read the current value from PublicFormat.generation_spend_cap_usd_micros. It is always a number.

A Format that never named a cap gets the platform default: $400, whatever it outputs.

Set your own with generation_spend_cap_usd when you register the Format — a finite number greater than 0 and no more than 500. Clearing it (null) returns the Format to the $400 default.

What the run spends against that cap is metered at the rates on the API pricing page.

generation_spend_cap_usd on a run request names that run's own ceiling, up to the platform maximum of $500. A number above the Format's own cap is honored rather than clamped down to it — a long production the Format was registered too small for is a budget you can raise per run. Above 500 is a 400, never a silent clamp. Omit it and the run gets the Format's cap. Pass null and the run gets the platform maximum of $500null lifts the ceiling to the highest we allow, it does not remove it.

Runs over the API are unattended

A Format's body is written for interactive chat, where it may stop and wait for a person — "approve these preview stills before I make the video" is a deliberate quality gate in the Agents UI.

Over the API there is nobody to ask. So an API run is told, in its prompt, that those approvals are already granted and that it should carry on to the paid step within the Format's spend cap. The same applies to scheduled runs. Only the interactive Agents UI still stops and waits.

A run that genuinely cannot finish comes back failed, not completed:

Treat status: "completed" as a real result. It will not be handed to you in place of a run that stopped halfway.

One qualification, and it is the only one: a completed run always did real work and always populates artifacts[], but it can still carry output_error when the structured projection did not match your output_schema. Check output_error before reading output — on that path output is null while the media is still there in artifacts[]. See When output cannot be produced.

401 vs 403 vs 404

HTTP class is the first branch; error.code is the second. A missing key is 401 unauthorized. A known key missing formats:read / formats:write is 403 insufficient_scopenever 404 format_not_found. Scopes cannot be patched onto an existing key; mint a new one at API Keys and rotate.

HTTPerror.codeWhenAuthority
401unauthorizedNo key, malformed key, two credentials at once, revoked, or unknown. next_action is authenticate.Official: RFC 9110 §15.5.2. Product SoT: OpenAPI 401.
403insufficient_scopeValid key missing formats:read / formats:write. details.required_scope names it. next_action is authenticate.Official: RFC 9110 §15.5.4; RFC 6750 insufficient_scope.
403insufficient_scopeService-account key on Format runs or package writes (details.reason is service_account_format_runs_unsupported or service_account_format_authoring_unsupported).Product SoT: same code, distinguished by details.reason. Not a new 403 code.
403workspace_key_requiredYou are a member of the team workspace but the key was not issued there. details.workspace_id names it.Product SoT. Distinct from 404 so a member knows to mint a team key.
404format_not_foundUnknown, archived, outside this key's workspace, or a team handle you are not a member of. A member's team key on the correct handle is not this 404.Product SoT / #2393. Intentional tenancy hide — do not treat this as "the Format exists elsewhere" or "you are not the author."
404format_run_not_foundUnknown run id, or a run belonging to another owner.Product SoT / #2393. Same hide.
404format_run_queue_not_foundUnknown bulk-run queue, or another owner's.Product SoT.
404previous_run_not_foundprevious_run_id is unknown or not yours.Product SoT. The Format address was valid; the continuity id was not.
404format_content_not_foundThe Format exists for this key but that package path does not.Product SoT.

Errors

CodeStatusWhat to do
unauthorized401Missing, malformed, revoked, or unknown API key. next_action is authenticate.
insufficient_scope403Your key is missing formats:read / formats:write. Keys issued before this feature do not carry them — create a new key. next_action is authenticate: editing the request body cannot fix it. Never format_not_found.
workspace_key_required403The Format belongs to a team workspace and your key is not issued in it. Use a key created in details.workspace_id. next_action is authenticate.
format_not_found404Unknown, archived, outside your key's workspace, or a team handle you are not a member of. A member's team key on the correct {handle}/{slug} is 200, not this 404.
format_not_forkable409You addressed a built-in capability rather than a Format. Call one of the Formats by Sume, or your own.
format_api_trigger_disabled409The API call trigger is off for this Format.
format_inactive409The Format is inactive. Set it active to accept API runs.
format_run_in_progress409on_active_run: "reject" and a run is already in flight.
previous_run_not_found404previous_run_id is unknown, or names a run belonging to another owner.
previous_run_format_mismatch400That run was created on a different Format. Continue it where it started.
previous_run_not_terminal409The run you want to continue has not finished. Poll it, then call again.
previous_run_not_resumable400That run did not complete — it was canceled, failed, or skipped — so it left nothing to continue. Start a fresh run.
idempotency_conflict409That Idempotency-Key was already used with a different payload.
idempotency_key_in_use409Another request with the same key is still in flight. retryable: true — retry in about a second to receive the original run.
output_schema_invalid400Your output_schema is outside the supported subset. details.violations[] names each problem — see Supported schemas.
invalid_attachment400Bad attachments[] item, or over the media budget attachments[] and input share — see Attachment errors and Media referenced from input.
attachment_not_found400asset_id is unknown in this workspace.
attachment_too_large413An image is over 30 MB, or the set is over 500 MB total.
attachment_fetch_failed502Sume could not fetch an attachment from the URL you gave. Despite the 5xx status this is your input: category is media, next_action is fix_input, and retryable is false. details.index names the attachment.
insufficient_credits402The workspace cannot fund this run. next_action is add_funds; retrying without topping up returns the same answer.
organization_wallet_not_provisioned402An organization workspace with no funded wallet. An admin has to fund it.
rate_limited429Too many requests for this key. Wait retry-after seconds. error.details.scope names which budget ran out (read or write) — see Rate limits.
studio_agent_upstream_unavailable503A Sume-side outage, not your key. Retry later; re-issuing the key will not help.
format_run_failed_to_start4xx/5xxThe run could not be started for a reason with no more specific code. message carries the detail.

error.code is always a lowercase token (^[a-z0-9_]+$) and is safe to switch on. message is written for a human and may change; never match on it.

Rate limits

Every response carries ratelimit-limit, ratelimit-remaining and ratelimit-reset, and a 429 also carries retry-after. Pace on those headers rather than on a fixed sleep.

Reads and writes have separate budgets, and the read budget is forty times the write one, so a polling loop cannot rate-limit the run creates that spawned it. A 429 names the budget it came from in error.details.scope. Prefer a run webhook over a fast poll when you can, and back off on 429 rather than treating it as fatal — the run is still executing. Full table: Rate limits.

Next