# Sume Developer Documentation > Sume provides public API and CLI surfaces for media generation, first-party media artifacts, jobs, input assets, usage, and agent workflows. This file contains the full text of the Sume developer documentation, intended for LLM consumption. ### Quick start Source: https://docs.sume.com/quick-start.md Start in the Agents tab, or call a saved Format over the Developer API. There are two entry paths. You can work with the agent in the **Agents** tab, or you can call a saved Format over the **API**. Most teams use the two paths. They author the recipe in Agents, and then invoke the recipe from their backend. If you do not know the product map, first read [Sume basics](/the-basics). It gives information on Agents, Formats, Models, and the primary surface for partners. Detailed guides: [Format API](/formats) · [Calling a Format](/formats/call) · [Agents overview](/agents) · [Models overview](/models). Agents #### Option 1: Start it from one brief Paste this text into the composer at [sume.com/agents](https://www.sume.com/agents). The thread keeps each artifact and approval. ```text Make me a 15-second vertical product promo for my product. Use a presenter avatar, add captions, and show me a draft before spending credits on the final render. Once I approve it, save this thread's recipe as a Format called product-promo so I can call it from my backend. ``` #### Option 2: Work the thread Use the Agents tab if you want a person to stay in the loop: brief, review, iterate, and then save the result as a reusable recipe. 1. Open the Agents tab Sign in. Then start a thread. No installation is necessary. ```text https://www.sume.com/agents ``` 2. Brief the agent Tell the agent what the deliverable is, not which tool calls to make. The agent selects the models, and it asks you before it spends. ```text Make a 15-second vertical product promo with a presenter avatar and captions. Show me a draft before the final render. ``` 3. Save the recipe as a Format A Format is the recipe of the same thread. You can address it by handle and slug. Your backend calls the Format later. ```text Save this as a Format called product-promo. ``` Next: [Agents overview](/agents). API #### Option 1: Start it with your agent Paste this text into an agent that can edit or run your server-side integration code. ```text Call a Sume Format from my server environment. Use SUME_API_KEY for auth and POST https://api.sume.com/v1/formats/{handle}/{slug}/runs with a JSON body of {"instruction": "..."} plus an Idempotency-Key header. The call returns 202 with a run receipt; poll status_url until the run leaves queued/processing, then read result_url and return the artifacts and any structured output. If SUME_API_KEY is missing, ask me to create one at https://www.sume.com/dashboard/api-keys with the formats:read and formats:write scopes. ``` #### Option 2: Use it manually Use the API when you want your backend to run a saved Format directly. One call does the orchestration, so you do not have to write it yourself. 1. Create your API key in the **[API Keys dashboard](https://www.sume.com/dashboard/api-keys)** with the `formats:read` and `formats:write` scopes. Then export the key: ```bash export SUME_API_KEY="sume_live_..." ``` 2. Start a Format run. Call the Format by handle and slug. This is the same address that its detail page shows. An accepted run returns `202` with a receipt: 3. Read the result. Poll `status_url` until the run status changes from `queued` / `processing`. Then read `result_url`: ```bash curl -sS https://api.sume.com/v1/format-runs/arun_.../status \ -H "Authorization: Bearer $SUME_API_KEY" ``` In production, we recommend a [webhook](/sdk/webhooks) and not a poll loop. Or let the [TypeScript SDK](/sdk/runs) wait for the result. Next: [Calling a Format](/formats/call). Do you want only one model call and not a packaged workflow? Then go directly to [Create new avatar](/models/avatar), [Image 1.0](/models/image), or the [API recipes](/api/cookbook). ### Sume basics Source: https://docs.sume.com/the-basics.md A product map of Sume’s surfaces — Agents, Formats, Models, SDK, Dashboard, and the Developer API. Sume is primarily a **video agent** platform. People author and iterate in chat. Partners invoke saved recipes over HTTP. Each generation tool (Image, Video, Avatar, TTS, timeline, and more) is also available as an HTTP API. But the primary partner path is a call to a **sandbox Agent / Format** that composes those tools. It is not a set of raw model calls that you connect yourself. Because of that composition, you can ship deliverables that a single [Video 1.0](/models/video) clip cannot give: multi-minute host footage, B-roll, voiceover, and timeline assembly into a **post-ready** video. For job admission, polls, and the mechanics of `media.sume.com`, refer to [Core concepts](/workflows/core-workflow). #### In a nutshell - **[Agents](#agents)** — a person in the loop authors with the Agent at [sume.com/agents](https://www.sume.com/agents) - **[Formats / Format API](#formats--format-api)** — the primary partner invoke surface - **[Models](#models)** — atomic generation APIs (components in a support role) - **[Agent Completions / Scheduled](#agent-completions--scheduled)** — ad-hoc and repeated agent runs - **[SDK](#sdk)** — the TypeScript client around the same Developer API - **[Dashboard](#dashboard)** — keys, jobs, usage, billing - **[Developer API + media](#developer-api--media)** — `api.sume.com` and `media.sume.com` - **[Workspaces](#workspaces)** — where keys and spend resolve #### Agents [Agents](https://www.sume.com/agents) is the chat UI where a person works with the sandbox Agent: write a brief, approve spend, examine artifacts, and shape a house style over a few turns. People author Formats here. A person can edit a Format in the library, or ask the Agent in chat to save a recipe (`SKILL.md` plus references). When a human must stay in the loop, interactive chat is the correct surface. Docs: [Agents overview](/agents) · [Safe automation](/agents/safe-automation) #### Formats / Format API A **Format** is a saved recipe to author a deliverable. Partners call it by handle and slug. Sume then starts a fresh sandbox, loads the recipe, runs the Agent with generation tools, and returns artifacts and optional structured JSON. **This is the recommended surface for most partners.** One HTTP call carries the judgment and orchestration that otherwise live in your own glue code. ```text Format run = fresh sandbox + recipe (SKILL) + instruction/input + tools → artifacts + optional structured output ``` Docs: [Format API overview](/formats) · [Calling a Format](/formats/call) · [Bulk runs](/formats/bulk-runs) · [Cookbook: embed a Format](/cookbooks/embed-a-format) #### Models **Models** are atomic generation endpoints. They create an avatar, render a talking clip, generate an image or a short video clip, add captions, and more. They are real product surfaces. Use them when you need one model invocation and nothing else. They **support** Formats. A Format decides *which* tools to call, in which order, and how to assemble the result. If you need only a single clip or image, call the model. If you need a packaged workflow, call a Format. Docs: [Models overview](/models) · [Video 1.0](/models/video) · [Avatar videos](/models/avatar-videos) #### Agent Completions / Scheduled It is not necessary to save each agent task as a Format. | Surface | When to use | |---|---| | [Agent Completions](/agents/completions) | A one-off backend task, with nothing to save as a recipe | | [Scheduled](/agents/actions) | The same saved task on a cadence (Actions) | Both run the Agent. Completions is ad-hoc, and Scheduled repeats. When the recipe is fixed and only the inputs change, Formats stay the path. #### SDK The [TypeScript SDK](/sdk) is a thin client over the same Developer API for Node / Bun / Deno / Workers. It includes [`subscribeFormatRun`](/sdk/runs), so you do not write your own poll loop. You can also do all of its operations over plain HTTP. Two more clients exist and still work, but they are not part of the primary path today. They are the [CLI](/cli) for local shells and scripts, and [hosted MCP](/mcp) for clients that speak remote MCP. Use them when your environment needs them, not as the default integration. #### Dashboard The dashboard is the surface for human operators. It shows the same workspace that the API key resolves to: - [API keys](https://www.sume.com/dashboard/api-keys) - [Jobs](https://www.sume.com/dashboard/jobs) - [Usage](https://www.sume.com/dashboard/usage) - [Billing & subscription](https://www.sume.com/dashboard/subscription) - [Playground](https://www.sume.com/playground) Docs: [API keys](/dashboard/api-keys) · [Jobs](/dashboard/jobs) · [Usage](/dashboard/usage) · [Billing](/dashboard/credits) #### Developer API + media | Domain | Role | |---|---| | `api.sume.com` | Public Developer API (`/v1`) and OpenAPI (`/reference/json`) | | `media.sume.com` | First-party generated media artifacts | Keys authenticate. The API is workspace-scoped. The API returns the generated outputs that belong to Sume as `media.sume.com` URLs. That is the public artifact contract. Docs: [Public API](/public-api) · [API reference](/api/reference) · [Authentication](/authentication) · [Media inputs](/workflows/asset-library) #### Workspaces API keys and spend resolve to a **workspace**. The key carries that context. Do not send `workspace_id` in request bodies. You can reach team-owned Formats on both vanity and opaque paths with a key **created in that team workspace**. A personal key fails with `403 workspace_key_required`. Refer to [Team Formats need a team key](/formats/call#team-formats-need-a-team-key). #### What next? - [Quick start](/) — your first run, in the Agents tab or over the API - [Format API](/formats) — why Formats exist and how a run works end to end - [Core concepts](/workflows/core-workflow) — jobs, admission, artifacts, usage ### Best practices Source: https://docs.sume.com/best-practices.md Short patterns for calling Formats and Agent Completions well — with copy-paste examples. This page gives a few patterns that keep partner integrations stable and reliable. Use the interactive examples below as a start. Replace the handle, the slug, and the instruction with the values for your account. #### Prefer a Format over raw model calls When you have a saved recipe, call the recipe by its handle and slug. The Format owns the tools, the spend gates, and the house style. Your client only sends the brief. For the full contract, refer to [Calling a Format](/formats/call). When you must get typed JSON back, refer to [Structured output](/formats/structured-output). #### Bind a schema when a system will consume the result If another service will read the receipt, bind `output_schema`. Then you get a validated object, not free text. Mark a field as required only if the Format actually produces that field. Full rules: [Structured output](/formats/structured-output). #### Use Agent Completions for one-off work If there is no saved Format or the brief changes each time, send an Agent Completion with a spend cap. #### Keep credentials scoped For Formats that are provisioned on a team handle, use a **team** API key. Select the narrowest scopes that your client needs. Rotate keys from the [Dashboard](/dashboard/api-keys). #### Cap spend on every run Always set `generation_spend_cap_usd` (and an agent cap when applicable). A missing cap is a bug in the client, not a convenience. ## Format API ### Overview Source: https://docs.sume.com/formats.md Run a saved video recipe from your backend with one HTTP call. Create a run, take the receipt by webhook or poll, and read back durable media plus JSON in a schema you supply. A Format is a saved production recipe: a house style, an output contract, and a playbook for one kind of video. Your backend calls the Format by name. Sume runs it in a fresh sandbox with the generation tools. You get finished media on `media.sume.com`. If you ask for it, you also get a JSON object in a shape that you defined. This page goes from an API key to a finished run. The pages after it give the reference for each step. #### Your first run You must have an API key that carries the `formats:read` and `formats:write` scopes. Create a key at [API keys](https://www.sume.com/dashboard/api-keys). Keep the key server-side. ##### 1. List the Formats your key can call ```bash export SUME_API_KEY="sume_live_..." curl -sS "https://api.sume.com/v1/formats" \ -H "Authorization: Bearer $SUME_API_KEY" ``` ```json { "data": [ { "id": "skl_…", "object": "format", "handle": "acme", "slug": "live-commerce", "title": "Live commerce", "status": "active", "api_trigger_enabled": true, "io": { "profile": "url_to_video", "input_kind": "url", "output_kind": "video" }, "generation_spend_cap_usd_micros": 120000000, "vanity_invoke_url": "https://api.sume.com/v1/formats/acme/live-commerce/runs", "invoke_url": "https://api.sume.com/v1/formats/skl_…/runs" } ], "has_more": false, "next_cursor": null } ``` The list shows the Formats that your key's workspace owns, and the ready-made [Formats by Sume](/formats/catalog). `vanity_invoke_url` is the address that you call next. ##### 2. Start a run ```bash curl -sS -X POST "https://api.sume.com/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: order-8823-v1" \ -d '{ "instruction": "Vertical 9:16 host video. Use the script as written. No BGM, no captions.", "input": { "product_url": "https://shop.example.com/p/8823", "host_image_url": "https://cdn.example.com/hosts/yura.png", "vo_language": "ko" }, "generation_spend_cap_usd": 120, "communication": { "webhook_url": "https://acme.example.com/hooks/sume" } }' ``` ```json { "data": { "id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "queued", "format": { "id": "skl_…", "slug": "live-commerce", "title": "Live commerce", "version": 23 }, "status_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/status", "result_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/result", "events_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/events", "cancel_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/cancel", "webhook_delivery": { "url": "https://acme.example.com/hooks/sume", "status": "not_armed" }, "usage": { "currency": "USD", "billable_amount_usd_micros": 0, "generation_spend_cap_usd_micros": 120000000 }, "thread_id": "thr_…", "next_action": "poll_status" } } ``` `202` means that Sume accepted a fresh run. Store `data.id`. The receipt gives all the other items that you need as URLs. Thus, you never build a path manually. For the full body reference, refer to [Create a run](/formats/call). ##### 3. Take the result When the run completes, Sume POSTs the same receipt to your `webhook_url` as one signed `format.run.terminal` event. To poll instead, read the run until `status` is terminal: ```bash curl -sS "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The response for a finished run is: ```json { "data": { "id": "arun_e43e6c5cb2b74052", "status": "completed", "primary_output_url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "artifacts": [ { "id": "artf_…", "type": "video", "url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "content_type": "video/mp4", "duration_ms": 73360 } ], "output": { "text": "…", "videos": [{ "type": "video", "url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "…": "…" }], "images": [], "audio": [], "files": [] }, "usage": { "currency": "USD", "billable_amount_usd_micros": 14959638, "generation_spend_cap_usd_micros": 120000000 }, "next_action": "none" } } ``` `primary_output_url` is the one item to show. `artifacts[]` is every file that the run made. Media URLs are durable and public. Store them. If you bind an [output schema](/formats/structured-output), `output` comes back in your own shape, not in the built-in shape. That is the whole loop: create, then webhook or poll, then read. Runs that make video take minutes, not seconds. Long-form host video usually completes in 15 to 30 minutes. Thus, design for the asynchronous path from the start. #### Two hosts, one contract | Host | Use it for | Keys | |---|---|---| | `https://api.sume.com` | Production. Runs spend real credits from your workspace. | Create them at [API keys](https://www.sume.com/dashboard/api-keys). | | `https://api.dev.sume.com` | Integration and staging. Same routes, same receipts, same webhook delivery. | Sume issues them for a development workspace. Get them from your Sume contact. | A key works only on the host that it was created for. The other host answers `401 unauthorized`. Keys for both hosts look like `sume_live_…`. Thus, name your environment variables by host, not by prefix. The examples on these pages use production. To point them at development, change the host and the key. #### How a run works ```text POST /v1/formats/{handle}/{slug}/runs -> 202, receipt with status_url / result_url / events_url / cancel_url -> fresh sandbox boots, the Format's package lands on disk -> the agent follows the recipe and calls generation tools (host takes, B-roll, voiceover, captions, timeline assembly …) -> media is mirrored to media.sume.com -> terminal receipt: webhook POST format.run.terminal, or your next poll -> read output, artifacts[], primary_output_url ``` Every call has four parts: | Piece | What it is | Where it lives | |---|---|---| | The recipe | The Format: a `SKILL.md` body and reference files. The *how*: house style, branch rules, quality bar. | You write it in the Agents dashboard, in chat, or over the [Contents API](/formats/contents). | | The call | Your `instruction` and `input`. The *what*: product URL, brief, script, prices. | The request body. | | The run | One fresh sandbox, one agent turn, one receipt. A run never gives a partial delivery. A run that could not finish comes back `failed`. | `arun_…`, at `/v1/format-runs/{run_id}`. | | The result | Durable media and `output`, in the built-in shape or in a shape projected onto your [schema](/formats/structured-output). | The terminal receipt. | This gives two properties. A run is one unit of work: a bulk request is a server-side queue of ordinary runs, not a different engine ([Bulk runs](/formats/bulk-runs)). Also, the recipe is set before your instruction. Thus, you do not send a system prompt again on every call with no guarantee that it holds. ##### This is not chat A Format run is one unattended turn with the Format attached. The run does not stop to ask a person questions. Approvals that a chat-authored recipe asks for in chat are pre-granted, and the run continues in its spend cap. The Agents chat UI at [sume.com/agents](https://www.sume.com/agents) is the surface for a human in the loop. A chat turn can have no Format attached. Build partner integrations on runs, not on chat threads. #### Find your Formats Three reads are available. Each read must have `formats:read`: | Call | Returns | |---|---| | `GET /v1/formats?limit=50` | The Formats that your key's workspace owns, and the first-party catalog. Keyset pages: while `has_more` is `true`, send `next_cursor` back as `cursor`. | | `GET /v1/formats/{handle}/{slug}` | One Format by its address. | | `GET /v1/formats/{format_id}` | The same Format by its opaque `skl_…` id. | The key sets the visibility. A personal key lists your personal Formats. A team key lists the Formats of that workspace, for every member. Neither key lists the Formats of the other key. A Format outside your key's workspace gets `404 format_not_found`, the same answer as an id that does not exist. If a Format that you expect is missing, you have the other key. These fields are important when you select a Format: | Field | Notes | |---|---| | `handle`, `slug`, `vanity_invoke_url` | The address to call. For a team Format, `handle` is the handle of the workspace that owns it. For a personal Format, it is your own handle. For the catalog, it is `sume`. | | `invoke_url` | The opaque `skl_…` path. It stays the same after a rename. If a stored URL must stay valid after a handle or slug change, persist this path. Renamed handles continue to resolve for 90 days. | | `status`, `api_trigger_enabled` | Both must allow API runs. `inactive` or `false` refuses a create with `409`. A Format that you never ran over the API can show `inactive` / `false` until its first run, and it still runs. Do not gate your integration on a poll that shows them as true. | | `io` | What the Format takes and makes. `input_kind` is `url`, `text`, `image` or `product`. `output_kind` is `video`, `image` or `text`. It is `null` on Formats saved before this field existed. | | `showcase` | A real output that the Format produced at registration, or `null`. | | `generation_spend_cap_usd_micros` | The cap that a run inherits when it names no cap of its own. It is $400 for a Format that never set a cap. | | `version` | Sume bumps it on every edit. The receipt's `format.version` shows which version ran. | | `package_sha`, `contents_url` | The package behind the Format, for the [Contents API](/formats/contents). | By design, the recipe body is not in this shape. The body goes to the agent, not to the caller. Formats by Sume answer at the reserved `sume` handle, `POST /v1/formats/sume/{slug}/runs`, with any key that carries the scopes. The run, its media and its spend belong to the key that made the call. Refer to the [Format catalog](/formats/catalog). Every Format with an address also has a call sheet on this site at `https://docs.sume.com/formats/{handle}/{slug}`. The call sheet is a share link with the curl, scopes and poll loop for that Format. It shows nothing from the Format body. #### Instruction composition The agent gets these parts, in this order: ```text [Format: live-commerce v23] <- a pointer at the recipe; the body is never inlined [Format attached: … SKILL.md] <- the whole package, on disk in the run's workspace [Format run instruction] <- your `instruction`, or the Format's default [Sume unattended run] <- API and scheduled runs only [Sume action input] <- a pointer at your `input`, written whole to a file [Attached files] <- your `attachments`, when present ``` The Format comes first because it is the *how*, and your instruction comes after it. Thus, where the two do not agree, the model does what you asked for. Sume writes your `input` to `/workspace/inputs/sume-action-input.json`, whole at any size up to the cap. Sume tells the agent to read it as data, never as instructions. If you open the run's `thread_id` in Agents, the first message is this text, with no changes. When a run did something that you did not expect, read this message first. ##### How big `SKILL.md` should be The body has no size limit other than the 100 MiB per file and per package that applies to every package file. Size changes how reliably the agent *follows* a recipe, not if the recipe runs. Keep `SKILL.md` a short index that the agent can hold at one time. Put detail into `references/*`. These files sit adjacent to the body and cost nothing until the agent opens them. The [Contents API](/formats/contents) covers authoring. #### Attachments A run can carry up to 30 images that the agent can look at, as `attachments[]` on the create body: ```json { "instruction": "Make a product hero from the attached photo.", "input": { "brand": "Acme" }, "attachments": [ { "type": "input_image", "image_url": "https://cdn.example.com/shot.jpg" }, { "type": "input_image", "asset_id": "asset_…", "filename": "packshot.png" } ] } ``` | Field | Required | Notes | |---|---|---| | `type` | yes | `input_image`, the only type at this time. | | `image_url` | one of | Public HTTPS URL. Sume fetches it when you create the run. Thus, Sume must be able to get it without auth. | | `asset_id` | one of | An image that you uploaded through the [Assets API](/api/reference#media-inputs). The image must be ready and in the same workspace. | | `filename` | no | The label that the agent sees. The default is the URL's basename. | Sume fetches every attachment at create time. Sume examines its real type and size and copies it into durable storage. Thus, a broken or private image fails the create with a `4xx`/`5xx` that you can act on. The image does not stop the run minutes later. Sume does not copy an `asset_id` or a URL that is already on `media.sume.com` again. | Limit | Value | |---|---| | Types | JPEG, PNG, WebP, GIF, AVIF | | Images per run | 30 | | Bytes per image | 30 MB | | Bytes per run | 500 MB | ##### Media referenced from `input` Many Formats take their references through `input` fields, not attachments: `host_image_url`, `product_image_urls[]`, `input_reference_image_urls[]`, narration URLs. These references also count against one budget that they share with `attachments[]`. The budget is 30 files per run in total, with a maximum of 30 images, 10 videos and 10 audio files. Sume checks by file type, not by field name. An HTTPS URL at any location or depth in `input` counts if its filename ends in a media extension. A product page URL does not count. The same URL counts one time, also if it occurs more than one time. Media that the agent finds itself during the run is not yours and does not count. If you go over any of these limits, the create fails with `400 invalid_attachment`. ##### Attachment errors | Status | Code | Cause | |---|---|---| | `400` | `invalid_attachment` | Wrong `type`, missing or non-HTTPS URL, both `image_url` and `asset_id`, too many items, or a source that is not a permitted image type. | | `400` | `attachment_not_found` | `asset_id` is unknown in this workspace. | | `413` | `attachment_too_large` | An image is over 30 MB, or the set is over 500 MB. | | `502` | `attachment_fetch_failed` | Sume could not fetch the image. Causes: unreachable host, hotlink protection, or a non-2xx answer. `details.index` names the attachment. | `Idempotency-Key` covers attachments. If you replay a key with a different image list, you get `409 idempotency_conflict`. A true replay does not fetch your images again. #### What the API does not do - **No push channel for progress.** The API has no SSE or WebSocket stream. `events_url` is a polled phase timeline (`preparing`, `running`, `finalizing`), not agent output or logs. Sume pushes the completion through the webhook. - **No list of all runs.** You list runs per Format (`GET /v1/formats/{handle}/{slug}/runs`) and read them one at a time at `/v1/format-runs/{run_id}`. There is no `GET /v1/format-runs`. - **Team Formats need a team key.** You can call a Format that a team workspace owns only with a key created in that workspace. This applies to both URL shapes. A personal key fails with `403 workspace_key_required`. Refer to [Create a run](/formats/call#team-formats-need-a-team-key). - **Images only as attachments.** `input_image` is the only attachment type. Send documents by URL in `input`. Send video or audio references in the same way. - **Authoring is a separate surface.** Create and edit the package over the [Contents API](/formats/contents), or in the dashboard. The run endpoints only execute. #### Next - [Create a run](/formats/call): the request body, idempotency, spend caps, keys, and every create error - [Runs and results](/formats/runs): the receipt, polls, webhooks, and how to continue and cancel a run - [Structured output](/formats/structured-output): bind a schema and get typed JSON back - [Errors and spend](/formats/errors): every code in one place, credits, rate limits - [Cookbook](/formats/cookbook): copy-paste recipes for a real-shaped run, a webhook receiver, a scene retry, a batch - [Bulk runs](/formats/bulk-runs): queue up to 100 runs with a concurrency window - [Format catalog](/formats/catalog): ready-made Formats by Sume - [Embed a Format in your product](/cookbooks/embed-a-format): key custody, spend tiers and artifacts for a multi-tenant product ### Create a run Source: https://docs.sume.com/formats/call.md POST /v1/formats/{handle}/{slug}/runs, field by field. The request body, idempotency, spend caps, keys and scopes, and every error the create call returns. One `POST` starts a run. Production integrations send this shape: your data in `input`, a schema for the result, a per-run spend cap, and a webhook. With the webhook, you do not have to poll. ```bash curl -sS -X POST "https://api.sume.com/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: order-8823-lc-v1" \ -d '{ "instruction": "Use the Intro/Mid/Fin script as written. Korean host, vertical 9:16, no BGM, no captions.", "input": { "product_url": "https://shop.example.com/p/8823", "product_name": "Aurora Headphones", "host_image_url": "https://cdn.example.com/hosts/yura.png", "vo_language": "ko", "on_card_name": "Aurora Headphones", "price": { "list": "31,000원", "sale": "22,940원", "discount_label": "26%" }, "script": { "segments": [ { "tag": "Intro", "text": "안녕하세요, …" }, { "tag": "Mid", "text": "…" }, { "tag": "Fin", "text": "…" } ] } }, "output_schema": { "name": "acme/live-commerce/v1", "strict": true, "schema": { "type": "object", "additionalProperties": false, "required": ["full_video"], "properties": { "full_video": { "$ref": "SumeMediaFile#" } } } }, "primary_output_key": "full_video", "generation_spend_cap_usd": 120, "communication": { "webhook_url": "https://acme.example.com/hooks/sume" } }' ``` `202 Accepted`: ```json { "data": { "id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "queued", "format": { "id": "skl_…", "slug": "live-commerce", "title": "Live commerce", "version": 23 }, "trigger": { "source": "api", "idempotency_key": "order-8823-lc-v1" }, "output_schema": { "name": "acme/live-commerce/v1", "strict": true, "source": "request_override" }, "usage": { "currency": "USD", "billable_amount_usd_micros": 0, "generation_spend_cap_usd_micros": 120000000 }, "webhook_delivery": { "url": "https://acme.example.com/hooks/sume", "status": "not_armed", "attempts": 0, "max_attempts": 10 }, "status_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/status", "result_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/result", "events_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/events", "cancel_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/cancel", "expires_at": "2026-08-24T00:53:39.094Z", "thread_id": "thr_9af848b9-…", "previous_run_id": null, "idempotency_hit": false, "next_action": "poll_status", "request_id": "req_6f560f0f004f4025be47d3d317b6ecf7" } } ``` Two status codes mean success. `202` is a fresh run. `200` is an idempotent replay (the same `Idempotency-Key` with the same body), and returns the original run with `idempotency_hit: true`. Both carry the full receipt. Store `data.id` and use the URLs on it. [Runs and results](/formats/runs) covers the next steps. #### Build your own call Fill in the fields to rewrite the cURL, TypeScript, JavaScript and Python snippets. Operations in the TypeScript SDK resolve with `{ data, error, response }`. They do not throw. Thus, examine `error` before you read `data`. #### Address a Format | Shape | Example | Use it when | |---|---|---| | `{handle}/{slug}` | `POST /v1/formats/acme/live-commerce/runs` | Every new integration. It is the address that the Format detail page shows, and the shape that production callers use. | | `{format_id}` | `POST /v1/formats/skl_…/runs` | A stored URL that must stay valid after a handle or slug rename. Permanent, with the same behavior. | | `sume/{slug}` | `POST /v1/formats/sume/sume-product-commercial/runs` | The [Formats by Sume](/formats/catalog) catalog. Any key with the scopes can call it. The run belongs to that key. | The two shapes resolve to the same Format and run the same pipeline: same body, headers, idempotency, caps and receipt. The receipt's `format.id` is always the opaque `skl_…` id, for each shape that you use. A renamed handle continues to resolve for 90 days. A team Format is at the **team's** handle. Any key created in that workspace with the correct scopes can `POST` it. Refer to [Team Formats need a team key](#team-formats-need-a-team-key). An unknown handle, an unknown slug, and a handle that you cannot see all answer the same `404 format_not_found`. `GET /v1/formats/{handle}/{slug}` and `GET /v1/formats/{handle}/{slug}/runs` take the same address. #### Request body Every field is optional on its own, but the body must name at least one of `instruction`, `input`, `previous_run_id` or `attachments`. `{}` or `{"input": {}}` is `400 invalid_request`, not a run on the Format's default. Unknown top-level fields are `400 unknown_parameter`, with a suggestion when the name is close (`webook_url` → `webhook_url`). | Field | Notes | |---|---| | `instruction` | The task in your words, up to 8000 characters. If you omit it, the run uses the Format's own default instruction. Sume puts it after the Format body, so it wins where the two do not agree. Refer to [What is accepted and what is carried](#what-is-accepted-and-what-is-carried). | | `input` | A JSON object of caller data: at most 64 top-level keys and 2 MiB (2097152 UTF-8 bytes, compact). You choose the shape. The Format reads the keys that it knows. Media URLs in it share the run's attachment budget. Refer to [`input` caller data](#input-caller-data). | | `attachments` | Up to 30 images that the agent can see: `{ "type": "input_image", "image_url": … }` or `{ …, "asset_id": … }`. Refer to [Attachments](/formats#attachments). | | `output_schema` | If you bind a JSON Schema, `output` comes back in that shape. `response_format` is the OpenAI-shaped alias. If you send both, you get `400 invalid_request`. For rules and failure modes, refer to [Structured output](/formats/structured-output). | | `primary_output_key` | The key in `output` that gives the URL for `primary_output_url`. Up to 64 characters. | | `generation_spend_cap_usd` | The generation ceiling of this run, up to the platform maximum of $500. If you omit it, the run inherits the Format's cap. Sume accepts a number above the Format's cap and does not clamp it. With `null`, the run uses $500. Sume rejects `0`. Refer to [Spend caps](#spend-caps). | | `communication.webhook_url` | Public HTTPS URL that receives one signed `format.run.terminal` POST when the run completes or fails. `callback_url` is an accepted alias. Sume normalizes top-level `webhook_url` / `callback_url` / `mode` into `communication`. Sume delivers on both hosts. Refer to [Webhook](/formats/runs#webhook). | | `communication.mode` | `async` (default) or `webhook`. Descriptive only. The URL arms delivery, not this field. | | `previous_run_id` | Continue an earlier run of this Format as one more turn of the same conversation, not as a fresh start. Refer to [Continue a run](/formats/runs#continue-a-run). | | `on_active_run` | What to do when a run of this Format is already in flight. Default `allow` (runs concurrently, and workspace generation concurrency still applies). `skip` records a `skipped` run. `reject` answers `409 format_run_in_progress`. Scheduled Actions default to `skip`, so do not copy their bodies here. | | `model` | Agents catalog id for the LLM that orchestrates the run. If you omit it, the run uses the `gpt-6-sol` default. A request for retired `gpt-5.6-sol` runs on `gpt-6-sol`. This field selects only the orchestrator. The Format's tools select the image, video and audio models. An id outside the catalog gets `400 invalid_request`. The receipt shows the id that ran. | | `idempotency_key` | Body form of the `Idempotency-Key` header. If you send both, the header wins. | Headers: `Authorization: Bearer $SUME_API_KEY` **or** `x-api-key: $SUME_API_KEY` (one, never both), `Content-Type: application/json`, and `Idempotency-Key` on every create. The maximum request body is 4 MiB (`413 payload_too_large`). #### `input` caller data `input` is the JSON object that your service gives to the run. It is not a wire schema, and Sume publishes no field list for it. You choose the shape, and the Format's recipe reads the keys that it recognizes. Two integrations that call the same Format can send fully different objects, and both are correct. The example bodies on these pages are one integrator's convenient shape, not a contract. The API checks only these items: | Check | Rule | On failure | |---|---|---| | Type | A JSON object. The API refuses arrays, strings and numbers. `null` and omission both mean no input. | `400` | | Property count | At most 64 top-level keys. The API does not count nested keys, so groups of keys are free. | `400` | | Size | At most 2097152 UTF-8 bytes (2 MiB) on the compact serialization. | `400` | | Media references | HTTPS URLs to image, video or audio files, at any depth, share the run's attachment budget: 30 in total, at most 30 images, 10 videos, 10 audio. Refer to [Media referenced from `input`](/formats#media-referenced-from-input). | `400 invalid_attachment` | Do not confuse the two JSON fields. `input` is loose data that goes in, and `output_schema` is a strict contract for the data that comes out. Sume writes `input` whole to a file in the run's workspace. Sume tells the agent that this file is caller-supplied data, not instructions. Thus, the file is the correct location for scraped product copy, a customer's message or a supplier's field. Do not concatenate that data into `instruction`. The file is a trust boundary, not a sandbox. Runs are spend-capped, so the cap sets a limit on the blast radius of a hostile payload. But do not pass raw untrusted text through on purpose. `input` does not reach the structured output. Sume makes `output` from what the run made and said. Thus, a value that you sent, for example an order id or a SKU, cannot come back in the output unless the run repeats it. Keep your identifiers on your side, keyed by `data.id` or by your `Idempotency-Key`. Refer to [Where your object comes from](/formats/structured-output#where-your-object-comes-from). ##### What is accepted and what is carried | Field | Accepted | Carried to the run | |---|---|---| | `instruction` | 8000 characters | The first ~4000 characters, as prompt text. Keep it well below that limit. Put data in `input`. | | `input` | 2 MiB | All of it, as a file that the agent reads. Sume never truncates it. | | The Format body | No cap other than 100 MiB per package file | All of it, attached as files. | An empty `input` (`{}`) adds no file and no block. The result is byte-identical to a request without the field. Then a Format that says "read `product_url` from the input" has nothing to read. #### Idempotency Send `Idempotency-Key` on every create. Derive it from the item that the run makes: your order id and a version that you bump only when you want a re-run. Do not derive it from the time of the request. If you use a new `uuidgen` per request, the header has no effect. | Replay | Result | |---|---| | Same key, same body | `200` with the original receipt and `idempotency_hit: true`. Sume does not start a second run or make a second charge. | | Same key, different body (a different `instruction` or attachment list is also a different body) | `409 idempotency_conflict`. Nothing runs. | | Same key, two requests at the same moment | One request wins. The other gets `409 idempotency_key_in_use`, which is retryable. Wait approximately one second, then send again to get the original run. | | Same key after a create that failed (`402`, `503`, …) | Sume released the key. Correct the cause and retry with the same key. | The scope of a key is one Format. If you send the same key to two Formats, you start two runs. A key is up to 255 characters. #### Spend caps Every Format has a generation spend cap. A run can never spend more than its own effective cap. Read the Format's cap from `generation_spend_cap_usd_micros` on [`GET /v1/formats/…`](/formats#find-your-formats). A Format that never named a cap reports the platform default of $400. `generation_spend_cap_usd` on the request names this run's own ceiling: | You send | The run's cap | |---|---| | Nothing | The Format's cap. | | A number up to 500 | That number. Sume accepts a number above the Format's own cap and does not clamp it. | | `null` | The platform maximum, $500. It lifts the ceiling, but it does not remove it. | | `0`, or above 500 | `400`. A run that cannot spend cannot deliver. | Every receipt gives the effective cap as `usage.generation_spend_cap_usd_micros`. It gives the actual spend of the run against the cap as `usage.billable_amount_usd_micros`. Production live-commerce integrations run with caps of approximately $120. A single-scene retry on the same thread needs a fraction of that. Sume meters the spend against the cap at the rates on the [API pricing page](https://www.sume.com/pricing/api). [Errors and spend](/formats/errors#credits-and-spend) covers what happens at the wallet. #### Keys and scopes ##### Scopes | Scope | Needed for | |---|---| | `formats:read` | List and read Formats, read and list runs, read queues. | | `formats:write` | Create a run, create a bulk queue, cancel a run, redeliver a webhook. | Sume fixes the scopes when it mints a key. Keys created before the release of the Formats API do not carry these scopes. A key without one of them fails every Format request with `403 insufficient_scope`, never a `404`. Create a new key at [API keys](https://www.sume.com/dashboard/api-keys) and rotate to it. Service-account keys cannot create Format runs. They fail with `403 insufficient_scope` and `details.reason` of `service_account_format_runs_unsupported`. ##### Team Formats need a team key To invoke a Format that a team workspace owns, use an API key created in that workspace. Membership is not sufficient. If a team member uses a personal key, the API refuses it with `403 workspace_key_required`, and `details.workspace_id` names the workspace that the key must come from. ```json { "error": { "code": "workspace_key_required", "message": "This Format belongs to a team workspace. Create an API key in that workspace and use it instead of a personal key.", "details": { "workspace_id": "org_…" } } } ``` The rule agrees with the flow of money. A team Format's runs bill the team wallet, count against the team's generation concurrency, and read their media back through the team workspace. A personal key can split those items. In the past, a personal key produced runs that made a real video and then reported `output_schema_unsatisfied` with nothing harvested. Create the key from the team's dashboard. Personal keys stay correct for personal Formats. Reads work the same way, by key. A team key lists the Formats of that workspace for every member, and never your personal Formats. A team handle that you are not a member of gets `404`. You cannot tell this answer apart from a handle that does not exist. Thus, a `403 workspace_key_required` always means "right team, wrong key". ##### Running a Format another workspace shared with you An owner can share a team Format with a different **workspace**, but never with a user. The method is the same as when a GitHub repository adds an outside collaborator. The owner workspace adds your team handle on the Format's **Access** tab (live immediately). Or, the owner invites your workspace with `POST /v1/formats/{handle}/{slug}/grants`. Then your admin accepts with `POST /v1/format-grants/{grant_id}/accept` and a key created in _your_ workspace. After that, you call the Format at the **owner's** address with **your own team key**: ```sh curl -sS -X POST "https://api.sume.com/v1/formats/{owner-handle}/{slug}/runs" \ -H "Authorization: Bearer $YOUR_TEAM_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{"instruction":"…"}' ``` The request has no workspace field. The key is the actor. The run, its spend, its concurrency slot and its media belong to your workspace, not to the owner. `GET …/runs` with your key lists only your runs. The owner still decides if the Format accepts any API calls. Its Status and API call trigger apply to every caller. The owner can also remove your access. After that, the address is a `404` for you again. The API refuses a personal key in the same way as on any team Format. Membership of the owner workspace does not replace a grant. If you have seats in both workspaces, the key that you use decides. Each workspace keeps its own run history and its own bill, and a key from one workspace never reads the runs of the other. #### Runs over the API are unattended A Format written for chat can pause and wait for a person. In the Agents UI, "approve these stills before I make the video" is a quality gate by design. Over the API, nobody is there. Thus, Sume tells the run that those approvals are already granted. The run then continues to the paid step in its spend cap. A run that really cannot finish comes back `failed`, never a half-finished `completed`: ```json { "data": { "status": "failed", "output": null, "output_error": { "code": "unattended_blocked", "message": "no avatar matched the brief, so no video was made." }, "error": { "code": "unattended_blocked", "message": "no avatar matched the brief, so no video was made." } } } ``` One qualification: a `completed` run always did real work and always fills `artifacts[]`. But it can still carry `output_error` when the projection did not match your `output_schema`. Examine `output_error` before you read `output`. Refer to [When output cannot be produced](/formats/structured-output#when-output-cannot-be-produced). #### 401 vs 403 vs 404 The HTTP class is the first branch. `error.code` is the second. A missing key gets `401 unauthorized`. A known key without `formats:read` / `formats:write` gets `403 insufficient_scope`, **never** `404 format_not_found`. Scopes cannot be patched onto an existing key. Mint a new key and rotate. | HTTP | `error.code` | When | |---|---|---| | 401 | `unauthorized` | No key, malformed key, two credentials at the same time, revoked, or unknown. `next_action` is `authenticate`. | | 403 | `insufficient_scope` | Valid key without `formats:read` / `formats:write`. `details.required_scope` names the scope. Also a service-account key on a run create or package write, with `details.reason` set. `next_action` is `authenticate`. | | 403 | `workspace_key_required` | You are a member of the team workspace, but you used a personal key. `details.workspace_id` names the workspace to mint a key in. For a team key from _a different_ workspace, the grant decides. The run starts when that workspace has an accepted grant. Otherwise, the answer is a `404`. | | 404 | `format_not_found` | Unknown, archived, outside this key's workspace, a team handle that you are not a member of, or a shared Format with a grant that is still pending or that the owner removed. A member's team key on the correct handle never gets this 404. | | 404 | `format_run_not_found` | Unknown run id, or a run that a different owner has. | | 404 | `format_run_queue_not_found` | Unknown bulk queue, or the queue of a different owner. | | 404 | `previous_run_not_found` | `previous_run_id` is unknown or not yours. | | 404 | `format_content_not_found` | The Format exists for this key but that package path does not. | Status classes obey RFC 9110. `insufficient_scope` is the RFC 6750 vocabulary. The `404` on a Format that someone else owns is a tenancy hide by design. It does not show that the Format exists at a different location (design record: #2393). #### Errors [Errors and spend](/formats/errors#errors-at-create) gives all the answers of the create call in one table. These are the errors that you will see first: | Status | `error.code` | What to do | |---|---|---| | 400 | `invalid_request` | The body named none of `instruction` / `input` / `previous_run_id` / `attachments`, sent both `output_schema` and `response_format`, or failed a size cap. Read `message`. | | 400 | `output_schema_invalid` | Your schema is outside the supported subset. `details.violations[]` names every problem. | | 401 | `unauthorized` | Correct the header, not the body. | | 402 | `insufficient_credits` | The workspace cannot fund the run. `next_action` is `add_funds`. | | 403 | `insufficient_scope`, `workspace_key_required` | Mint the correct key. Refer to the text above. | | 404 | `format_not_found` | Examine the address and the key that you use. | | 409 | `format_inactive`, `format_api_trigger_disabled` | The owner workspace sets these on the Format page's **API** tab. They apply to every caller, and also to the workspaces that the owner shared the Format with. | | 409 | `idempotency_conflict`, `idempotency_key_in_use` | Unstable key derivation, or a concurrent duplicate. Refer to [Idempotency](#idempotency). | | 429 | `rate_limited` | Wait `retry-after` seconds. | | 503 | `studio_agent_upstream_unavailable` | A Sume-side outage. Retry with the same `Idempotency-Key`. | `error.code` is a lowercase token that you can `switch` on. `message` is for humans and can change. #### Rate limits Every response carries `ratelimit-limit`, `ratelimit-remaining` and `ratelimit-reset`, and a `429` adds `retry-after`. Reads and writes have separate budgets. The read budget is forty times the write budget, so a poll loop cannot starve your own creates. A `429` names the budget that it came from in `error.details.scope`. For the full table, refer to [Errors and spend](/formats/errors#rate-limits). #### Next - [Runs and results](/formats/runs): the receipt, polls, webhooks, and how to continue and cancel a run - [Structured output](/formats/structured-output): bind a schema and get typed JSON back - [Errors and spend](/formats/errors): every code, credits, and rate limits - [Cookbook](/formats/cookbook): real-shaped bodies you can paste - [Bulk runs](/formats/bulk-runs): the same body, up to 100 times, with a concurrency window ### Structured output Source: https://docs.sume.com/formats/structured-output.md Bind a JSON Schema to a Format run and get typed, validated JSON back — supported schema rules, the post-run projection, and every failure mode. By default, a completed Format run gives you media and a paragraph of text. This result is good for a person, but it is not easy to use in a database. If you bind a schema, you get a typed object. It is the same run, but with a shape that you can write directly into your own records. ```json { "headline": "Aurora Headphones, all day quiet.", "hero_image": { "type": "image", "url": "https://media.sume.com/artifacts/artf_.../image-0.png", "content_type": "image/png", "file_name": "image-0.png", "size_bytes": null, "width": null, "height": null, "duration_ms": null, "expires_at": null }, "alt_text": "Aurora Headphones on a warm studio backdrop." } ``` The exact request and response schemas come from the live OpenAPI (`https://api.sume.com/reference/json`). The tables on this page are a summary that is easy to read. They are not a second schema. #### Not the same kind of thing as `input` A run request has two JSON-shaped fields, and they operate in very different ways. Confusion between the two fields is the most frequent cause of errors in a first integration. | | [`input`](/formats/call#input-caller-data) | `output_schema` | |---|---|---| | What it is | Caller data | A contract for the receipt | | Direction | You → the run | The run → you | | Shape | Any JSON object that your backend needs | JSON Schema, inside the [supported subset](#supported-schemas) | | Checked for | Object type, key count, byte size | Every rule in the subset | | A shape Sume does not expect | The run starts. Unknown keys are only more data. | `400 output_schema_invalid`. Nothing runs, and Sume charges nothing. | | Where it lands | The agent's prompt, as a fenced data block | The post-run projection | Thus, `input` is a flexible concatenation, and `output_schema` is a strict typed receipt. You can make your `input` as loose as you want. You cannot make your `output_schema` loose at all. `strict: false` does not make it loose, and there is no other way to bypass the rules. The two fields do not interact. **The projection never sees your `input`** (refer to the sections below). Thus, a value that you sent cannot come back in `output`, unless the run repeats it in its final text. #### Where your object comes from This part is different from a chat completion. Make sure that you understand it before you design a schema. There are two paths, and the receipt tells you which path you got. ```text 1. the run executes in its sandbox <- the recipe, your instruction, your input 2. the run submits your object <- filled_by: "agent" (preferred) ...or does not, and then: 3. the result is harvested <- generated media + the run's closing text 4. the harvest is projected onto your schema <- filled_by: "projection" 5. either way: gated, then `output` appears on the receipt ``` **`filled_by: "agent"` — the run answered.** Sume gives your schema to the run as a tool that the run must call before it stops. Your schema is the argument shape of that tool. The model that does the work is also the model that fills your object. It fills the object while it still knows what it made and why. It can see your `input` and your `instruction`, because they are part of the run. **`filled_by: "projection"` — the fallback.** If the run stops and did not submit a valid object, a separate constrained pass builds one after the run. The pass uses the data that the run left. This pass is an OpenAI-strict `json_schema` completion at temperature 0. It gets only two facts: | Fact | Detail | |---|---| | The run's generated media | All the artifacts that the run produced, with their durable URLs and metadata. | | The run's closing text | The final assistant message, truncated to the first 8000 characters. | The pass does **not** get your `input`, **not** your `instruction`, **not** the Format body, and **not** the intermediate steps of the run. On this path, all data that you want in `output` must be in one of these two facts. Thus, it is important to read `filled_by`. It shows if the run wrote the object, or if a pass built the object again from the data that the run left. It explains most unexpected results: - **On the projection path, your own identifiers do not round-trip.** The projection cannot see an `order_id` that you sent in `input`. Keep your identifiers on your side, with `run.id` or the `Idempotency-Key` that you sent as the key. Use `output` only for the data that the run made. - **Nothing in `output` is invented, on either path.** The same checks gate both paths before the object gets to you (refer to the sections below). An object that the run wrote does not get more trust. - **On the projection path, an empty text field is a signal.** The projection never sees your `input`. Thus, if the titles, descriptions or ids in a schema come from the brief, these fields come back `null`. At the same time, the media fields are full. `filled_by: "projection"` with null prose shows a run that stopped early. It does not show a Format that did not write copy. If you score runs automatically (a smoke matrix, a partner integration, a dashboard), read `filled_by` before you count a run as delivered. For video, read the file, not the number next to it. `ffprobe` on `primary_output_url` takes two seconds. It is the only method that shows the difference between an assembled cut and a clip that has the shape of one. The practical rule that comes from this is the same: **require only what the Format actually makes**. If a schema requires a field that the recipe never produces, the schema will fail on the projection path each time. The failure is silent until you read `output_error`. #### Bind a schema There are two spellings, with one behavior. Use the spelling that is applicable to your client. **Native — `output_schema`:** ```json { "instruction": "Make one hero image for the linked product.", "input": { "product_url": "https://example.com/p" }, "output_schema": { "name": "acme/promo-hero/v1", "strict": true, "schema": { "type": "object", "additionalProperties": false, "required": ["headline", "hero_image", "alt_text"], "properties": { "headline": { "type": "string" }, "hero_image": { "$ref": "SumeMediaFile#" }, "alt_text": { "type": ["string", "null"] } } } }, "primary_output_key": "hero_image" } ``` **OpenAI-shaped alias — `response_format`:** ```json { "response_format": { "type": "json_schema", "json_schema": { "name": "acme/promo-hero/v1", "strict": true, "schema": { "...": "as above" } } } } ``` Sume normalizes `response_format` into `output_schema` on the receipt. Thus, the receipt always shows the native spelling. **If you send both, the result is `400 invalid_request`.** The alias uses the **Chat Completions** spelling: `type` at the top, and the binding nested under `json_schema`. The OpenAI *Responses* API flattens the same fields into `text.format: { type, name, strict, schema }`. Sume does **not** accept this flattened shape. If your client builds Responses-style bodies, move the fields into `output_schema`, or nest them again under `json_schema`. | Field | Rules | |---|---| | `name` | Required. 1–64 characters, `^[A-Za-z0-9._/-]+$`. Give it a namespace, because it shows on every receipt. | | `strict` | The default is `true`. Refer to the note under [Supported schemas](#supported-schemas). | | `schema` | Required. A JSON Schema object inside the supported subset. | A Format can also have its own default schema, which you bind in the dashboard. A per-request `output_schema` overrides it for that run. The receipt shows which schema applied, in `output_schema.source`: | `source` | Meaning | |---|---| | `default` | No schema is bound. `output` is [the built-in schema](#the-built-in-schema). | | `action_default` | The Format's own bound schema. (`action_` is the wire spelling. [Scheduled](/agents/actions) uses the same spelling.) | | `request_override` | The `output_schema` that you sent on this run request. | #### Supported schemas Schemas must agree with the OpenAI strict-mode subset. This subset is a requirement, not a recommendation. If a schema is outside the subset, Sume rejects it at submit with `400 output_schema_invalid`. The response has a `details.violations[]` array that names each problem. Nothing runs. Thus, Sume does not charge you. ##### The supported keywords, exactly The subset uses an **allowlist**. A keyword that is not on this list is a violation. Sume does not ignore it silently. An ignored constraint is a part of the schema that Sume cannot promise to satisfy. | Group | Accepted | |---|---| | Structure | `type`, `properties`, `required`, `additionalProperties`, `items`, `$defs`, `$ref`, `anyOf` | | Values | `enum`, `const` | | Strings | `format`, `pattern`, `minLength`, `maxLength` | | Numbers | `minimum`, `maximum`, `exclusiveMinimum`, `exclusiveMaximum`, `multipleOf` | | Arrays | `minItems`, `maxItems` | | Annotation | `title`, `description`, `default`, `examples`, `$schema`, `$id` | Types are `string`, `number`, `integer`, `boolean`, `object`, `array`, `null`. Sume rejects all other keywords. These are the keywords that cause problems for real integrations: | Rejected | Instead | |---|---| | `oneOf` | `anyOf`. Only `anyOf` is on the list. Schemas ported from OpenAPI usually use `oneOf`. | | `allOf` | Flatten the branches into one object. | | `not`, `if` / `then` / `else`, `dependentRequired`, `dependentSchemas` | You cannot express these. Model the alternatives as `anyOf`, or validate on your side after you read `output`. | | `nullable: true` | A nullable union: `"type": ["string", "null"]`. | | `patternProperties`, `propertyNames`, `unevaluatedProperties`, `additionalItems` | Declare the properties that you want. `additionalProperties: false` covers the rest. | ##### Every node needs a `type` A node can also have a `$ref` or an `anyOf` in its place. A bare `{ "description": "…" }` is legal JSON Schema, and it means "anything". But here it is a `missing_type` violation. An `array` must also declare `items`. `$ref` and `anyOf` each **short-circuit** the node that contains them. Sume checks the sibling keywords next to them against the allowlist, but these keywords have no other effect. Put constraints inside the `anyOf` branches or inside the `$defs` entry, not next to the `$ref`. ##### The root must be an object ```json { "type": "object", "additionalProperties": false, "required": [], "properties": {} } ``` The `type` of the root must be `"object"`, as a single type. Thus, Sume rejects even `{ "type": ["object", "null"] }`. Sume also rejects a top-level array, string, or union. Wrap it: ```json // rejected { "type": "array", "items": { "type": "string" } } // accepted { "type": "object", "additionalProperties": false, "required": ["captions"], "properties": { "captions": { "type": "array", "items": { "type": "string" } } } } ``` ##### Every object needs `additionalProperties: false` This rule is applicable to all objects in the schema, not only to the root. This includes the objects nested inside array `items` and inside `$defs`. ```json // rejected: the nested object is open { "type": "object", "additionalProperties": false, "required": ["scene"], "properties": { "scene": { "type": "object", "properties": { "title": { "type": "string" } } } } } ``` ##### Every property must be listed in `required` There is no optional property. Each property that you declare must be present. Use a **nullable union** to express optionality: ```json // rejected: `subtitle` is declared but not required { "type": "object", "additionalProperties": false, "required": ["title"], "properties": { "title": { "type": "string" }, "subtitle": { "type": "string" } } } // accepted: `subtitle` is always present, and may be null { "type": "object", "additionalProperties": false, "required": ["title", "subtitle"], "properties": { "title": { "type": "string" }, "subtitle": { "type": ["string", "null"] } } } ``` This rule causes problems for more ported schemas than any other rule. Read `null` as "the run had nothing to put here". This is the same case for which you wanted `optional`. ##### Size and nesting limits | Limit | Value | Violation | |---|---|---| | Nesting depth | 10 levels | `max_depth` | | Total properties | 5000, counted across the whole document | `max_properties` | | Enum values | 1000 per enum | `max_enum_values` | | Total string length | 120,000 characters, summed over every property name, key, and string value in the document | `max_string_length` | The last limit is a document-wide budget, not a per-field cap. Thus, long `description` annotations on a large schema can use all of it, even when no single string is very long. ##### `$ref` is limited Only two targets resolve: | Target | Use | |---|---| | `#/$defs/*` | Your own definitions, declared at the **root** of the schema document. | | `SumeMediaFile#` | Sume's media shape. Refer to the section [below](#sumemediafile). | Sume rejects an external `$ref` (a URL, a sibling document, `#/components/...`). Sume also rejects `$ref: "#"`. OpenAI strict mode permits root recursion in this way, but Sume does not. Sume also rejects a `#/$defs/*` target that has no root `$defs` entry of that name. This rule finds a `$defs` block that is nested inside a sub-schema and not declared at the root. Recursion through a named definition is permitted. A `$defs` entry can `$ref` itself. The depth limit counts literal nesting in the document. Thus, a self-referential definition does not increase the depth count. ```json { "type": "object", "additionalProperties": false, "required": ["scenes"], "properties": { "scenes": { "type": "array", "items": { "$ref": "#/$defs/scene" } } }, "$defs": { "scene": { "type": "object", "additionalProperties": false, "required": ["caption", "clip"], "properties": { "caption": { "type": "string" }, "clip": { "$ref": "SumeMediaFile#" } } } } } ``` ##### `strict: false` does not relax any of this Sume accepts and stores it, but it does not change the subset above. If a schema is outside the subset, Sume rejects it, whether `strict` is `true` or `false`. Do not use it to bypass the subset. There is no way to bypass it. There is also no equivalent of the OpenAI JSON mode (`{"type": "json_object"}`), the loose "valid JSON, any shape" option. Bind a schema or use [the built-in one](#the-built-in-schema). These are the only two options. ##### Reading `details.violations[]` Each entry is `{ path, rule, message }`. `path` is a JSON-Pointer-style location, for example `#/properties/scenes/items/properties/clip`. `message` is text for a person, and it can change. `rule` is a stable lowercase token that you can safely `switch` on: | `rule` | Meaning | |---|---| | `root_must_be_object` | The root is missing, is not an object, or its `type` is not exactly `"object"`. | | `not_an_object` | A schema node is not a JSON object. | | `missing_type` | A node has no `type`, `$ref`, or `anyOf`. | | `unsupported_type` | A `type` that is not one of the seven types above. | | `unsupported_keyword` | A keyword that is not on the allowlist. | | `additional_properties_false` | An object node without `additionalProperties: false`. | | `required_completeness` | A declared property that is not in `required`, or a `required` entry that has no property of that name. | | `missing_items` | An `array` node with no `items`. | | `unsupported_ref` | A `$ref` that is neither `#/$defs/` nor `SumeMediaFile#`, or one that names a definition that does not exist. | | `invalid_defs` | `$defs` is present but is not an object of named schemas. | | `max_depth`, `max_properties`, `max_enum_values`, `max_string_length` | The [limits above](#size-and-nesting-limits). | Sume reports all the problems, not only the first. Thus, one `400` gives you sufficient data to fix the schema. #### Coming from OpenAI structured outputs If you used `response_format: { type: "json_schema", … }` or the Responses API's `text.format`, most of what you know is applicable here. The subset rules are the same, and they come from the same guide. The difference is *where Sume applies the schema*. | OpenAI | Sume | Note | |---|---|---| | `response_format.json_schema` | `output_schema`, or `response_format` verbatim | Chat Completions spelling only. Sume does not accept `text.format`. | | `json_schema.name` | `output_schema.name` | Required in both places. Give it a namespace, because it shows on every receipt. | | `json_schema.strict` | `output_schema.strict` | Accepted, the default is `true`, and it changes nothing. Sume always enforces the subset. | | `{"type": "json_object"}` (JSON mode) | *(no equivalent)* | Bind a schema, or use the built-in one. | | The model emits the JSON | A post-run projection emits it | The schema constrains the projection, never the run. | | `refusal` on the message | `output_error` on the receipt | Different mechanism: not a safety refusal but a failed projection. | | `incomplete_details.reason: "max_output_tokens"` | *(not applicable)* | The projection is small and bounded. There is no truncated-JSON case to handle. | | Streamed partial JSON | *(not applicable)* | `output` appears once, on the terminal receipt. | | `$ref: "#"` root recursion | Rejected | Recurse through a named `#/$defs/*` entry instead. | | Nothing comparable | [The URL gate](#the-url-gate) | Sume checks each URL in `output` against the media that the run really produced. | | Nothing comparable | [`SumeMediaFile#`](#sumemediafile) | A built-in `$ref` target for the run's media. | The change in how you think is this: with OpenAI, you constrain **what the model says**. Here, you constrain **how Sume reads back a completed run**. All other differences come from this. It is why the schema cannot make the Format produce a video, and why a hallucinated URL cannot get into `output`. It is also why a run can give you `output: null`. #### `SumeMediaFile` Use `{ "$ref": "SumeMediaFile#" }` at each location where you want a piece of the media of the run in your output. All fields are required. All fields except `type` and `url` are nullable. | Field | Type | Notes | |---|---|---| | `type` | `"image" \| "video" \| "audio" \| "file"` | | | `url` | string (uri) | Must be a URL that this run actually produced. Refer to [the URL gate](#the-url-gate). | | `content_type` | string \| null | For example, `video/mp4`. | | `file_name` | string \| null | | | `size_bytes` | integer \| null | | | `width`, `height` | integer \| null | Images and video. | | `duration_ms` | integer \| null | Video and audio. | | `expires_at` | string (date-time) \| null | **`null` for durable `media.sume.com` URLs**, which is the normal case. It has a value only when Sume returns a signed URL. | Sume serves Sume-hosted media from `media.sume.com` as `public, max-age=31536000, immutable`, and the media does not expire. Store the URL with your own record and render it later. It is not necessary to refresh it. A durable URL is also a *public* URL. If this is important to your product, refer to [Map artifacts into your UI](/cookbooks/embed-a-format#5-map-artifacts-into-your-ui). #### The built-in schema If you do not bind a schema, Sume projects `output` onto `sume/action-run-output/v1`: ```json { "text": "Short summary written by the Agent.", "images": [], "videos": [ { "type": "video", "url": "https://media.sume.com/...", "content_type": "video/mp4", "file_name": "teaser.mp4", "size_bytes": 4210233, "width": 1080, "height": 1920, "duration_ms": 12000, "expires_at": null } ], "audio": [], "files": [] } ``` `text` is nullable. The four arrays are always present, and they can be empty. Sume fills the built-in schema **deterministically** from the generated media and the final text of the run. Sume does not use a model. Thus, it cannot fail in the way that a custom schema can fail. If you want the media but do not want to design a schema, this schema is sufficient to ship with. #### The URL gate Before a custom `output` gets to you, Sume checks each URL in it against the set of media that this run actually produced. The comparison is exact string equality. Two passes find these URLs. The first pass collects each `http(s)://` string in the object, at any depth. It does this whether or not you declared the field as media. The second pass collects the `url` of each [`SumeMediaFile`](#sumemediafile)-shaped value, **whatever it contains**. Thus, a placeholder such as `"none"` or `""` in a `url` cannot get past the gate only because it does not look like a URL. If this run did not produce a URL, the projection fails, and Sume does not return the URL. This is also true for a well-formed `media.sume.com` URL that looks correct. Then Sume validates the object against your schema. Thus, a completed run returns output that agrees with your schema and has real media URLs. Or, it returns `output: null` and tells you why. **It never returns a schema-shaped guess.** ##### The deliverable is not one of its parts A schema that describes a whole made of parts (scenes, shots, segments) usually also has a field for the completed, assembled file. These are different files. Sume rejects a receipt that fills the whole with one of its own parts. The check is narrow intentionally. It needs **two or more** parts that report `"status": "succeeded"` with their own video. It also needs a video **outside** all parts that uses one of those files again. The check does not affect a Format whose single clip really is the deliverable. It also does not affect a poster or thumbnail that the run intentionally reused from a part. A run that could not assemble the file has an honest answer available. That answer is not "here is a scene". Report the parts that you made and leave the assembled field null. Or, report the real status of the assembly. ##### Durations are checked against the file A `duration_ms` inside a [`SumeMediaFile`](#sumemediafile) comes from the same ledger that fills `artifacts[]`. If that ledger recorded a length, the value in `output` must agree with it within 10%. A value that does not agree describes a different file. It fails the projection and does not get to you. If the ledger did not record a length, Sume does not check the value. In this case, `null` means "not measured", not "zero". Sume does not fail a run because of a fact that nobody measured. Thus, a `duration_ms` that you read back is the value of the artifact, or it is not verified. It is never a number calculated from other data. ##### What the gate does not check The gate checks only URLs, durations, and the shape. The projection pass reads all other values (ids, labels, captions, counts) from the media metadata and the closing text of the run. These values come from what the run reported, but Sume does not verify them against the run. Thus, treat these fields as the report of the run about its work, not as measurements. The projection writes a `duration_seconds` number that you declared yourself. The `duration_ms` on the media file is the value that Sume checks. The gate also admits only media that the run **generated**. A file that the run only uploaded is not in that set. Thus, if a schema field holds the URL of an upload, the projection fails, and all of the `output` fails with it. Do not put uploads in your schema. #### `primary_output_key` and `primary_output_url` Most integrations have one item to show. If you name it, the receipt resolves it for you: ```json { "primary_output_key": "hero_image" } ``` Then `primary_output_url` on the receipt is the URL at that key. The resolution order is: 1. The `primary_output_key` on the run request. 2. The Format's own `primary_output_key`. 3. The first top-level key that holds a media object, or an array whose first element is a media object. For the built-in schema, the fallback order is `videos` → `images` → `audio` → `files`. If you named a key in step 1 or 2, the receipt echoes it back when `output` has a value under it. This includes a value that is not media. If you point at a `headline` string, you get `primary_output_key: "headline"` with `primary_output_url: null`, because there is no URL to resolve. The resolution goes to step 3 only when the key is missing from `output`, or empty there. `primary_output_key` has a maximum of 64 characters. Both fields are `null` on any non-terminal status. They are also `null` when `output_error` is set. #### When output cannot be produced Before you read `output`, examine `output_error`. **On a run over the API, a projection failure is a run failure.** These runs are unattended, as [Call a Format](/formats/call#runs-over-the-api-are-unattended) states. A `completed` receipt whose `output` is `null` looks like a success, but it is not one. Thus, `status` is `failed`, and `error` has the same reason as `output_error`. In the Agents UI, a person reads the thread. There, the same shape stays `completed`, because it is a draft, not a receipt. **A failed run still publishes what it produced.** `output` is null when nothing satisfied your schema, not only because the run failed. For example, a live-commerce show rendered 20 of 40 scenes and then stopped. The show reports those 20 on `output`. The other scenes have the "failed" value that your own schema defines. The `output_error` next to them tells why the run stopped. A failure never gets a **pointer**: `primary_output_key` and `primary_output_url` are null on every non-`completed` run. Thus, `if (run.primary_output_url)` stays a safe test for "the deliverable exists", and a partial result cannot make it give a wrong answer. For this to work, your schema must permit the partial result. Refer to [Make a partial result legal](#make-a-partial-result-legal). **`artifacts[]` is populated either way**, with all the items that the run made. You always have the media, even when the shape failed. The exception is a projection that did not get access to the run. `output_extraction_failed` records a transport failure, not a verdict. Thus, the run stays `completed`, and Sume projects it again automatically on your next read. | `output_error.code` | Meaning | `details` | |---|---|---| | `output_schema_unsatisfied` | The projection did not agree with your schema, or it referenced media that this run did not produce. | `rejected_urls[]` (**first 10 only**) or `violations[]`, and a `harvested` count by media type. | | `output_extraction_failed` | The projection could not run. `status` stays `completed`. A `reason` of `harvest_unavailable` means that Sume could not read the media of the run while the run finalized. The receipt fills in on the next read. | `reason` | | `unattended_blocked` | The run stopped and did not claim a deliverable that it did not produce. `message` is the reason in the words of the run. | A `harvested` count by media type, and `assembled_deliverable: false`. The harvested media are intermediates, not the completed cut. | | `deliverable_missing` | The Format produces media (`io.output_kind`) that the run never made. Thus, Sume did not let any structured output claim it. | `declared_output_kind`, and a `harvested` count by media type. | | `primary_output_missing` | The run satisfied your schema, but the `primary_output_key` that you declared has no value. `output` still has the partial result. The run is `failed` because the item that you named is not in it. | `primary_output_key` | | `agent_reported_failure` | The accepted receipt of the run says that the run did not deliver: an explicit-fail payload, `failed` / `stand-in` media slots, or a primary of the wrong media type for the `io.output_kind` of the Format. `output` still has the ledger. The run is `failed` because the receipt says so. | `reason` (`explicit_failed`, `slots_failed`, `primary_not_deliverable`), `non_delivered_slots[]` or `primary_media_type`, and a `harvested` count by media type. | Treat this set as open, because new codes can appear. Branch on the codes that you handle, and use a default path for the rest. Do not use an exhaustive switch. ```json { "status": "failed", "output": null, "output_error": { "code": "output_schema_unsatisfied", "message": "The structured output referenced media URLs that this run did not produce.", "details": { "rejected_urls": ["https://media.sume.com/artifacts/artf_fake/image.png"], "harvested": { "images": 1, "videos": 0, "audio": 0, "files": 0 } } }, "error": { "code": "output_schema_unsatisfied", "message": "The structured output referenced media URLs that this run did not produce." }, "primary_output_key": null, "primary_output_url": null, "artifacts": [{ "type": "image", "url": "https://media.sume.com/artifacts/artf_t1/frame.png" }], "next_action": "none" } ``` Handle each case as this list shows: - **`output_schema_unsatisfied`, repeatedly, on the same Format.** The cause is almost always a schema that requires media that the recipe does not generate. Compare `details.harvested` with your required fields. The example above requires an image and got one, but it wanted a second image. Make the field a nullable union, or change the `instruction` so that the run makes the media. - **`output_schema_unsatisfied` with `violations[]`.** The problem is the shape, not the media. The violations name the paths that cause the problem. - **`output_extraction_failed`.** This failure is transient. Before you do anything else, read the run one more time. This step alone clears a `harvest_unavailable`. If the failure continues, retry the run with a **new** idempotency key. The old key is bound to the receipt that you already have. - **Any of them, in your UI.** You still have `artifacts[]`. Show the media and log the shape failure. This is better than an error for a customer whose video exists. #### Make a partial result legal A run can produce some of its output and then fail. The run can report this only if **your schema says a partial is a legal shape**. The platform does not add a minimum of its own. Ajv enforces the keywords that you wrote, and only those keywords. Thus, if a schema requires every scene, a 20-of-40 show returns `output: null`. Two rules do all the work. Both rules are not intuitive under the strict subset: 1. **Optional means nullable, not absent.** You must still list each property in `required` (refer to [Every property must be listed in `required`](#every-property-must-be-listed-in-required)). To show that a value can be missing, use `"type": ["string", "null"]`. 2. **Do not use `minItems` on the arrays that you want to receive partially.** Sume really enforces it. Thus, `minItems: 1` on a scene list will reject the ledger that you want to read. ```json { "name": "live-commerce-ledger/v1", "strict": true, "schema": { "type": "object", "additionalProperties": false, "required": ["full_video", "scenes", "notes"], "properties": { "full_video": { "type": ["string", "null"], "description": "The assembled show. null when it was never assembled." }, "scenes": { "type": "array", "description": "Every planned slot, in order. May be empty.", "items": { "$ref": "#/$defs/scene" } }, "notes": { "type": ["string", "null"] } }, "$defs": { "scene": { "type": "object", "additionalProperties": false, "required": ["id", "status", "video_url", "failure_reason"], "properties": { "id": { "type": "string" }, "status": { "type": "string", "enum": ["completed", "failed", "skipped"] }, "video_url": { "type": ["string", "null"] }, "failure_reason": { "type": ["string", "null"] } } } } } } ``` Use it together with `"primary_output_key": "full_video"`. This keeps the loose schema honest. A run that fills `scenes` but leaves `full_video` null satisfied the schema, but it did not produce the deliverable. Thus, the run terminalizes as `failed` with `primary_output_missing`, and it does not report a success. The loose schema lets you see the partial result. It does not let the run pass. When you read one of these receipts, branch on `full_video` for "did I get a show". Branch on each `scenes[].status` for "what do I need to retry". To retry, send the id of the failed run as `previous_run_id` on a new run (refer to [Continue a run](/formats/runs#continue-a-run)). The next turn continues the same conversation, and those clips are already in it. #### Failures at submit Sume finds these schema problems before anything runs. Sume does not charge you. | Code | Status | What to do | |---|---|---| | `output_schema_invalid` | 400 | Your schema is outside the supported subset. `details.violations[]` names each problem. | | `invalid_request` | 400 | This includes a request that sends both `output_schema` and `response_format`. | #### Checklist - [ ] The root is `{"type": "object"}` only. `additionalProperties: false` is on **every** object, also in `$defs`. - [ ] Every node has a `type`, a `$ref`, or an `anyOf`. Every `array` declares `items`. - [ ] Every declared property is in `required`. Optionality is a nullable union. - [ ] No `oneOf`, `allOf`, `not`, `if`/`then`/`else`, or `nullable: true`. - [ ] `$ref` targets are only `#/$defs/*` (declared at the root) and `SumeMediaFile#`, and never `#`. - [ ] The schema requires only media that the Format actually produces. - [ ] No field expects a value that you sent in `input`, because the projection cannot see it. - [ ] `name` is namespaced and stable. Thus, you can grep the receipts. - [ ] `primary_output_key` names the one item that your UI shows. - [ ] Your reader handles `output: null` with `output_error` set, on a `failed` run and on a `completed` one. - [ ] When `output` is null, your reader uses `artifacts[]` as a fallback. #### Next - [`input` — caller data](/formats/call#input-caller-data) — the other half of the run body, and the one with no schema - [Runs and results](/formats/runs) — the receipt that contains all of this - [Calling a Format](/formats/call) — the invoke contract - [Embed a Format in your product](/cookbooks/embed-a-format) — how to map output into your own records ### Runs and results Source: https://docs.sume.com/formats/runs.md The Format run receipt, and the two ways to get it. Poll the run, or take one signed webhook. Lifecycle, progress, continuing, canceling, listing. `POST …/runs` answers immediately with a receipt. The run takes minutes. This page is about the time after the `202`. It tells you how to know that the run finished, and what the finished receipt contains. There are two ways to know the result, and the two ways give the same receipt: | | Webhook | Poll | |---|---|---| | You do | Send `communication.webhook_url` on the create. Verify the signature. Answer `2xx`. | Read the run until `status` is terminal. | | You get | One signed POST for each run, when it completes or fails. | The same receipt, on your schedule. | | Costs you | One public HTTPS endpoint. | One timer for each run in progress, and read budget. | | Available | `api.dev.sume.com` and `api.sume.com`. | On all hosts. | Production integrations use the two ways. The webhook is the fast path. A read of `result_url` is the backup for the day when your endpoint is down. When a delivery fails, the run does not change. #### Lifecycle ```text queued -> processing -> completed | failed | canceled | skipped ``` | Status | Meaning | |---|---| | `queued` | Accepted, not started. | | `processing` | The run is in progress. | | `completed` | Finished. `output`, `artifacts[]`, and `primary_output_url` are populated. | | `failed` | Finished with an error. `error` gives the cause. `artifacts[]` still holds all the media that the run made. | | `canceled` | `POST …/cancel` stopped the run. The word has one `l`. | | `skipped` | The run never started: you sent `on_active_run: "skip"`, and a different run was in progress. | `next_action` on the receipt tells you what to do. It is `poll_status` while the run is `queued` or `processing`, `retry_later` on a `skipped` run, and `none` on each terminal run. Those are the only three values that a Format run emits. #### Poll The receipt holds its own URLs. Use these URLs. Do not build the paths yourself. | URL on the receipt | Endpoint | Returns | |---|---|---| | (the receipt itself) | `GET /v1/format-runs/{run_id}` | The full receipt at all statuses. Most integrations poll this URL. | | `status_url` | `GET /v1/format-runs/{run_id}/status` | The small poll payload: `status`, `next_action`, `cancelable`, `expires_at`, `queue`, timestamps, and the URLs again. When the run is terminal, it also holds `usage`, and `error` on a failure. It never holds `output`, `artifacts`, or `primary_output_url`. | | `result_url` | `GET /v1/format-runs/{run_id}/result` | The full receipt when the run is terminal. While the run is in progress, it returns `409 run_not_completed` with `details.status`. | | `events_url` | `GET /v1/format-runs/{run_id}/events` | The phase timeline. Refer to [Watch a run progress](#watch-a-run-progress). | | `cancel_url` | `POST /v1/format-runs/{run_id}/cancel` | Stops the run. Refer to [Cancel](#cancel). | A loop that reads the full receipt and stops on a terminal status: ```bash RUN_ID="arun_e43e6c5cb2b74052" SLEEP=5 while :; do RUN=$(curl -sS "https://api.sume.com/v1/format-runs/$RUN_ID" \ -H "Authorization: Bearer $SUME_API_KEY") STATUS=$(echo "$RUN" | jq -r '.data.status') case "$STATUS" in queued|processing) sleep "$SLEEP"; SLEEP=$(( SLEEP < 60 ? SLEEP * 2 : 60 )) ;; *) break ;; esac done echo "$RUN" | jq '{status: .data.status, primary: .data.primary_output_url, error: .data.error, output_error: .data.output_error}' ``` Use these three practices to keep a poll loop correct: - **Back off.** Long-form video is 15 to 30 minutes of work. Thus, a poll each second gives you nothing and uses read budget. Double the gap, up to one minute. A `429` or `503` during the loop is temporary, because the run continues to execute and to spend. Thus, wait and poll again, and do not think that the run failed. - **Use `expires_at` as your ceiling.** A non-terminal receipt holds the deadline after which Sume force-finalizes the run as `failed`. The deadline is 90 minutes from `created_at`, or earlier when the run is older than 25 minutes and was silent for 10. Use this value to set your own timeout, and do not invent a different one. When the run is terminal, it is `null`. - **Read `queue` if `queued` lasts.** `queue.state` is `waiting` while the pickup is in the normal window. It is `runtime_unavailable` when the run waited longer than that window and nothing claimed it. Then `retry_after_seconds` tells you how long to back off. `position` is always `null`, because Sume does not publish the queue depth. If a `runtime_unavailable` lasts more than a few minutes, send a support ticket with the `request_id`. In TypeScript, [`subscribeFormatRun`](/sdk/runs) creates the run and runs this loop for you. [`waitForRun`](/sdk/runs) does the loop for a run id that you already have. #### Webhook If you send `communication.webhook_url` on the create, Sume POSTs the terminal receipt to it one time. This occurs when the run completes or fails, on `api.dev.sume.com` and on `api.sume.com`. A canceled or skipped run never delivers. The cancel call answers you directly, and a skipped run is already terminal on the create response. ```json { "event": "format.run.terminal", "request_id": "arun_e43e6c5cb2b74052", "run_id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "OK", "outcome": "ok", "created_at": "2026-08-23T23:41:02.118Z", "payload": { "id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "completed", "format": { "id": "skl_…", "slug": "live-commerce", "title": "Live commerce", "version": 23 }, "output": { "full_video": { "type": "video", "url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "…": "…" } }, "primary_output_url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "artifacts": [{ "id": "artf_…", "type": "video", "url": "https://media.sume.com/artifacts/artf_…/full_video.mp4", "content_type": "video/mp4" }], "usage": { "currency": "USD", "billable_amount_usd_micros": 14959638, "generation_spend_cap_usd_micros": 120000000 }, "result_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/result", "thread_id": "thr_…" }, "error": null } ``` | Field | Branch on it for | |---|---| | `event` | Always `format.run.terminal` for a Format run. Route on it. You do not need to examine the body. | | `request_id`, `run_id` | Equal, and stable across retries. They are the dedupe key. | | `status` | `OK` when the run completed, `ERROR` when it failed. | | `outcome` | `ok`: completed with output. `degraded`: completed and billed, with real media in `artifacts[]`. But `output` is `null` because the projection did not match your schema, and `output_error` gives the cause. `error`: the run did not complete. Branch here when the question is "did I get usable output". | | `payload` | The run receipt. It is byte-identical to `data` from `GET /v1/format-runs/{run_id}`, thus one handler serves the two transports. It is `null` only when the receipt was more than 1 MiB. Then `error.code` is `payload_too_large`, and `error.result_url` gives the address to fetch it. | | `error` | `null` on `OK`. On other statuses, `{ code, message }`, the same as `payload.error`. | | `created_at` | The time when Sume built this delivery body. Use it to put deliveries in order. You cannot use `request_id` for the order, because it repeats on retries. | ##### Verify every delivery Each POST has three headers: ```text x-sume-webhook-timestamp: 1785000000 x-sume-webhook-signature: sume-v1= x-sume-webhook-secret-fingerprint: <12 hex chars> ``` The signature is HMAC-SHA256 over `.` with the signing secret of your workspace. You can read this secret on the Webhooks tab of the dashboard, or from `GET /v1/webhooks/signing-secret` (each key with `account:read` can read it). Before you parse the body, verify the signature against the raw bytes. Reject timestamps outside a five-minute window. Compare the fingerprint header with the fingerprint next to the secret. This makes sure that the two sides hold the same secret. The [Cookbook](/formats/cookbook#a-webhook-receiver) has complete receivers in Node and Python. In TypeScript, the check is one call: ```ts import { verifyWebhook } from "@sume-com/sdk"; const ok = await verifyWebhook({ body: rawBody, headers: request.headers, secret: process.env.SUME_COM_WEBHOOK_SIGNING_SECRET!, }); ``` ##### Delivery rules | Property | Value | |---|---| | When | One time for each run, on `completed` or `failed`. Never on `canceled` or `skipped`. | | Success | All `2xx` codes, in not more than 10 seconds. Record the event in durable storage. Then answer. Then do the work. | | Retries | Up to 10 attempts. The backoff is the longer of two values: exponential (30 s × 2^(attempt−1), with jitter), or your `Retry-After` on a `429`/`503`. The maximum backoff is one hour. | | Redirects | Sume does not follow redirects. A `3xx` is a failed attempt, thus register the final URL. | | URL rules | Public HTTPS only. Localhost, private ranges, credentials in the URL, and plain HTTP give `400 invalid_request` at create. Sume examines the URL again at delivery time. | | Dedupe | On `request_id`. Each retry repeats it. | A delivery outcome never changes the run. After ten refused attempts, you have a failed *delivery* and a run that is still `completed`. Fetch the run from `result_url`. ##### Check what happened to a delivery Each receipt for a run created with a `webhook_url` has a `webhook_delivery` block: ```json { "webhook_delivery": { "url": "https://acme.example.com/hooks/sume", "event_type": "format.run.terminal", "status": "delivered", "attempts": 1, "max_attempts": 10, "next_attempt_at": null, "last_attempt_at": "2026-08-23T23:41:04.000Z", "delivered_at": "2026-08-23T23:41:04.000Z", "last_status_code": 200, "last_error": null, "signature_version": "sume-v1", "signing_secret_fingerprint": "2b3844659f5e" } } ``` | `status` | Means | |---|---| | `not_armed` | Sume stored the URL, and no delivery is scheduled yet. The run is still in progress. | | `pending`, `retrying` | Armed. `next_attempt_at` is the time of the next attempt. | | `delivered` | Your endpoint answered `2xx`. | | `failed`, `exhausted` | Sume stopped the attempts. `last_status_code` and `last_error` (our transport error, never your body) give the cause. The run did not change. | You can replay a terminal delivery after you repair your receiver, or to do a test of a receiver. To replay, call `POST /v1/format-runs/{run_id}/webhook/redeliver` with `formats:write` and an empty body. It re-POSTs the current receipt with a new timestamp and signature. It does not use one of the ten automatic attempts. The call returns `409 webhook_not_configured` when the run had no URL. It returns `409 run_not_terminal` while the run is still in progress. The full contract is on [Run webhooks](/agents/run-webhooks). That page also gives the generation-job webhooks that `POST /v1/models/…` emits on a different event set. #### Watch a run progress `status` tells you if a run is done. `GET /v1/format-runs/{run_id}/events` tells you what the run does now: ```json { "data": [ { "at": "2026-08-23T23:23:41.000Z", "phase": "preparing", "status": "done", "duration_ms": 1840 }, { "at": "2026-08-23T23:31:12.000Z", "phase": "running", "status": "running", "duration_ms": null }, { "at": "2026-08-23T23:40:57.000Z", "phase": "finalizing", "status": "done", "duration_ms": 620 } ] } ``` `preparing` is all the work before the agent gets the run. `running` is the time when the agent does the recipe, and most of the time goes there. `finalizing` is teardown and output harvest. Each entry has a `status` (`pending`, `running`, `done`, `warning`, `error`, `skipped`), and a `duration_ms` when the phase measured itself. Consecutive entries with the same phase and status become one entry, and its `at` continues to increase. Thus, `at` on the last entry is the progress clock of the run. If this clock does not move for several minutes, the run is stalled, not slow. Sume will finalize the run at the bounds above. It is a phase timeline, not a log stream. Sume does not publish agent output, tool calls, or sandbox internals here, and will not publish them in the future. There is no push channel for progress. Poll this endpoint, or ignore it and use the terminal webhook. #### The run receipt `GET /v1/format-runs/{run_id}` returns the full shape at all statuses. | Field | Notes | |---|---| | `id`, `object` | `arun_…`, and `format.run`. | | `format` | `{ id, slug, title, version }`. `id` is the opaque `skl_…`. `version` is the Format version that ran. Later edits do not change it. | | `status`, `next_action`, `cancelable` | Refer to [Lifecycle](#lifecycle). `cancelable` is `true` while the run is `queued` or `processing`. | | `trigger` | `{ source: "api", idempotency_key }` for each run that you create. | | `created_at`, `started_at`, `finished_at` | The last two are `null` until the events occur. | | `expires_at`, `queue` | The deadline and the pickup state. Refer to [Poll](#poll). | | `output_schema` | `{ name, strict, source }`: the schema that gave `output` its shape. `source` is `request_override` when you sent a schema. | | `output` | The structured result. `null` on each non-terminal status, and `null` when no result satisfied `output_schema`. A `failed` run still publishes the partial result that it wrote, if there is one. | | `output_error` | `{ code, message, details }` when the run could not produce `output`. Examine it before you read `output`. | | `primary_output_key`, `primary_output_url` | The one item to show. Both are `null` unless the run is `completed`. | | `artifacts[]` | Each durable file that the run generated: `{ id, type, url, content_type, size_bytes, width, height, duration_ms, checksum_sha256 }`. Empty until the run is terminal. Also populated on failures. | | `usage` | `{ currency, billable_amount_usd_micros, generation_spend_cap_usd_micros, debited_usd_micros, held_usd_micros, refunded_usd_micros, final, cap }`. Refer to [Reading `usage`](#reading-usage). `null` when the API could not read the spend. | | `error` | `{ code, message }`. Non-null only when the run is `failed`. | | `skip_reason` | Set on a `skipped` run. | | `webhook_delivery` | The delivery state for the URL that you registered, or `null` if you registered no URL. | | `idempotency_hit` | `true` when this receipt is an idempotency replay, not a new run. | | `thread_id`, `previous_run_id` | The conversation that this turn occurred in, and the run that it continued. Refer to [Continue a run](#continue-a-run). | | `model` | The catalog id that the orchestrator ran on. | | `request_id` | Log it. Support asks for this value. | | `status_url`, `result_url`, `events_url`, `cancel_url` | Refer to [Poll](#poll). | Media URLs are durable `media.sume.com` HTTPS URLs. They do not expire. Each person who has the URL can open it. Thus, if your product needs per-customer access control, proxy or copy the media. #### Continue a run A Format run is one agent turn. If you send `previous_run_id` on a new `POST …/runs`, the next turn continues the same conversation. Sume replays to the agent what the agent produced. Thus, the agent can do one part again and keep the other parts as they are. Live-commerce integrations use this method to retry a single scene, and they do not pay for the full show again. ```bash curl -sS -X POST "https://api.sume.com/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: order-8823-lc-v1-retry-sc7" \ -d '{ "previous_run_id": "arun_e43e6c5cb2b74052", "instruction": "Retry the selected scene only. Keep every other scene and the voice track unchanged.", "input": { "scene_id": "sc_7" }, "output_schema": { "…": "the same schema you bound on the first run" }, "primary_output_key": "full_video", "generation_spend_cap_usd": 8 }' ``` The run that you name must be continuable. Its receipt shows this: `thread_id` is not `null`, and the run `completed` or has a non-empty `artifacts[]`. You can continue a `failed` run that left work. You cannot continue a run that left nothing. | Refusal | Means | |---|---| | `404 previous_run_not_found` | The id is unknown, or a different owner has the run. | | `400 previous_run_format_mismatch` | That run started on a different Format. Continue it on the Format where it started. | | `409 previous_run_not_terminal` | The run did not finish yet. Poll it. Then call again. | | `400 previous_run_not_resumable` | There is nothing to continue: no `thread_id`, or the run is not `completed` and has no artifacts. `details` shows `previous_run_status`, `has_thread`, and `artifact_count`. Start a new run. | A continuation is a new run, with a new id, a new receipt, its own spend cap, and its own single webhook. The original run never changes. The two runs share `thread_id`, which is read-only. To continue, name `previous_run_id`, not a thread id (a thread id gives `400 unknown_parameter`). Bind the same `output_schema` on each turn, because it is per run and not inherited. `artifacts[]` on a continued run lists all the media that the full conversation generated, but `usage` stays per run. #### Cancel ```bash curl -sS -X POST "https://api.sume.com/v1/format-runs/$RUN_ID/cancel" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The cancel call needs `formats:write`, and it is idempotent. In the two cases, the current receipt comes back, and `cancel_effect` tells you which case occurred. `canceled` means that this call stopped a run in progress. `no_op` means that the run already finished before the call. You pay for the generation that the run completed before the cancel, and `usage` shows it. A canceled run never delivers a webhook. If your integration is webhook-only, cancel is the one path where no delivery will arrive. Use the receipt that this call returns. #### List runs for a Format ```bash curl -sS "https://api.sume.com/v1/formats/acme/live-commerce/runs?limit=20" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The list shows the newest runs first. `limit` is 1–100, and the default is 20. `GET /v1/formats/{format_id}/runs` is the opaque twin. If no run of the Format ever started over the API, the call returns an empty list, not a `404`. The response is a page: send `next_cursor` back as `cursor` until `has_more` is `false`. The cursor is opaque and keyset over `(created_at, id)`. Thus, runs created while you page do not move rows. A cursor that is not ours gives `400 invalid_request`. ##### Routes that do not exist Three paths that callers guess, and what to use in their place: | Guess | Use | |---|---| | `GET /v1/format-runs` | There is no cross-Format list. List the runs of each Format, or keep your own index, with the `data.id` that you stored at create as the key. | | `GET /v1/formats/{handle}/{slug}/runs/{run_id}` | Read runs at `/v1/format-runs/{run_id}`. The Format path only creates and lists runs. | | `GET /v1/format-runs/{run_id}/messages` | Sume does not publish the conversation over the API. `events_url` gives the phase timeline. `output` and `artifacts[]` hold the result. | #### Reading `usage` `usage.billable_amount_usd_micros` is the generation spend of this run. Sume enforces the cap of the run against this total, and the total counts reserved and captured amounts. This value increases while the run is in progress, and settles when the run terminates. It does not include the LLM turn of the agent, thus it is not the total cost of the run. `usage.cap` gives the same calculation in parts: `limit_usd_micros`, `counted_usd_micros` (this value), and `remaining_usd_micros`. The cost is `usage.debited_usd_micros`: the real amount that the wallet deducted for the run and its thread. This amount contains the captured ledger rows of all operation types, and also the LLM row of the turn. `held_usd_micros` are holds that are still open (not spend yet). `refunded_usd_micros` are holds that Sume gave back (not spend). `final` changes to `true` when no hold is open. [`GET /v1/usage?run_id=`](/dashboard/usage) uses the same rows and the same fold. Thus, the receipt, the ledger, and the answer of an agent always agree. `usage` is `null` when the API could not read the spend at all. This is different from `0`. The three wallet fields are `null` on receipts written before the ledger answered. #### Errors on the run endpoints | Status | `error.code` | What to do | |---|---|---| | 401 | `unauthorized` | The key is missing, malformed, revoked, or unknown. | | 403 | `insufficient_scope` | The key does not have `formats:read` (reads) or `formats:write` (cancel, redeliver). Mint a new key. | | 404 | `format_run_not_found` | The run id is unknown, or a different owner has the run. A run that you cannot see gives the same result as a run that does not exist. | | 409 | `run_not_completed` | You sent `GET …/result` before the run was terminal. `details.status` holds the current status. Poll `status_url`. Then retry. | | 429 | `rate_limited` | Wait for `retry-after`. Polls use the read budget. This budget is separate from the write budget, and much larger. | | 503 | `studio_agent_upstream_unavailable` | An outage on the Sume side, not a problem with your key. Retry at a later time. The run continues. | [Errors and spend](/formats/errors) gives all the codes, together with the run-time failures (`error` and `output_error`). #### Next - [Errors and spend](/formats/errors): all codes, what a `failed` run holds, credits, and rate limits - [Structured output](/formats/structured-output): how to shape `output`, and what to do when it is null - [Cookbook](/formats/cookbook): a webhook receiver, a scene retry, a batch - [Run webhooks](/agents/run-webhooks): the complete delivery contract - [Waiting for runs](/sdk/runs): `subscribeFormatRun`, `waitForRun`, and the phase timeline from TypeScript ### Bulk runs Source: https://docs.sume.com/formats/bulk-runs.md Queue up to 100 Format runs with a concurrency window — POST …/bulk-runs, poll GET /v1/format-run-queues/{id}, and read each child at GET /v1/format-runs/{run_id}. You can leave a list of Format runs to run overnight. You do not have to drive the fan-out from your laptop. A bulk request is a **server-side queue of ordinary Format runs**, not a different execution engine. Each item is the same unit of work as `POST …/runs`: one sandbox, one agent turn, one [run receipt](/formats/runs). The accurate request and response schemas come from live OpenAPI (`https://api.sume.com/reference/json`). The tables here are a readable summary, not a second schema. #### Endpoints | Method | Path | Scope | |---|---|---| | `POST` | `/v1/formats/{format_id}/bulk-runs` | `formats:write` | | `POST` | `/v1/formats/{handle}/{slug}/bulk-runs` | `formats:write` | | `GET` | `/v1/format-run-queues/{queue_id}` | `formats:read` | The two POST paths are twins. For new integrations, we recommend `{handle}/{slug}`. The opaque `skl_…` path stays valid forever. The request body, headers, scopes, and queue receipt are the same for both paths. The API has no public list-queues or cancel-queue endpoint. To cancel a child, use `POST /v1/format-runs/{run_id}/cancel`. Refer to [Cancel](/formats/runs#cancel). #### Auth The API-key rules are the same as for a single [Format run](/formats/call): 1. Use a Bearer API key (`Authorization: Bearer $SUME_API_KEY`). 2. The key carries `formats:write` to create a queue and `formats:read` to poll it. 3. If a **team workspace** owns the Format, use a key issued **in that workspace**. 4. Service-account keys cannot create Format runs or bulk queues. They fail with `403 insufficient_scope` and `details.reason` of `service_account_format_runs_unsupported`. Keys created before the release of the Format API-call trigger do not carry these scopes. Create a new key. You cannot add scopes to a key that already exists. Refer to [Calling a Format](/formats/call#scopes) and [Team Formats need a team key](/formats/call#team-formats-need-a-team-key). | Scope | Needed for | |---|---| | `formats:write` | `POST …/bulk-runs` (and single-run create / cancel). | | `formats:read` | `GET /v1/format-run-queues/{queue_id}` (and Format / run reads). | #### Create a queue ```bash curl -sS -X POST "https://api.sume.com/v1/formats/chase/product-promo/bulk-runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{ "concurrency": 3, "items": [ { "instruction": "clip 1", "input": { "url": "https://example.com/1.jpg" } }, { "instruction": "clip 2", "input": { "url": "https://example.com/2.jpg" } }, { "instruction": "clip 3", "input": { "url": "https://example.com/3.jpg" } }, { "instruction": "clip 4", "input": { "url": "https://example.com/4.jpg" } } ] }' ``` The opaque twin: ```bash curl -sS -X POST "https://api.sume.com/v1/formats/skl_.../bulk-runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "concurrency": 3, "items": [ { "instruction": "clip 1" }, { "instruction": "clip 2" } ] }' ``` An accepted create returns **`202`** with `{ "data": { …queue } }`. If the list is longer than the window, the first `concurrency` items are already in flight on that receipt. ##### Worked example: a production batch The small example above keeps the item bodies one line long. A real batch usually holds one long item for each spreadsheet row. The envelope still has only two keys: ```json { "concurrency": 2, "items": ["{/* one Format-run body per row */}"] } ``` Mobidoo uses this method for its live-commerce sheet against the vanity path `POST /v1/formats/mobidoo/live-commerce/bulk-runs`: `concurrency: 2` and 20 items. Each item is one broadcast row. The client skips rows with no finished draft. Each item binds its own `output_schema` and `generation_spend_cap_usd`. [Mobidoo → Bulk runs](/enterprise/mobidoo/bulk-runs) shows the **unabridged** item body from that batch: full `instruction`, full script, schema, and per-item webhook. That batch shape makes three items clear: - **The queue has no webhook.** `communication.webhook_url` is per item. Poll `status_url` for queue-level progress. - **`completed` is not "all succeeded".** It means that every item is terminal. Branch on `counts.failed`. - **Mint a fresh `Idempotency-Key` per batch.** If you replay a spent key, the API returns `202` with the *old* queue. ##### Request body | Field | Notes | |---|---| | `concurrency` | Required integer **1–16**. The number of child Format runs that stay in flight at the same time. | | `items` | Required array, **1–100** entries, in order. Each entry is the same body as [`POST …/runs`](/formats/call#request-body). | | `idempotency_key` | Optional body form of the `Idempotency-Key` header. If you send both, the header wins. | The API rejects unknown top-level fields. You must send `items`. `{ "concurrency": 3 }` is `400`, not an empty queue. Each item must name at least one of `instruction`, `input`, `previous_run_id`, or `attachments`. This is the same rule as for a single run. A bad item fails the **create** (`400 invalid_request`, `details.index`) **before** a queue exists. Sume dispatches nothing. The bulk controller owns single-flight. The controller executes every item with `on_active_run: "allow"`. Sending `skip` or `reject` on an item does not stall the window on the first in-flight run of this Format. Workspace generation concurrency still applies to the children. Per-item `communication.webhook_url` is the same as for a single run. Each child can register its own terminal webhook. The **queue object has no webhook**. Do not expect a queue-level callback. #### Queue receipt `data` is a `format.run_queue`: ```json { "data": { "id": "frq_...", "object": "format.run_queue", "format": { "id": "skl_...", "slug": "product-promo", "title": "Product promo", "version": 3 }, "concurrency": 3, "status": "running", "counts": { "total": 4, "queued": 1, "running": 3, "completed": 0, "failed": 0, "canceled": 0 }, "items": [ { "index": 0, "status": "running", "run_id": "arun_...", "error": null }, { "index": 1, "status": "running", "run_id": "arun_...", "error": null }, { "index": 2, "status": "running", "run_id": "arun_...", "error": null }, { "index": 3, "status": "queued", "run_id": null, "error": null } ], "created_at": "2026-08-24T00:00:00.000Z", "updated_at": "2026-08-24T00:00:00.000Z", "finished_at": null, "status_url": "https://api.sume.com/v1/format-run-queues/frq_..." } } ``` | Field | Notes | |---|---| | `id` | Queue id (`frq_…`). | | `object` | Always `format.run_queue`. | | `format` | `{ "id", "slug", "title", "version" }`. `id` is the opaque `skl_…`. | | `concurrency` | The window you sent (1–16). | | `status` | `queued` / `running` / `completed`. See below. | | `counts` | `total`, `queued`, `running`, `completed`, `failed`, `canceled`. All required. | | `items` | One row per submitted item, in the same order. | | `created_at`, `updated_at` | ISO-8601. | | `finished_at` | Set when the queue status becomes `completed`. Otherwise, `null`. | | `status_url` | `GET /v1/format-run-queues/{id}`. Poll this URL for progress. | ##### Queue status | Status | Meaning | |---|---| | `queued` | The queue did not dispatch any item yet. | | `running` | The concurrency window drains the list. | | `completed` | **Every item is terminal.** Examine `counts` for failures. `completed` is not "all succeeded". | `counts.total` is `items.length`. `counts.running` includes items that the API still treats as in flight (a child that the API claimed but did not yet give a `run_id` still counts as `running`). #### Items | Field | Notes | |---|---| | `index` | Zero-based position in the submitted `items` array. | | `status` | `queued` / `running` / `completed` / `failed` / `canceled`. | | `run_id` | Child Format run id after dispatch. `null` while queued, and `null` if the item failed before a child run started. Poll `GET /v1/format-runs/{run_id}` for the full receipt. | | `error` | `{ "code", "message" }` or `null`. | | Item status | Meaning | |---|---| | `queued` | Not started. `run_id` is `null`. | | `running` | In the concurrency window. | | `completed` | Child run `completed`. Terminal. Frees a slot. | | `failed` | Child run `failed` (or `skipped`, recorded here as `failed`), or the child could not start. Terminal. Frees a slot. | | `canceled` | Child run is canceled. Terminal. Frees a slot. | A failed item that never started a run still keeps that `index`, with `run_id: null` and `error` set to the create-run failure (`code` / `message` from that attempt, for example `format_run_failed_to_start`). The rest of the queue continues. When a child run settles, `error` is: | Child run | Item `error` | |---|---| | `completed` | `null` | | `failed` | `{ "code": "format_run_failed", "message": "The Format run failed." }` | | `canceled` | `{ "code": "format_run_canceled", "message": "The Format run was canceled." }` | To find **why** a child failed, read the run receipt (`GET /v1/format-runs/{run_id}`), not only the queue item. #### Concurrency window The server keeps `concurrency` child runs in flight. When a slot opens, the server immediately starts the next queued item, until the list is drained. This is not client-driven fan-out. In-flight slots are items with the public status `running`. `completed`, `failed`, and `canceled` are terminal and **free a slot**. When an in-flight child gets to `completed` or `failed` (the usual drain path), the next queued item starts immediately. Thus, the window stays full. A `canceled` child does the same. Create already fills the window. With `concurrency: 3` and 8 items, the `202` receipt shows three `running` and five `queued`. When item 0 completes, item 3 starts. The window stays at three until fewer than three items remain. The window never goes above `concurrency`, even if a poll or an advance of progress occurs more than once. Child runs still go through ordinary Format-run admission (wallet, workspace generation concurrency, spend caps). If a child fails to start, that **item** becomes `failed`. The queue create already returned `202`. #### Poll the queue Use `status_url`, or build `GET /v1/format-run-queues/{queue_id}` from `id`: ```bash curl -sS "https://api.sume.com/v1/format-run-queues/$QUEUE_ID" \ -H "Authorization: Bearer $SUME_API_KEY" ``` `200` returns the same queue object as create. Use `counts` for a dashboard. Use `items` when you need per-row `run_id` / `error`. The queue is `completed` when every item is terminal. Then the API sets `finished_at`. Branch on `counts.failed` and `counts.canceled`. Queue `completed` does not mean success. A queue that you cannot see gets the same answer as a queue that does not exist: `404 format_run_queue_not_found` (unknown id, or a queue that a different user owns). #### Child runs Each dispatched item is an ordinary Format run: ```bash curl -sS "https://api.sume.com/v1/format-runs/$RUN_ID" \ -H "Authorization: Bearer $SUME_API_KEY" ``` `status_url` / `result_url` / `events_url` / `cancel_url` on that receipt operate the same as on [Runs and results](/formats/runs). The queue does not replace those endpoints. It adds counts and per-item status on top. When you cancel a child (`POST /v1/format-runs/{run_id}/cancel`), the queue marks that item `canceled` and frees its slot for the next queued item. #### Idempotency Send `Idempotency-Key` on create as a header. The API also accepts body `idempotency_key`, but the header wins. The scope of a key is one Format. | Replay | Result | |---|---| | Same key, same `{ concurrency, items }` | `202` and the queue that already exists. | | Same key, different payload | `409 idempotency_conflict` (`details.queue_id` names the original). | Unlike a single run, a bulk replay stays **`202`**. The queue object has no `idempotency_hit` field. The queue does not add a separate idempotency-in-use lock. It uses only what the control plane already stores for that key. There is no queue-level webhook. There is no queue-level idempotency behavior other than the header / body key above. #### Errors Create (`POST …/bulk-runs`): | Code | Status | What to do | |---|---|---| | `unauthorized` | 401 | Missing, malformed, revoked, or unknown API key. | | `insufficient_scope` | 403 | The key does not have `formats:write`, or it is a service-account key (`details.reason` is `service_account_format_runs_unsupported`). `next_action` is `authenticate`. You cannot patch scopes. Mint a new key. Missing scopes never give `format_not_found`. | | `workspace_key_required` | 403 | Team Format, personal key. Use a key created in `details.workspace_id`. | | `invalid_request` | 400 | `concurrency` is not an integer 1–16. Or, `items` is missing, empty, or longer than 100. Or, an item is not an object. Or, `items[i]` names none of `instruction` / `input` / `previous_run_id` / `attachments`. `details.index` names a bad item. | | `format_not_found` | 404 | Unknown, archived, outside your key's workspace, or a team handle that you are not a member of. | | `format_api_trigger_disabled` | 409 | The API call trigger is off for this Format. | | `format_inactive` | 409 | The Format is inactive. | | `idempotency_conflict` | 409 | You already used that `Idempotency-Key` with a different bulk payload. | | `invalid_attachment` / `attachment_not_found` / `attachment_too_large` / `attachment_fetch_failed` | 400 / 413 / 502 | The API sends these codes when it resolves an item's `attachments`, **before** it creates the queue. These are the same codes as in [Calling a Format](/formats/call#errors). | | `rate_limited` | 429 | Wait `retry-after` seconds. Create spends the write budget. | | `studio_agent_upstream_unavailable` | 503 | A Sume-side outage. Retry later. | Poll (`GET /v1/format-run-queues/{queue_id}`): | Code | Status | What to do | |---|---|---| | `unauthorized` | 401 | Missing or invalid API key. | | `insufficient_scope` | 403 | The key does not have `formats:read`. `details.required_scope` names the scope. | | `format_run_queue_not_found` | 404 | Unknown queue id, or a queue that a different owner has. | | `rate_limited` | 429 | Wait `retry-after`. A poll spends the **read** budget, which is separate from the create budget. | | `studio_agent_upstream_unavailable` | 503 | Retry later. The queue continues to drain. | A `429` or `503` during a poll loop is transient. The queue continues to work. Back off. Do not think of it as a failed queue. Wallet / admission failures on a **child** after `202` do not fail the create. That item becomes `failed` with the create-run error, and the window refills from the remaining `queued` items. #### End to end ```bash export SUME_API_KEY="sume_live_..." QUEUE=$(curl -sS -X POST "https://api.sume.com/v1/formats/chase/product-promo/bulk-runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{ "concurrency": 3, "items": [ { "instruction": "clip 1", "input": { "url": "https://example.com/1.jpg" } }, { "instruction": "clip 2", "input": { "url": "https://example.com/2.jpg" } }, { "instruction": "clip 3", "input": { "url": "https://example.com/3.jpg" } }, { "instruction": "clip 4", "input": { "url": "https://example.com/4.jpg" } } ] }') QUEUE_ID=$(echo "$QUEUE" | jq -r '.data.id') STATUS=$(echo "$QUEUE" | jq -r '.data.status') while [ "$STATUS" = "queued" ] || [ "$STATUS" = "running" ]; do sleep 5 QUEUE=$(curl -sS "https://api.sume.com/v1/format-run-queues/$QUEUE_ID" \ -H "Authorization: Bearer $SUME_API_KEY") STATUS=$(echo "$QUEUE" | jq -r '.data.status') done echo "$QUEUE" | jq '.data.counts' # Queue status completed means every item is terminal — inspect counts.failed. ``` In production, use exponential backoff, not a fixed five-second sleep. When the Format makes video, each child is still minutes of work. If you poll the queue every second, you get no benefit and you use your rate limit. To read a finished child's media, get `run_id` from `items[]`. Then use the procedure in [Runs and results](/formats/runs). #### Next - [Calling a Format](/formats/call) — the per-item invoke contract and every submit error - [Runs and results](/formats/runs) — child receipts, polls, cancelation - [Structured output](/formats/structured-output) — bind a schema on each item - [Run webhooks](/agents/run-webhooks) — per-child `communication.webhook_url`, not a queue callback - [Embed a Format in your product](/cookbooks/embed-a-format) — partner integration around a single run ### Errors and spend Source: https://docs.sume.com/formats/errors.md Every error the Formats API returns, by stage. The envelope, what to do per code, what a failed run carries, webhook delivery outcomes, credits and spend caps, and rate limits. All errors on the API have the same envelope, and each code is a lowercase token that you can branch on. This page lists all the codes of the Formats surface in one place. It puts the codes in groups, by the time when they occur. The groups are: before a run exists, while you read it, when the run fails, and when a webhook delivery failed. Spend and rate limits are at the end. #### The error envelope ```json { "error": { "code": "workspace_key_required", "message": "This Format belongs to a team workspace. Use an API key created in that workspace.", "request_id": "req_0123456789abcdef0123456789abcdef", "category": "auth", "stage": "auth", "retryable": false, "retry_after_seconds": null, "public_reason": "workspace_key_required", "next_action": "authenticate", "details": { "workspace_id": "org_…" } } } ``` | Field | Use it for | |---|---| | `code` | The stable token to `switch` on. It matches `^[a-z0-9_]+$` and is never a sentence. | | `message` | Text for a human. It can change. Log it. Never match on it. | | `request_id` | The API also sends it as the `x-sume-request-id` header. Give it to support. | | `retryable`, `retry_after_seconds` | Tells you if the same request can succeed when you send it again, and how long to wait first. | | `next_action` | `authenticate`, `fix_input`, `add_funds`, `retry_later`, `poll_status`, `inspect_events` or `contact_support`. | | `category`, `stage`, `public_reason` | Less specific labels for dashboards and alerts. | | `details` | Code-specific data: `required_scope`, `workspace_id`, `violations[]`, `index`, `status`, `scope`, …. Each code below names its fields. | Branch on the HTTP status first. Then branch on `code`. A `4xx` at create means that nothing ran and the API charged nothing. Thus, correct the call. Do not retry it. The most frequent and most expensive mistake is to retry a `403 insufficient_scope` in a loop. #### Errors at create These errors apply to `POST /v1/formats/{handle}/{slug}/runs` and `POST …/bulk-runs`. Nothing runs, and the API charges nothing. A failed create releases its `Idempotency-Key`. | Status | `error.code` | Meaning | What to do | |---|---|---|---| | 400 | `invalid_request` | One of these conditions is true. The body names none of `instruction` / `input` / `previous_run_id` / `attachments`. The body names both `output_schema` and `response_format`. `input` is not an object, has more than 64 top-level keys, or is more than 2 MiB. `generation_spend_cap_usd` is `0` or more than 500. `model` is not in the catalog. A bulk body has bad `concurrency` or `items`. A webhook URL is not public HTTPS. The `communication` aliases do not agree. | Read `message` and `details`. | | 400 | `unknown_parameter` | A top-level field that the API does not know. This also applies to `thread_id`. `details.errors[].suggestion` names the field that you probably intended. | Rename the field, or remove it. | | 400 | `output_schema_invalid` | The schema is not in the supported subset. `details.violations[]` lists each `{ path, rule, message }`. | Correct each violation. Refer to [Supported schemas](/formats/structured-output#supported-schemas). | | 400 | `invalid_attachment` | An `attachments[]` item is bad, or the attachments go above the media budget. `attachments[]` shares this budget with `input` URLs. | Refer to [Attachment errors](/formats#attachment-errors). | | 400 | `attachment_not_found` | `asset_id` is unknown in this workspace. | Upload it, or use a URL. | | 400 | `previous_run_format_mismatch` | `previous_run_id` names a run that started on a different Format. | Continue it on that Format. | | 400 | `previous_run_not_resumable` | That run left nothing to continue. `details` shows `previous_run_status`, `has_thread`, and `artifact_count`. | Start a new run. | | 401 | `unauthorized` | There is no key, a malformed key, a revoked key, a key for the other host, or two credentials at the same time. | Correct the header. | | 402 | `insufficient_credits` | The workspace wallet cannot fund the run. `next_action` is `add_funds`. | Add funds to the wallet. If you retry before you add funds, you get the same answer. | | 402 | `organization_wallet_not_provisioned` | The organization workspace has no funded wallet. | An admin must fund it. | | 403 | `insufficient_scope` | The key does not have `formats:write` (`details.required_scope`), or it is a service-account key (`details.reason: service_account_format_runs_unsupported`). | Mint a new key with the scopes. You cannot add scopes to a key that you already have. | | 403 | `workspace_key_required` | The Format is a team Format, and the key is a personal key. `details.workspace_id` names the workspace. For a team key from a different workspace, the API uses its grant to make the decision (the run starts, or `404`). | Create a key in that workspace, or ask its owner for a grant. | | 404 | `format_not_found` | The handle or slug is unknown, or the Format is archived. Or the Format is not in the workspace of your key. Or it is a team Format, and your workspace has no accepted grant for it (a pending or revoked grant gives the same result). | Examine the address and the key. For a shared Format, ask the owner for a grant. | | 404 | `previous_run_not_found` | `previous_run_id` is unknown or not yours. | Examine the id. | | 409 | `format_inactive` | The owner workspace set the Format to Inactive. This applies to all callers, and also to the workspaces that the owner shared it with. | Owner workspace: Format page → **API** tab → Status **Active**. | | 409 | `format_api_trigger_disabled` | The owner workspace disabled the API call trigger. This applies to all callers, and also to the workspaces that the owner shared it with. | Owner workspace: Format page → **API** tab → API call trigger **On**. | | 409 | `format_run_in_progress` | `on_active_run: "reject"` is set, and a run is already in progress. | Wait, or remove `reject`. | | 409 | `format_not_forkable` | You addressed a built-in capability, not a Format. | Call a Format by Sume, or call your own Format. | | 409 | `previous_run_not_terminal` | The run that you want to continue is still in progress. | Poll the run. Then call again. | | 409 | `idempotency_conflict` | You already used that `Idempotency-Key` with a different body. For bulk runs, `details.queue_id` names the original queue. | Correct how you derive your keys. Do not retry without a change. | | 409 | `idempotency_key_in_use` | A different request with the same key is in progress. `retryable: true`. | Wait approximately one second. Then send the request again. | | 413 | `payload_too_large` | The request body is more than 4 MiB (`details.limit_bytes`). | Make `input` smaller. Send media by URL. | | 415 | `unsupported_media_type` | You did not send the body as `application/json` (`details.received_content_type`). | Set `Content-Type: application/json`. | | 413 | `attachment_too_large` | An image is more than 30 MB, or a set is more than 500 MB. | Resize. | | 429 | `rate_limited` | No write budget remains for this key. `error.details.scope` is `write`. | Wait for `retry-after`. Refer to [Rate limits](#rate-limits). | | 502 | `attachment_fetch_failed` | Sume could not fetch an attachment (`details.index`). The status is `5xx`, but the cause is your input: `next_action` is `fix_input`. | Make sure that the URL is public. | | 503 | `studio_agent_upstream_unavailable` | An outage on the Sume side. | Retry at a later time with the same `Idempotency-Key`. | | 4xx/5xx | `format_run_failed_to_start` | The run could not start, and no more specific code is applicable. | Read `message`. Retry one time. If the error occurs again, contact support with `request_id`. | A `202` never changes to one of these errors at a later time. After you get a receipt, failures arrive on it as `status: "failed"`. #### Errors while reading These errors apply to `GET /v1/format-runs/{run_id}`, `/status`, `/result`, `/events`, `POST …/cancel`, `POST …/webhook/redeliver`, `GET /v1/format-run-queues/{queue_id}`, and the Format reads. | Status | `error.code` | Meaning | What to do | |---|---|---|---| | 401 | `unauthorized` | The same as above. | Correct the header. | | 403 | `insufficient_scope` | Reads need `formats:read`. Cancel and redeliver need `formats:write`. | Mint a new key. | | 404 | `format_run_not_found` | The run id is unknown, or a different owner has the run. | Examine the id and the key. A run that you cannot see gives the same result as a run that does not exist. | | 404 | `format_run_queue_not_found` | The queue is unknown, or a different owner has it. | The same as above. | | 404 | `format_not_found` | The same as above. | The same as above. | | 409 | `run_not_completed` | You sent `GET …/result` before the run was terminal. `details.status` is the current status. `retry_after_seconds` gives a recommended time for the next poll. | Poll `status_url`. Then read `result_url`. | | 409 | `webhook_not_configured` | You sent redeliver for a run created without a `webhook_url`. | There is nothing to redeliver. | | 409 | `run_not_terminal` | You sent redeliver while the run is still in progress. | Wait for the terminal receipt. | | 429 | `rate_limited` | No read budget remains (`details.scope: read`). | Wait for `retry-after`. The run continues. | | 503 | `studio_agent_upstream_unavailable` | An outage on the Sume side. | Retry. The run continues. | A `429` or `503` in a poll loop is temporary. If you stop the loop, the run and its spend do not stop. Thus, back off and poll again. #### When the run itself fails A run that could not finish comes back with `status: "failed"`, `error: { code, message }`, and usually `output_error` with more data. `artifacts[]` still lists all the media that the run generated. `output` still holds the partial result that satisfied your schema, if there is one. A failure never gets a pointer: `primary_output_url` is `null`. Thus, `if (run.primary_output_url)` stays a safe test for "the deliverable exists". | `error.code` | Meaning | What to do | |---|---|---| | `unattended_blocked` | The run stopped at a gate that it could not pass without a person. No avatar matched the brief, or an input was missing that the run asks for in chat. `message` is text that you can show to a user. | Correct the input or the brief. Retry with a new `Idempotency-Key`. | | `output_schema_unsatisfied` | The run finished, but its result did not match your `output_schema`, or the result referenced media that the run did not produce. The data is in `details.rejected_urls[]` or `details.violations[]`, with a `harvested` count by media type. | Compare `details.harvested` with the requirements of your schema. Usually, the schema asks for a file that the Format never makes. Change that field to nullable, or change the instruction. | | `deliverable_missing` | The Format declares that it produces media (`io.output_kind`), but this run made no media. | Retry. If the error occurs again, the input is not what the recipe expects. | | `primary_output_missing` | The result satisfied your schema, but the `primary_output_key` that you named is empty. `output` holds the partial result. | Continue the run to fill the gap, or retry. | | `agent_reported_failure` | The accepted `return_format_output` of the run said that the run did not deliver. The cause is an explicit-fail payload, or media slots that report `failed` / `stand-in`. Or the primary is not the declared deliverable of the Format (audio or a still under a video key). `output` holds the ledger of what the run made. `primary_output_url` is null. | Read `details.reason` and `output`. The clips on `output` are real, and a retry does not generate them again. Continue the run, or re-fire it with a new `Idempotency-Key`. | | `incomplete_assembly` | The run reached its time limit before all its generation jobs finished. Thus, the delivered media is not all the media that the run paid for. `details.pending_job_count` and `details.pending_jobs[]` name the unfinished jobs. When your schema permits it, the partial ledger is on `output`. | Continue the run with `previous_run_id`. The finished clips are on the thread, and the continuation does not generate them again. | | `output_extraction_failed` | The projection could not run. `details.reason: harvest_unavailable` means that the media was not readable while the run finalized. In that case, `status` stays `completed`, and the receipt fills in on the next read. `details.reason: harvest_threw` means that the harvest of the host crashed after the run finished. In that case, the run is `failed`, and `details.thrown` holds the frame that threw and the build. If the ledger of the run holds none of the declared media of the Format, the run reports `deliverable_missing`, not this code. | `harvest_unavailable`: read the run one more time. Then retry with a new key. `harvest_threw`: this is a host defect. Report the run id. The clips on the thread are real, and a retry does not generate them again. | | `mcp_unavailable` | The per-turn Sume MCP tools that this run needed did not attach. Thus, the host failed the run before the model started, and did not run a turn without tools (#7378). `details.retryable` is `true` and `details.charged` is `false`. No generation ran, and there is no charge. | Retry with a new `Idempotency-Key`. If the error occurs again, the cause is the tools, not your request. | | `provider_unavailable` | The provider stream of the model stopped, and all its reconnects failed before the run produced its deliverable. Your input did not cause this error. `details.retryable` is `true`. | Retry with a new `Idempotency-Key`. The finished clips are on the thread, and the retry does not generate them again. | | `provider_credits_exhausted` | The account of Sume at the model provider used all of its credit, thus the run stopped. The cause is on the Sume side. It is not your Sume balance and not your input. `details.retryable` is `false`. | Do not re-fire immediately. When Sume reports that the provider account is restored, retry with a new `Idempotency-Key`. | | `format_run_failed` | The generic failure. | Read `message`. A run that tried to spend more than its cap also gets this code. Thus, compare `usage.billable_amount_usd_micros` with `usage.generation_spend_cap_usd_micros` before you raise the brief. | This set of codes is open, and new codes can appear. Thus, handle the codes that you know, and use a fallback for all other codes. Retry a failed run with a **new** `Idempotency-Key`. The old key is bound to the receipt that you already have. When the failure left clips, it is better to [continue the run](/formats/runs#continue-a-run) than to start a new run. On a webhook, the same run arrives with `status: "ERROR"`, `outcome: "error"`, and an `error.code` that is the same as `payload.error.code`. A run that completed but could not fill your schema arrives as `status: "OK"`, `outcome: "degraded"`. It has real media and `output: null`. #### Webhook delivery failures A delivery outcome never changes the run. The `webhook_delivery` block on each receipt tells you the result: | `webhook_delivery.status` | Meaning | What to do | |---|---|---| | `retrying` | An attempt failed. `next_attempt_at` is the time of the next attempt. For HTTP `429`/`503` from your endpoint, the API obeys `Retry-After` up to one hour. | Nothing, unless `last_status_code` shows a problem on your side. | | `failed`, `exhausted` | Ten attempts failed because of a refusal, a timeout (10 s each), or a redirect. Or the URL failed re-validation. `last_status_code` and `last_error` tell you which cause. | Read the receipt from `result_url`. Correct the endpoint. To replay, send `POST …/webhook/redeliver`. | | Envelope with `payload: null` | The receipt was more than 1 MiB. `error.code` is `payload_too_large`, and `error.result_url` gives the address to fetch it. `status` still shows the real outcome of the run. | Fetch `result_url`. A handler that always reads `payload` as an object will throw on your largest runs. | For the full delivery rules, refer to [Runs and results](/formats/runs#webhook). #### Credits and spend A run spends from the workspace of the key. Two gates apply, at different times: | Gate | When | On failure | |---|---|---| | Wallet | At create. The workspace must be able to fund the run. | `402 insufficient_credits` (`next_action: add_funds`), or `402 organization_wallet_not_provisioned`. Nothing ran. | | Spend cap | During the run. The run cannot spend more than its effective cap. | The run ends as `failed`. `usage` shows how near to the cap the spend got. | The cap is the control that you own. Each Format has a cap (`generation_spend_cap_usd_micros` on the Format, or $400 when the Format never set one). `generation_spend_cap_usd` on the request names the ceiling of this run, up to the $500 platform maximum. The API accepts a value above the cap of the Format, runs `null` at $500, and rejects `0`. Production long-form runs usually have caps of approximately $120, and a single-scene retry has a few dollars. For details, refer to [Spend caps](/formats/call#spend-caps). You pay for metered generation (video, image, avatar, voice, timeline work) at the rates on the [API pricing page](https://www.sume.com/pricing/api). The receipt shows this amount as `usage.billable_amount_usd_micros`. This value increases while the run is in progress. It counts reserved and captured amounts, and it settles when the run terminates. This value does not include the LLM turn of the agent. Thus, it is not the total cost of the run. It is also a receipt value, not an invoice: [`GET /v1/usage`](/dashboard/usage) and `GET /v1/balance` are the billing records. `usage` is `null` when the API could not read the spend. This is different from `0`. You pay for generation that finished before a cancel or a failure. If a later step fails, you do not get a refund for that generation. A `4xx` at create, an idempotent `200` replay, and a `skipped` run cost nothing. #### Rate limits Each key has a per-minute request budget for all of `/v1`. The plan of the workspace sets this budget. Reads and writes have separate budgets. Reads get forty times the write number, thus polls cannot block your own creates. | Plan | Writes per minute | Reads per minute | |---|---:|---:| | Free | 120 | 4800 | | Pro | 300 | 12000 | | Startup | 600 | 24000 | | Scale | 1200 | 48000 | | Enterprise | Contracted (Scale until provisioned) | Contracted | A **read** is a `GET` of all types: the receipt, `status_url`, `events_url`, `result_url`, and the Format and run lists. A **write** is all other requests: create calls for runs and queues, cancel, and redeliver. Each response holds the current state. A `429` names the budget that it came from: | Header | Meaning | |---|---| | `ratelimit-limit` | The number of requests permitted in the current window, for the budget that this request used. | | `ratelimit-remaining` | The number of requests that remain in that window. | | `ratelimit-reset` | The number of seconds until the window resets. | | `retry-after` | The number of seconds to wait. The API sends it on `429`. `error.details.scope` is `read` or `write`. | Use the headers to set your pace. Do not count the requests yourself. The request rate is not the generation capacity. The concurrency limit of the plan controls how many generations run at the same time. If you increase your request rate, this limit does not increase. #### Getting help Each response has `x-sume-request-id`, and each receipt has `request_id`. When you contact support, include these values, the run id, and the `error.code`. Do not send API keys, signing secrets, or raw media URLs. #### Next - [Create a run](/formats/call): the create contract, with the `401` / `403` / `404` dialect in detail - [Runs and results](/formats/runs): polls, webhooks, `webhook_delivery`, and how to continue a failed run - [Structured output](/formats/structured-output): schema rules and each `output_error` - [Authentication](/authentication): keys, rotation, and the full rate-limit notes ### Cookbook Source: https://docs.sume.com/formats/cookbook.md Copy-paste recipes shaped like production traffic. A live-commerce run with typed output and a webhook, a webhook receiver in Node and Python, a scene retry on the same thread, a sheet-driven batch, and a wait loop. Each recipe on this page has the shape of a real production request. Placeholders replace the data of the customer. Replace `acme/live-commerce`, the URLs, and the copy with your values. Keep the structure. Do this setup one time: ```bash export SUME_API_KEY="sume_live_..." # a workspace key with formats:read + formats:write export SUME_API="https://api.sume.com" # or https://api.dev.sume.com with a development key ``` #### A live-commerce run with typed output and a webhook This is the most frequent production call. It has a product page, a host image, a tagged script, and a schema that names the assembled cut and each scene. It also has a per-run cap and a webhook. Keep the body in a file. Real scripts are too long for a shell heredoc. ```bash curl -sS -X POST "$SUME_API/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: sheet-45-v1" \ -d @run-body.json ``` `run-body.json`: ```json { "instruction": "Use the Intro/Mid/Fin script as written; do not shorten or add sentences. Korean host, vertical 9:16, 30fps, no BGM, no captions. Keep card typography inside the top 40% of the frame.", "input": { "sheet_no": 45, "product_url": "https://shop.example.com/p/4438469916", "product_name": "Aurora French Terry Crewneck", "brand_name": "Aurora", "on_card_name": "Aurora daily crewneck", "highlights": [ "Soft brushed french terry", "Sizes S to XL", "31,000 → 24,810 (20% off)" ], "product_image_urls": [ "https://cdn.example.com/products/4438469916/still-800.jpg" ], "host_image_url": "https://cdn.example.com/hosts/chaerin.png", "vo_language": "ko", "price": { "currency": "KRW", "list": 31000, "sale": 24810, "discount_label": "20%" }, "script": { "segments": [ { "tag": "Intro", "text": "안녕하세요, …" }, { "tag": "Mid", "text": "…" }, { "tag": "Fin", "text": "…" } ] } }, "output_schema": { "name": "acme/live-commerce/v1", "strict": true, "schema": { "type": "object", "additionalProperties": false, "required": ["full_video", "scenes"], "properties": { "full_video": { "$ref": "SumeMediaFile#" }, "scenes": { "type": "array", "items": { "$ref": "#/$defs/scene" } } }, "$defs": { "scene": { "type": "object", "additionalProperties": false, "required": ["id", "role", "status", "video"], "properties": { "id": { "type": "string" }, "role": { "type": "string", "enum": ["talk", "broll", "under_banner"] }, "status": { "type": "string", "enum": ["succeeded", "stand-in", "failed"] }, "video": { "$ref": "SumeMediaFile#" } } } } } }, "primary_output_key": "full_video", "generation_spend_cap_usd": 120, "communication": { "mode": "webhook", "webhook_url": "https://acme.example.com/hooks/sume" } } ``` What each key does: | Key | Why it is there | |---|---| | `instruction` | Decisions: use the script as written, the framing, and what not to add. It is prose, much shorter than 4000 characters. | | `input` | Data: the recipe reads the keys that it knows (`product_url`, `host_image_url`, `vo_language`, `script`, `price`). The other keys go to the run as context. You can put your own reference data (`sheet_no`) here, but it will not come back in `output`. | | `output_schema` | `full_video` is the deliverable. `scenes[]` gives you each clip with a stable `id` that you can retry. `SumeMediaFile#` is the built-in media shape. Each property is necessary. An optional property is a nullable property. | | `primary_output_key` | Makes `primary_output_url` the assembled cut. If a run fills `scenes` but not `full_video`, the receipt is `failed`, not a false success. | | `generation_spend_cap_usd` | The ceiling for this run. For production live-commerce runs, this value is approximately $120. | | `communication.webhook_url` | The API sends one signed POST when the run ends, thus no poll is necessary. Keep `result_url` as the backup. | From the `202`, store `data.id` and `data.thread_id` with your row. Use the first value to get the receipt. Use the second value to group retries. When the webhook arrives, `payload.output.full_video.url` is the show and `payload.output.scenes[]` are the clips. The retry recipe below repairs a scene with `status: "stand-in"` or `"failed"`. #### A webhook receiver The receiver does four steps in this sequence. First, it verifies the signature against the raw bytes. Then it answers `2xx` quickly. Next, it dedupes on `request_id`. Only after these three steps, it acts on `outcome`. The two versions below handle the oversized-receipt case (`payload: null`). ##### Node This version uses only the Web `Request` API. Thus, it works in a Next.js route handler, Hono, Workers, or Deno. It uses `verifyWebhook` from `@sume-com/sdk`. ```ts import { verifyWebhook } from "@sume-com/sdk"; const SECRET = process.env.SUME_COM_WEBHOOK_SIGNING_SECRET!; // Webhooks tab of the dashboard export async function POST(request: Request) { const raw = await request.text(); // the raw string, never a re-serialized object if (!(await verifyWebhook({ body: raw, headers: request.headers, secret: SECRET }))) { return new Response("bad signature", { status: 401 }); } const event = JSON.parse(raw); if (event.event !== "format.run.terminal") return new Response(null, { status: 204 }); // Dedupe: retries repeat request_id. Insert-or-ignore, then decide whether to process. const fresh = await db.webhookEvents.insertIfAbsent({ id: event.request_id, body: raw }); if (!fresh) return new Response(null, { status: 204 }); // Answer now. The run is already terminal, and the work below can take as long as it needs. queueMicrotask(() => handle(event).catch(console.error)); return new Response(null, { status: 204 }); } async function handle(event: any) { // Over 1 MiB the receipt is not inlined. Fetch it instead. const receipt = event.payload ?? (await fetch(event.error.result_url, { headers: { Authorization: `Bearer ${process.env.SUME_API_KEY}` }, }) .then(r => r.json()) .then(r => r.data)); switch (event.outcome) { case "ok": return markReady(receipt.id, receipt.output, receipt.primary_output_url); case "degraded": // Real media in artifacts[], but output is null: usually a schema asking for // something the Format never produces. Show the media, log output_error. return markNeedsReview(receipt.id, receipt.artifacts, receipt.output_error); case "error": return markFailed(receipt.id, receipt.error, receipt.artifacts); } } ``` If you cannot use the SDK, you can write the check in a dozen lines. First, reject a timestamp that is more than five minutes off. Then calculate the HMAC-SHA256 of `${timestamp}.${raw}` with your secret, hex-encoded. Compare the result in constant time against the value after `sume-v1=` in `x-sume-webhook-signature`. The full function is on [Run webhooks](/agents/run-webhooks#signature). ##### Python This version uses FastAPI. It reads the raw body before it parses the JSON. ```python import hashlib import hmac import json import os import time from fastapi import FastAPI, Request, Response SECRET = os.environ["SUME_COM_WEBHOOK_SIGNING_SECRET"].encode() TOLERANCE_SECONDS = 300 app = FastAPI() def verify(raw: bytes, timestamp: str | None, signature: str | None) -> bool: if not timestamp or not signature: return False try: ts = int(timestamp) except ValueError: return False if abs(time.time() - ts) > TOLERANCE_SECONDS: return False digest = hmac.new(SECRET, f"{ts}.".encode() + raw, hashlib.sha256).hexdigest() return hmac.compare_digest(f"sume-v1={digest}", signature) @app.post("/hooks/sume") async def sume_webhook(request: Request): raw = await request.body() if not verify( raw, request.headers.get("x-sume-webhook-timestamp"), request.headers.get("x-sume-webhook-signature"), ): return Response(status_code=401) event = json.loads(raw) if event.get("event") != "format.run.terminal": return Response(status_code=204) if not record_once(event["request_id"], raw): # your insert-or-ignore return Response(status_code=204) enqueue(handle, event) # answer first, work later return Response(status_code=204) def handle(event: dict) -> None: receipt = event["payload"] or fetch_receipt(event["error"]["result_url"]) outcome = event["outcome"] if outcome == "ok": mark_ready(receipt["id"], receipt["output"], receipt["primary_output_url"]) elif outcome == "degraded": mark_needs_review(receipt["id"], receipt["artifacts"], receipt["output_error"]) else: mark_failed(receipt["id"], receipt["error"], receipt["artifacts"]) ``` Two problems occur in each first receiver. First, a framework that parses JSON for you destroys the signed bytes. Thus, read the raw body on this route. Second, a receiver that renders video before it answers uses all of the 10-second attempt budget. Then the API retries the delivery while the receiver works. Thus, record and then answer before you process the event. You can do a test of the receiver without a real run. `POST /v1/webhooks/test-deliveries` (or **Send test** on the dashboard) sends a `webhook.test` payload to your URL. To replay a real delivery, use `POST /v1/format-runs/{run_id}/webhook/redeliver`. #### Retry one scene on the same thread A run is one turn of a conversation. To do a clip again, continue that conversation with `previous_run_id`. Name the scene in the request. The voice track, the other clips, and the script do not change. The API returns the full scene list, re-assembled. ```bash curl -sS -X POST "$SUME_API/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: sheet-45-v1-retry-sc_7" \ -d '{ "previous_run_id": "arun_e43e6c5cb2b74052", "instruction": "Retry the selected scene only. Keep every other scene and the voice track unchanged. Do not change the script. New take only.", "input": { "scene_id": "sc_7" }, "output_schema": { "…": "identical to the first run" }, "primary_output_key": "full_video", "generation_spend_cap_usd": 8, "communication": { "webhook_url": "https://acme.example.com/hooks/sume" } }' ``` Put the note of an operator in `instruction` (which scene, what is wrong, how it must change). Keep `input.scene_id` as the machine-readable pointer. For two scenes at the same time, use `"scene_ids": ["sc_7", "sc_9"]`. The API returns a new run (`arun_…`, a new receipt, its own webhook) on the same `thread_id`. `output.scenes[]` is the full list again. The retried scene has a new URL, the other scenes keep their URLs, and the API re-assembles `full_video` at a new URL. Plan the budget for a single-scene retry as a fraction of the create spend. Measured production retries cost approximately one twentieth of the spend of the first run. Always send a cap. A retry is a new take, not a re-encode. The API generates all the generative parts of that scene again. If the look changes, retry the scene. If the words, the host, or the product change, start a new production with new scene ids. The API refuses the continuation with `400 previous_run_not_resumable` when the earlier run left nothing. It refuses with `409 previous_run_not_terminal` while the earlier run is still in progress. It refuses with `400 previous_run_format_mismatch` if you address a different Format. Refer to [Continue a run](/formats/runs#continue-a-run). #### Batch a sheet with bulk runs One row of a broadcast sheet becomes one item. The sheet becomes one `POST …/bulk-runs`. The queue keeps `concurrency` runs in progress. When a slot becomes free, the queue starts the next run. ```bash curl -sS -X POST "$SUME_API/v1/formats/acme/live-commerce/bulk-runs" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: sheet-2026-09-03-v1" \ -d @bulk-body.json ``` `bulk-body.json` has two keys. Each item is the same as the body of a single run (the first recipe on this page), one time for each row. On the client side, skip rows that do not have a finished script. There is no empty item. ```json { "concurrency": 2, "items": [ { "instruction": "…", "input": { "sheet_no": 45, "product_url": "…" }, "output_schema": { "…": "…" }, "primary_output_key": "full_video", "generation_spend_cap_usd": 120, "communication": { "webhook_url": "https://acme.example.com/hooks/sume" } }, { "instruction": "…", "input": { "sheet_no": 46, "product_url": "…" }, "output_schema": { "…": "…" }, "primary_output_key": "full_video", "generation_spend_cap_usd": 120, "communication": { "webhook_url": "https://acme.example.com/hooks/sume" } } ] } ``` `202` returns a queue (`frq_…`). In this queue, the first `concurrency` items are already `running`. Keep your own sheet-row ↔ `index` map. `items[i].index` is the position that you submitted. To get the progress, poll the queue, not the children: ```bash QUEUE_ID="frq_…" while :; do Q=$(curl -sS "$SUME_API/v1/format-run-queues/$QUEUE_ID" -H "Authorization: Bearer $SUME_API_KEY") echo "$Q" | jq -c '.data.counts' [ "$(echo "$Q" | jq -r '.data.status')" = "completed" ] && break sleep 30 done # Every item is terminal now. Read each child's receipt, and look at counts.failed before calling it done. echo "$Q" | jq -r '.data.items[] | select(.run_id != null) | .run_id' | while read -r RUN_ID; do curl -sS "$SUME_API/v1/format-runs/$RUN_ID" -H "Authorization: Bearer $SUME_API_KEY" \ | jq -c '{id: .data.id, status: .data.status, primary: .data.primary_output_url, output_error: .data.output_error.code}' done ``` The batch shape makes three things clear. The queue has no webhook, because `communication.webhook_url` is per item. Queue `completed` means that each item is terminal, not that each item succeeded. Thus, branch on `counts.failed` and on the `output_error` of each child. A used `Idempotency-Key` returns `202` with the *old* queue, thus mint a new key for each batch. For the full contract, refer to [Bulk runs](/formats/bulk-runs). #### Wait for a run without a webhook If you cannot expose an endpoint (a script, a CI job, a one-off), poll with backoff. Stop when the run gets to a terminal status. ```bash RUN_ID=$(curl -sS -X POST "$SUME_API/v1/formats/acme/live-commerce/runs" \ -H "Authorization: Bearer $SUME_API_KEY" -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" -d @run-body.json | jq -r '.data.id') SLEEP=5 while :; do RUN=$(curl -sS "$SUME_API/v1/format-runs/$RUN_ID" -H "Authorization: Bearer $SUME_API_KEY") STATUS=$(echo "$RUN" | jq -r '.data.status') [ "$STATUS" = "queued" ] || [ "$STATUS" = "processing" ] || break sleep "$SLEEP"; SLEEP=$(( SLEEP < 60 ? SLEEP * 2 : 60 )) done echo "$RUN" | jq '{status: .data.status, primary: .data.primary_output_url, error: .data.error}' ``` In TypeScript, `subscribeFormatRun` does the create and the loop in one call. It resolves on all terminal statuses. Thus, a failed run is a result that you branch on, not an exception: ```ts import { readFile } from "node:fs/promises"; import { createSumeClient, subscribeFormatRun } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const run = await subscribeFormatRun({ client, path: { handle: "acme", slug: "live-commerce" }, idempotencyKey: "sheet-45-v1", body: JSON.parse(await readFile("run-body.json", "utf8")), timeout: 45 * 60_000, timeline: true, // also read the phase timeline on every poll onStatus: (status, snapshot) => console.log(status, snapshot.timeline?.at(-1)?.phase), }); if (run.status === "completed") console.log(run.primary_output_url); else console.error(run.status, run.error, run.output_error); ``` Its default timeout is 20 minutes. For long-form video, increase the timeout. Use `expires_at` on the receipt as the true ceiling. A timeout does not cancel the run. The run continues, and you continue to pay for it. Thus, keep the run id and read the run again at a later time. #### Read a Format before you call it This call helps in a settings screen or a preflight. With it, you can make sure that the address resolves for the key that you hold. You can also see what the Format takes and its cap. ```bash curl -sS "$SUME_API/v1/formats/acme/live-commerce" \ -H "Authorization: Bearer $SUME_API_KEY" \ | jq '.data | {id, handle, slug, version, status, api_trigger_enabled, io, cap_usd: (.generation_spend_cap_usd_micros / 1000000), vanity_invoke_url}' ``` You can get a `404 format_not_found` here with a key that you think is right. This almost always means that you need the other key. Team Formats answer only to keys created in the team workspace. A `403 workspace_key_required` tells you the same thing, with a more helpful message. You get it when the owner of your key is a member of that team. #### Next - [Create a run](/formats/call): all the fields of the body - [Runs and results](/formats/runs): the receipt, the poll rules, and the webhook rules - [Errors and spend](/formats/errors): what each code means and what to do - [Embed a Format in your product](/cookbooks/embed-a-format): key custody, spend tiers, and how to handle artifacts for a multi-tenant product ### Format catalog Source: https://docs.sume.com/formats/catalog.md Ready-made Sume Formats you can call from your own backend, with the scopes, curl, and run lifecycle for each. A [Format](/formats) is a saved authoring recipe that an Agent applies to a run. Sume ships a first-party catalog of Formats. Thus, a production workflow can be one HTTP call, not a prompt that you maintain. These pages are for the caller. They cover what the Format does, what to send, and how to read the result. They do not cover how to author a Format. The catalog answers at the reserved `sume` handle. You can call these slugs today: `sume-close-camera-ugc`, `sume-virtual-try-on`, `sume-product-usage-demo`, `sume-mobile-app-ugc`, `sume-before-after`, `sume-product-commercial`, `sume-beauty-studio`, `sume-logo-motion-design`, `sume-cinematic-studio-commercial`, `sume-fashion-editorial`, `sume-virtual-fitting`, `sume-water-splash-hero`, `sume-model-product-portrait`, `sume-magazine-cover-campaign`, `sume-formula-texture-hero`, `sume-sunscreen-splash`, `sume-cream-squeeze`, `sume-serum-drip`, `sume-toner-pour`, `sume-editorial-product-set`, `sume-slideshow`, `sume-wall-of-text`, `sume-green-screen`, `sume-video-hook`, `sume-fruits-drama`, `sume-recreate`, `sume-restyle`. Before you call a Format, read it with `GET /v1/formats/sume/{slug}`. This call returns its `description` and the `io` profile that the next section describes. Any slug not on this list answers `404 format_not_found` at `sume/{slug}`. #### Discover what a Format takes `GET /v1/formats` and `GET /v1/formats/{format_id}` return two fields that describe the Format, not the call: ```json { "io": { "profile": "url_to_video", "input_kind": "url", "output_kind": "video" }, "showcase": { "kind": "video", "media_url": "https://media.sume.com/...", "thumbnail_url": "https://media.sume.com/...", "created_at": "2026-08-01T00:00:00.000Z" } } ``` `io` is the Format's declared IO profile. `input_kind` is one of `url`, `text`, `image` or `product`. `output_kind` is one of `video`, `image` or `text`. Use it to select a Format from a list without a call, and to know the shape of `input` that the Format expects. By design, the run body's `input` is a free-form object, so this profile is the only declared contract between a Format's author and its callers. `showcase` is a worked example that the Format really produced during registration. Before Sume stores it, Sume verifies it against the generated-media ledger. Thus, it is output from a real run of this Format, not a picture that someone attached. **Both are `null` for Formats saved before registration existed.** That means "not declared", not "takes no input". In that case, use the Format's `description` and its page. #### Call the catalog directly The catalog is different in what the Format *knows*. The wire contract is the same as on [Calling a Format](/formats/call), with no changes. Formats by Sume answer at the reserved `sume` handle — `POST /v1/formats/sume/{slug}/runs` — and any key carrying `formats:write` can call one. The run, its media and its spend belong to the key that made the call. The catalog Format itself stays shared and unowned. Thus, you do not have to fork, install, or copy anything first. To *change* a catalog Format, fork it in the [Format library](https://www.sume.com/agents/format). Then the address of your copy is `{your_handle}/{slug}`. You call Formats that you author yourself in the same way. #### Call sheet for any Format Every Format with a handle and a slug also has a page on this site at the address that the API uses: ```text https://docs.sume.com/formats/{handle}/{slug} ``` The page has the scopes, curl, spend cap, poll and cancel sections for that Format. These pages are share and handoff links, not nav. Thus, the sidebar does not list them. They render for any well-formed address and never show the Format's body. Thus, if you give one to a partner, the partner learns only how to call the Format. #### Where to go next - [Format API](/formats) — what a Format is, and how the instruction is composed - [Calling a Format](/formats/call) — the full invoke contract and every error - [Bulk runs](/formats/bulk-runs) — queue many calls to the same Format - [Structured output](/formats/structured-output) — bind a schema and get typed JSON back - [Runs and results](/formats/runs) — polls, receipts, cancelation - [OpenAPI](https://api.sume.com/reference) — accurate request and response schemas ### Editing a Format package Source: https://docs.sume.com/formats/contents.md Read and commit the files behind your own Format over /v1, the same way you would use the GitHub Contents API. Each custom [Format](/formats) is a small package of files. The package has a `SKILL.md` entry file and optional notes under `references/`. The package also has a real commit history. With the Contents API and an API key, you can read those files and commit changes to them. Thus, a coding agent can edit a Format in the same way that it edits a repository. The shape of this API is the shape of the GitHub Contents API. We selected this shape on purpose. If you know `gh api repos/{owner}/{repo}/contents/{path}`, you know this API: ``` POST /v1/formats # create the Format GET /v1/formats/{handle}/{slug}/contents GET /v1/formats/{handle}/{slug}/contents/{path} PUT /v1/formats/{handle}/{slug}/contents # several files, one commit PUT /v1/formats/{handle}/{slug}/contents/{path} DELETE /v1/formats/{handle}/{slug}/contents/{path} ``` `{handle}/{slug}` is the same address that you use to invoke the Format. A read needs `formats:read`. A write needs `formats:write`. The author of the commit is always the owner of the key. The API ignores an `author` or `committer` in the body. This API changes the Format package. It does not run the Format. A change to a package has no effect on a run that already started. Each run reads the package that it started with. #### Create a Format `POST /v1/formats` opens a Format and its package repository. This is the same as how `POST /repos` opens a repository on GitHub. The call needs `formats:write`. The API creates the Format in **your key's workspace**. There is no `{handle}` in the address. Thus, you cannot create a Format in a workspace that is not yours. ```bash export SUME_API_KEY=sume_… # a workspace key; never commit it curl -sS -X POST "https://api.dev.sume.com/v1/formats" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "content-type: application/json" \ -d '{"slug":"plate-shots","title":"Plate shots","description":"Studio plate photography."}' ``` ```json { "data": { "id": "skl_…", "object": "format", "handle": "mobidoo", "slug": "plate-shots", "version": 1, "package_sha": "3f2b1c…", "contents_url": "https://api.dev.sume.com/v1/formats/mobidoo/plate-shots/contents", "vanity_invoke_url": "https://api.dev.sume.com/v1/formats/mobidoo/plate-shots/runs" } } ``` `auto_init` (default `true`) commits a minimal valid `SKILL.md`. Thus, you can read and write the Format through the endpoints below immediately. The intended next call replaces that file. `false` is a `400`: each Format package must contain `SKILL.md`, thus you cannot create an empty Format. The body does not accept package files, because the Contents API is for package files. All the files of a Format must obey one set of rules, for all writers. Keep `package_sha` and `contents_url` from the reply. The first value is the `If-Match` precondition for your next write. The second value is the address for that write. If your workspace already uses the slug, the call answers `409`. The call never overwrites a Format. The API opens the repository **before** the catalog row. If the API cannot reach the package history, the call fails with `503 format_git_unavailable`. In this condition, the API creates no Format. This prevents a Format with commits that no one can resolve. #### Read the package ```bash curl -sS "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents" \ -H "Authorization: Bearer $SUME_API_KEY" ``` ```json { "data": [ { "type": "dir", "name": "references", "path": "references", "sha": "9c1a…", "size": 0 }, { "type": "file", "name": "SKILL.md", "path": "SKILL.md", "sha": "0ee8…", "size": 4213 } ] } ``` A request for one path returns the file, base64-encoded: ```bash curl -sS "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents/SKILL.md" \ -H "Authorization: Bearer $SUME_API_KEY" ``` ```json { "data": { "type": "file", "name": "SKILL.md", "path": "SKILL.md", "sha": "0ee8784832cd08154436f264ddeb38703dc89cd7", "size": 4213, "encoding": "base64", "content": "LS0tCm5hbWU6IGxpdmUtY29tbWVyY2UK…" } } ``` **Keep the `sha`.** It is the git blob sha of the stored file. Each write uses it as the precondition. If the path is a directory (`…/contents/references`), the API returns the entries of that directory, not a file. #### Read the whole package at once If you list the root and then get each path, you send one request for each file. Sometimes you want the full package, for example to load it into the context of an agent. In that condition, ask for the package in one call: ```bash curl -sS "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents?recursive=1" \ -H "Authorization: Bearer $SUME_API_KEY" ``` ```json { "data": [ { "type": "file", "name": "SKILL.md", "path": "SKILL.md", "sha": "0ee8784832cd08154436f264ddeb38703dc89cd7", "size": 4213, "encoding": "base64", "content": "LS0tCm5hbWU6IGxpdmUtY29tbWVyY2UK…" }, { "type": "file", "name": "plan.md", "path": "references/plan.md", "sha": "3f7b1c0d9e2a4b6c8d0e1f2a3b4c5d6e7f809a1b", "size": 128, "encoding": "base64", "content": "IyBQbGFuCg==" } ] } ``` Each row is a file with its body, sorted by `path`. There are no `dir` rows, because the paths show the directories. `recursive=true` also works. All other values, and no value, give the plain root list above. The `sha` on each row is the same precondition that a write uses. Thus, one recursive read is sufficient before you start to edit. This parameter applies only to the list. `…/contents/{path}` gives the same answer with or without the parameter. #### Commit a change `PUT` writes one full file and makes one commit. It replaces the file. It does not patch the file. Send the full new body, base64-encoded. ```bash curl -sS -X PUT \ "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents/references/plan.md" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "message": "tighten gate 3", "content": "IyBQbGFuCg==", "sha": "3f7b1c0d9e2a4b6c8d0e1f2a3b4c5d6e7f809a1b" }' ``` ```json { "data": { "content": { "type": "file", "path": "references/plan.md", "sha": "5a2f…", "size": 7 }, "commit": { "sha": "b41c9a…", "message": "tighten gate 3", "tree": { "sha": "77de8a…" }, "parents": [{ "sha": "0a91f2…" }] }, "version": 12 } } ``` `commit.tree.sha` is the new package identity of the Format. It is the same value that the record of the Format shows. `version` is the display counter that the dashboard shows. If the path already exists, send `sha`. To create a new file, do not send it. Two agents that edit the *same file* cannot overwrite the work of each other without an error. After the first write, the `sha` of the second agent does not agree with the file. For agents that edit *different* files of one Format, refer to [`If-Match`](#guard-the-whole-package-with-if-match). `DELETE` takes `message` and `sha`. It answers with `content: null`: ```bash curl -sS -X DELETE \ "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents/references/plan.md" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{"message": "drop the old plan", "sha": "5a2f1b3c4d5e6f708192a3b4c5d6e7f809a1b2c3"}' ``` #### Commit several files at once If you edit a folder one path at a time, each file costs one commit and one round trip. `PUT` at the package root (with no `{path}`) takes a `files` list. It writes all the files as **one** commit, with one `version` increase and one new `package_sha`: ```bash curl -sS -X PUT \ "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "message": "edit plan + facts", "files": [ { "path": "references/plan.md", "content": "IyBQbGFuCg==", "sha": "3f7b1c0d9e2a4b6c8d0e1f2a3b4c5d6e7f809a1b" }, { "path": "references/facts.md", "content": "IyBGYWN0cwo=", "sha": "9b1d0e2f3a4b5c6d7e8f90a1b2c3d4e5f6a7b8c9" }, { "path": "references/new-note.md", "content": "IyBOZXcK" } ] }' ``` Each entry obeys the single-file rules. `content` is the full file, base64-encoded. `sha` is necessary when the path already exists. To create a path, do not send it. `content` in the reply is the list of files that this commit wrote. **`files` is a change set, not the package**. If the Format holds a path and this body does not name it, the API keeps that path as it is. Thus, an edit to two of twelve files never puts the other ten at risk. As a result, you cannot delete a file with this call. To remove a file, use `DELETE …/contents/{path}`. Thus, the API removes a file only when you ask for it, never because you forgot to name the file. If the `sha` of an entry is stale, or if the new package breaks a rule, the API commits **nothing**. The API commits the full batch or no part of it. #### Guard the whole package with `If-Match` A per-file `sha` cannot say "the package has not moved since I read it". Two agents can edit *different* files. Each agent holds a `sha` that is still current, thus the two writes land. But the second agent planned its edit against an old tree. That tree was not there at the time of the write. To close that gap, send the sha of the package. It is `commit.tree.sha` from your last write. It is also the same value that the Format record shows as `package_sha`: ```bash curl -sS -X PUT \ "https://api.dev.sume.com/v1/formats/mobidoo/live-commerce/contents" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "If-Match: 77de8a9b1c2d3e4f5061728394a5b6c7d8e9f001" \ -H "Content-Type: application/json" \ -d '{"message": "edit plan", "files": [{"path": "references/plan.md", "content": "IyBQbGFuCg==", "sha": "3f7b…"}]}' ``` If the package changed after your read, the API refuses the write with `409 format_package_sha_mismatch`. `error.details.package_sha` holds the current sha. Thus, you can read again and retry in a single round trip. The header also works on `PUT …/{path}` and `DELETE …/{path}`. It is an *additional* check, and the per-file `sha` checks still apply. Here, `If-Match` is not an opaque ETag. It is the package sha, as a bare 40-character hex string. Formats published before `package_sha` existed do not have a package sha. Those Formats ignore the header. They do not answer with a `409` that you can never satisfy. #### What a package may contain These are the same rules that the dashboard editor enforces: - `SKILL.md` at the package root is necessary. Its frontmatter `name` must be equal to the slug of the Format. You cannot delete it. - Files are at the root, or one directory level under `references/` or `agents/`. A path must not contain `..`. A path must not be absolute. - File names must match `^[A-Za-z0-9][A-Za-z0-9._-]*$`. A file name must not start with `_` or `.`. - Only `.md`, `.json`, `.yaml`, `.yml`, and `.txt` files are permitted. - The maximum is 100 MiB for each file and 100 MiB for each package. A Contents batch can name a maximum of 1000 paths. If a package breaks one of these rules, the API rejects it before it commits anything. For a rejected path, the API answers with `skill_path_invalid` and gives you the full allowlist. Thus, the first rejection gives sufficient data to correct the name. #### Errors worth handling | Status | `error.code` | What happened | |---|---|---| | `403` | `insufficient_scope` | The key does not have `formats:read` / `formats:write`, or it is a service-account key. A service-account key cannot create or edit packages. `next_action` is `authenticate`. You cannot patch scopes. Mint a new key. Missing scopes never give `format_not_found`. | | `409` | `skill_slug_taken` | Your workspace already has a Format with that slug. A create call never overwrites a Format. Select a different slug. | | `409` | `skill_slug_reserved` | A Format by Sume holds that slug for all workspaces. Select a different slug. | | `404` | `format_not_found` | The Format is unknown, or it is not in the workspace of the key. For a team Format, you need a key created in that workspace. | | `404` | `format_content_not_found` | The Format exists but holds nothing at that path. | | `409` | `format_content_sha_required` | The path already exists and you did not send a `sha`. Read the path. Then retry. | | `409` | `format_content_sha_mismatch` | Your `sha` is stale, because a different writer committed first. Read the path again. Then retry. | | `409` | `format_package_sha_mismatch` | The package sha in your `If-Match` is stale. `error.details.package_sha` holds the current package sha. | | `400` | `skill_path_invalid`, `skill_frontmatter_invalid`, `skill_limit_exceeded`, … | The new package broke a rule above. The API committed nothing. | | `503` | `format_git_unavailable` | The package history did not accept the commit, thus **the API saved nothing**. Retry. | We designed the last error on purpose. A write either commits or fails. The Format cannot change without a commit. #### What this is not There is no public git endpoint and no clone URL. You cannot reach the history writer directly. Only this API makes commits. At this time, this API does not let you revert commits, read blame, or browse the history. ### Embed a Format in your product Source: https://docs.sume.com/cookbooks/embed-a-format.md Run a Sume Format from your own dashboard on behalf of your customer — key custody, idempotency, webhooks, spend caps, artifacts, and failures. You have a product with your own customers. You want a button in *your* UI that makes a Sume-generated video or image for the customer who clicked the button. This page gives the end-to-end procedure for that. The flow is always the same: ```text your customer's browser │ (your own auth, your own request) ▼ your server ───── POST /v1/formats/{handle}/{slug}/runs ─────▶ Sume API ▲ │ │ POST /hooks/sume (signed, sume-v1) │ └────────────────────────────────────────────────────────────┘ ``` Your customer never talks to Sume. Your server holds one Sume API key. The server runs Formats for your customers and maps the results onto your own records. For the invoke contract, read [Calling a Format](/formats/call). To shape the result, read [Structured output](/formats/structured-output). This page gives the integration around those two topics. #### 0. The whole flow With [`@sume-com/sdk`](/sdk) (`0.2.0+`), the happy path is **`subscribeFormatRun`**: one call creates the Format run and waits for the terminal receipt. In production, we recommend a webhook (§4). Sume has no SSE stream yet. `events_url` (`/v1/format-runs/{id}/events`) is a polled phase timeline. Thus, `onStatus` gives the status from a poll, not from a log feed. ```ts import { createSumeClient, subscribeFormatRun } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const run = await subscribeFormatRun({ client, // Your team's vanity path — the same {handle}/{slug} the Format detail page shows. // Team Formats need a team (workspace) key — see Calling a Format. path: { handle: "acme", slug: "product-promo" }, idempotencyKey: runKey(customer, order), body: { input: { product_url: order.productUrl }, generation_spend_cap_usd: 3, }, onStatus: (status, snapshot) => console.log(status, snapshot.next_action), }); if (run.status === "completed") { await attachOutputs(order.id, run); // run.primary_output_url, run.artifacts } ``` The SDK makes the work easier, but you do not have to use it. Each call here is one HTTP request that you can make with `fetch`. The [API reference](/api/reference) stays the source of truth for fields. The SDK gives you the loops that nobody likes to write: `subscribeFormatRun` / `waitForRun` and `verifyWebhook`. In production, use the push path. Refer to §4 below, and to [Waiting for runs](/sdk/runs) for the tradeoff. #### 1. Key custody **One Sume account and one server-side key serve many of your customers.** Sume has no per-end-user credential that you can give to a customer. Sume also has no browser-safe key. | Rule | Why | |---|---| | Keep the key in your server's environment. Never put the key in client JavaScript, a mobile bundle, or a `NEXT_PUBLIC_*` variable. | A Sume key spends *your* credits. Anyone who has the key can run any Format that you own, up to your caps. | | Never proxy the key. Proxy the *call*. | A "pass-through" endpoint forwards the browser's payload and attaches your key. That is the same leak, one hop later. Your endpoint must accept your customer's identifiers and make the Sume request itself. | | Give your endpoint your own authorization check. | Sume authenticates you, not your customer. Your product must decide if *this* customer can run *that* Format. | | To rotate a key, create a new key and retire the old key. | You cannot add scopes to an existing key. Refer to the text below. | Create the key at [API Keys](https://www.sume.com/dashboard/api-keys) with the `formats:read` and `formats:write` scopes. **Keys that Sume minted before the release of Format API triggers do not have those scopes.** You cannot add the scopes later. An old key fails every run with `403 insufficient_scope`. Create a new key. Service-account keys cannot create any Format runs. Their requests fail with `details.reason` of `service_account_format_runs_unsupported`. ```ts // server-only module. Importing this from a client component is the bug. import { createSumeClient, createFormatRunByVanityPath } from "@sume-com/sdk"; if (!process.env.SUME_API_KEY) throw new Error("SUME_API_KEY is not configured"); const client = createSumeClient({ apiKey: process.env.SUME_API_KEY }); export async function startRunForCustomer(customer: Customer, order: Order) { // 202 on a fresh run, 200 on an idempotent replay. Both carry the receipt. const { data, error } = await createFormatRunByVanityPath({ client, path: { handle: "acme", slug: "product-promo" }, headers: { "idempotency-key": runKey(customer, order) }, body: { input: { product_url: order.productUrl }, generation_spend_cap_usd: spendCapForPlan(customer.plan), communication: { mode: "webhook", webhook_url: "https://acme.example.com/hooks/sume", }, }, }); if (error) throw new Error(`Sume rejected the run: ${JSON.stringify(error)}`); return data!.data; } ``` The client sends `x-api-key`. Do not add your own `Authorization` header. If a request has both credentials, the API rejects it with `401 unauthorized`. Without the SDK, this call is a plain `POST` to `https://api.sume.com/v1/formats/{handle}/{slug}/runs`. Refer to [Calling a Format](/formats/call). #### 2. Derive `Idempotency-Key` per customer Your customers will double-click. Your job queue will redeliver. Each of these events makes two paid runs. To prevent this, derive the key from the item that the run makes. Do not derive the key from the time of the request. ```ts import { createHash } from "node:crypto"; /** Stable for one (customer, order, format version) — not for one HTTP call. */ function runKey(customer: Customer, order: Order) { return createHash("sha256") .update(`${customer.id}:${order.id}:product-promo:v1`) .digest("hex") .slice(0, 40); } ``` | Do | Do not | |---|---| | Hash your own stable identifiers: tenant id, order id, Format slug, and a version. Bump the version only when you want a re-run. | `uuidv4()` per request. Then the header has no effect. | | Namespace the key by customer. | A key that you build only from the order id. Two tenants with the same ids then share a run. | | Before you return to the browser, store the returned `run_id` with your record. | Derive the key again later to find the run. You can do this, but a stored id is one lookup, not one replay. | These are the accurate replay semantics: | Replay | Result | |---|---| | Same key, same body | `200` with the **original** run receipt and `idempotency_hit: true`. Sume does not start a second run or make a second charge. | | Same key, different body (a different `instruction` is also a different body) | `409 idempotency_conflict`. Nothing runs. | | No key | Every call starts a new paid run. | #### 3. Pick a spend cap per run Every Format has a generation spend cap. A run can never spend more than its own effective cap. If a run names no cap, the run inherits the Format's cap. `generation_spend_cap_usd` on the run request names the ceiling of that run, up to the platform maximum of $500. Sume accepts a number above the Format's own cap. A number above $500 gets a `400`. Thus, the cap is the natural place for your own plan tiers: ```ts function spendCapForPlan(plan: Plan) { switch (plan) { case "free": return 0.5; case "pro": return 3; case "enterprise": return undefined; // inherit the Format's own cap } } ``` Read the Format's own cap from `PublicFormat.generation_spend_cap_usd_micros`. This value is always a number. If a Format never named a cap, the default is $400. The receipt gives the effective cap of a run as `usage.generation_spend_cap_usd_micros`. Caps set a limit on *generation* spend. The terminal receipt reports the actual spend of the run against that ceiling as `usage.billable_amount_usd_micros`. This value is sufficient to show a per-run cost in your own UI, but it does not include the agent's own LLM turn. Thus, it is not the total cost of the run, and it is not an invoice. Bill your customer from your own records. Reconcile against [`GET /v1/usage`](/dashboard/usage). #### 4. Receive the result A Format run is asynchronous. You have two ways to know when the run is complete. Both ways give the same receipt. | | Webhook | Poll | |---|---|---| | You do | Register `communication.webhook_url`, verify the signature, and return `2xx`. | Loop on `status_url` until the run is terminal. | | Available | `api.dev.sume.com` and `api.sume.com`. | Everywhere. | | Costs you | One public HTTPS endpoint. | One timer per in-flight run. | Build the webhook receiver first. Keep the poll path (or [`subscribeFormatRun`](/sdk/runs)) wired as the backup for a time when your endpoint is down. For the full contract, refer to [Run webhooks](/agents/run-webhooks). For a complete receiver in Node and Python, refer to [Cookbook](/formats/cookbook#a-webhook-receiver). ##### Verify every delivery Sume signs the raw body with HMAC-SHA256 over `.`. Sume sends `sume-v1=` in `x-sume-webhook-signature`. Verify the signature **before** you parse the body. ```ts import { verifyWebhook } from "@sume-com/sdk"; export async function handleSumeWebhook(req: Request) { const raw = await req.text(); // raw string, not a re-serialized object const secret = process.env.SUME_COM_WEBHOOK_SIGNING_SECRET!; const ok = await verifyWebhook({ body: raw, headers: req.headers, secret }); if (!ok) return new Response("bad signature", { status: 401 }); const event = JSON.parse(raw); if (event.event !== "format.run.terminal") return new Response(null, { status: 204 }); // Dedupe on request_id — it repeats across retries of the same run. await recordTerminalRun(event.request_id, event); return new Response(null, { status: 204 }); // fast 2xx, work afterwards } ``` `verifyWebhook` is async because it runs on WebCrypto. WebCrypto lets you use it from Workers and Deno, and also from Node. For options and header names, refer to [Verifying webhooks](/sdk/webhooks). You can also write the check yourself, in a different language or at a gateway in front of your app. That check is a dozen lines against a published scheme: [Run webhooks](/agents/run-webhooks). Read your signing secret on the **Webhooks** tab of the dashboard (`/dashboard/webhooks`). Or, get it from `GET /v1/webhooks/signing-secret` with any API key that has `account:read`. Sume derives the secret for your workspace. Thus, a valid signature proves that Sume signed the delivery for you, not for anyone who has a shared secret. Store the secret as `SUME_COM_WEBHOOK_SIGNING_SECRET`, the name that Sume's delivery worker uses. Store it in the same way as the API key. Four items frequently cause problems for integrators here: - **Verify against the raw body.** Some frameworks parse JSON for you and give you an object. Such a framework already destroyed the signed bytes. In Express, mount `express.raw({ type: "application/json" })` on this route only. - **Return `2xx` fast, then work.** The delivery attempt budget is 10 seconds. If a receiver renders video before it responds, Sume retries the delivery while the receiver works. - **Dedupe on `request_id`.** Retries send the same value again. Ten attempts against an unreliable endpoint must not become ten rows in your database. - **`3xx` is not a delivery.** Sume does not follow redirects. Register the final URL, not a redirector, and not an HTTP URL. Sume rejects non-HTTPS, localhost, and private-range URLs at submit with `400 invalid_request`, and checks them again at delivery time. ##### Job webhooks are a different surface If you also call `POST /v1/models/...` directly, those calls emit **generation-job** webhooks (`job.completed` and related events) with `job_id`. [Webhooks](/workflows/webhooks) gives the details. These webhooks have different events, a different payload, and a different lifecycle. The signature scheme is the same, so one verifier covers both. But route on `event`, and never assume that a body has `run_id`. A single receiver for both must first switch on the event name. The receiver must send a `204` for any event that it does not know. Thus, a new event type does not cause a 500 and a retry storm. #### 5. Map artifacts into your UI A terminal `completed` receipt has three fields that contain media: | Field | Use it for | |---|---| | `primary_output_url` | The one item to show. It is `null` when the Format produced no single primary file. | | `artifacts[]` | All the items that the run generated: `{ id, type, url, content_type, size_bytes, width, height, duration_ms, checksum_sha256 }`. | | `output` | The Format's structured result, projected onto `output_schema`. Media in the result points to the same URLs. Refer to [Structured output](/formats/structured-output). | Every URL is a durable `media.sume.com` HTTPS URL. **These URLs do not expire.** Thus, an embed is practical. You can store the URL with your record and render it forever, with no refresh step. Design your integration for these two results: - **A durable URL is a public URL.** Anyone who has the URL can fetch it. The URL will go into your logs, your error reports, and your customer's browser history. If your product must never let customer A see customer B's output, proxy the bytes through your own authenticated route. Or, at receipt time, copy the bytes into your own storage and serve them from there. - **Copy, or link, but decide.** A link is free and instant. A copy costs you storage, but the copy stays available if you stop the use of Sume. If you want that guarantee, copy on the webhook, before you mark the record ready. ```ts async function attachOutputs(orderId: string, receipt: FormatRunReceipt) { const videos = receipt.artifacts.filter(a => a.type === "video"); await db.orders.update(orderId, { previewUrl: receipt.primary_output_url, assets: videos.map(a => ({ sumeArtifactId: a.id, url: a.url, contentType: a.content_type, durationMs: a.duration_ms, })), }); } ``` `artifacts[]` is empty until the run is terminal. Sume fills it from the same job ledger that fills `output`. Thus, the two fields always agree. #### 6. Failure taxonomy Runs fail in four different places. Your UI must have a different message for each place. If you show only "something went wrong" for all four, you will quickly get a support ticket that you cannot answer. ##### At submit — nothing ran, nothing was charged | Code | Status | What it means for your integration | |---|---|---| | `insufficient_scope` | 403 | Your key does not have `formats:read` / `formats:write`, or it is a service-account key. Correct your key, not your request. | | `format_not_found` | 404 | Unknown handle, unknown slug, or a Format that this key does not own. A Format by Sume also returns this code at an *account* handle. Such a Format answers at `sume/{slug}`. | | `format_not_forkable` | 409 | You addressed a built-in capability, not a Format. Call one of the Formats by Sume, or one of your own Formats. | | `format_api_trigger_disabled` | 409 | The API trigger is off for that Format. | | `format_inactive` | 409 | The Format is inactive. | | `format_run_in_progress` | 409 | Occurs only with `on_active_run: "reject"`. Retry later, or show "already running". | | `idempotency_conflict` | 409 | Same key, different body. Your key derivation is not stable. Correct it before you retry. | | `invalid_request` | 400 | Includes a `webhook_url` that is not a public HTTPS URL. | A 4xx here is a bug in your call, not a transient error. If you retry an `insufficient_scope` forever, you make a frequent and expensive mistake. ##### At run — a run existed, and did not produce a result `status` is `failed`, and the receipt has `error` and `output_error`. Important codes: | `error.code` | Meaning | |---|---| | `unattended_blocked` | The run hit a gate that it could not pass without a human. Examples: no avatar that matches, or a missing input that the run asks about in chat. Sume writes the message so that you can show it. | | `format_run_failed` | The generic failure. Read `error.message`. A run that tried to spend more than its cap also gets this code. Thus, if the runs of a plan tier fail again and again, first compare them with `usage.generation_spend_cap_usd_micros`. | To retry a failed run, use a **new** idempotency key. The old key is bound to the run that failed. If you use the old key again, you get the same failed receipt. **API runs are unattended.** A Format for interactive chat can pause to ask a human for approval. Over the API, those approvals are pre-granted, and the run continues in its spend cap. Thus, `completed` is a real result. You will never get a half-finished run that shows as done. ##### Terminal, but not a failure | `status` | Handle it as | |---|---| | `canceled` | Someone called `POST /v1/format-runs/{id}/cancel`. **Sume sends no webhook.** Use the cancel response and poll `status_url`. | | `skipped` | You passed `on_active_run: "skip"`, and a run was already in flight. **A skipped run never delivers a webhook.** The create response already told you, with `skip_reason` populated. Read the status from that response. Do not wait for a POST. The Format default is **`allow`** (concurrent runs). The Action default is **`skip`**. Do not copy Action examples into Format calls. Refer to [Calling a Format](/formats/call#request-body). | ##### At delivery — the run is fine, your endpoint was not A delivery result never changes the run. After ten refused attempts, you have a failed *delivery* and a run that is still `completed`. Fetch the run from `result_url`. One delivery case needs code. A receipt over **1 MiB** arrives with `payload: null` and `error.code` of `payload_too_large`. That delivery gives the `result_url`, and you fetch the receipt from it. If a handler assumes that `payload` is an object, it will throw on your largest, most valuable runs. ```ts const receipt = event.payload ?? (await fetchRun(event.error.result_url)); ``` #### Checklist before you ship - [ ] `SUME_API_KEY` is server-only and is not in any client bundle. - [ ] Your run endpoint authorizes your own customer before it calls Sume. - [ ] You derive `Idempotency-Key` from stable identifiers, and you do not generate it per request. - [ ] `generation_spend_cap_usd` is set per plan tier. - [ ] The webhook receiver verifies `sume-v1` against the raw body and returns `2xx` in less than a second. - [ ] Your receiver dedupes deliveries on `request_id`. - [ ] For `payload: null` (oversized receipt), your handler uses `result_url`. - [ ] You set `SUME_COM_WEBHOOK_SIGNING_SECRET` from the value on `/dashboard/webhooks`, and its fingerprint matches `x-sume-webhook-secret-fingerprint` on a delivery. - [ ] Poll on `status_url` stays wired as a backup (webhooks are live on `api.sume.com`). - [ ] Every error code above maps to a message that your support team can use. #### Next - [TypeScript SDK](/sdk) — the client that this page uses, from install to first run - [Calling a Format](/formats/call) — the invoke contract - [Bulk runs](/formats/bulk-runs) — a server-side queue of those runs - [Structured output](/formats/structured-output) — schemas, the projection, failure modes - [Runs and results](/formats/runs) — the receipt, field by field - [Run webhooks](/agents/run-webhooks) — full details of delivery, signatures, and retries - [Webhooks](/workflows/webhooks) — generation-job webhooks, the other surface ## API ### Overview Source: https://docs.sume.com/public-api.md Current Sume Developer API surface for catalog, jobs, usage, media inputs, and generation workflows. The Sume Developer API is the public server-side API for `api.sume.com`. It is workspace-scoped and key-authenticated. It is for generation workflows that must have durable jobs, public media artifacts, usage tracking, and dashboard observability. #### Base URL ```text https://api.sume.com/v1 ``` All endpoint paths on this page include `/v1`, because the API service root serves the live OpenAPI schema. #### Product, API, and media domains | Domain / URL | Purpose | |---|---| | `https://www.sume.com` | Public company/product site. | | `https://www.sume.com/dashboard` | Dashboard home. | | `https://www.sume.com/dashboard/api-keys` | Create and manage Developer API keys. | | `https://www.sume.com/dashboard/jobs` | Examine jobs. | | `https://www.sume.com/dashboard/usage` | Usage and balance summary. | | `https://www.sume.com/dashboard/subscription` | Billing & subscription (plans and credit top-ups). | | `https://www.sume.com/pricing/api` | Public metered API rate card (source of truth for prices). | | `https://www.sume.com/playground` | Avatar playground (human experiments). | | `https://www.sume.com/agents` | Agents product surface. | | `https://api.sume.com/v1` | Public Developer API. | | `https://api.sume.com/reference` | Swagger UI. | | `https://api.sume.com/reference/json` | Live OpenAPI JSON (schema source of truth). | | `https://media.sume.com` | First-party media and artifact URLs that completed jobs return. | #### Authentication Create an API key in the [API Keys dashboard](https://www.sume.com/dashboard/api-keys). Send the key from a server-side environment. ```bash export SUME_API_KEY="sume_live_..." curl https://api.sume.com/v1/me \ -H "Authorization: Bearer $SUME_API_KEY" ``` The API also accepts `x-api-key: sume_live_...`. Do not send workspace or user identifiers in request bodies. Sume resolves the scope from the API key. #### Start with the Format API Most partner integrations are one call to a saved recipe, not a chain of model invocations that you write yourself. [Format API](/formats) runs an authored recipe in a sandbox. It returns durable media and JSON in a schema that you supply. Refer to [Structured output](/formats/structured-output). The model endpoints below are the layer below the Format API. Call them directly when you want only one model invocation and you control the orchestration yourself. #### TypeScript clients From Node, Bun, Deno, or Workers, [`@sume-com/sdk`](/sdk) (`@sume-com/sdk@0.2.0`) wraps each operation on this page with generated types from the same OpenAPI schema. It also adds the helpers that you must otherwise write yourself: `subscribeFormatRun` (Format create + wait), `waitForRun`, `waitForJob` (generation jobs), `uploadFile`, and `verifyWebhook`. Start with `subscribeFormatRun` + [webhooks](/agents/run-webhooks). There is no SSE stream. Thus, you get the progress when you poll `events_url` (a phase timeline). ```bash npm install @sume-com/sdk ``` The SDK is a convenience layer, not a second contract. This page and the [API reference](/api/reference) stay the source of truth for fields. #### Current endpoint map This table is only a navigation summary. The **accurate request/response schemas** come from live OpenAPI (`https://api.sume.com/reference/json`). For method/path notes, we recommend the readable [API reference](/api/reference). For workflow prose, we recommend the model guides. Do not use duplicated Markdown tables as a second schema. For new integrations, we recommend **canonical product paths** (`/v1/{family}-1.0/...`). Sume continues to support the `/v1/models/sume/.../runs` aliases for compatibility. | Area | Canonical | Compatibility / aliases | Use for | |---|---|---|---| | Formats | `GET /v1/formats`, `GET /v1/formats/:handle/:slug`, `POST /v1/formats/:handle/:slug/runs`, `GET /v1/formats/:handle/:slug/runs`, `POST /v1/formats/:handle/:slug/bulk-runs`, `GET /v1/format-runs/:id`, `GET /v1/format-runs/:id/status`, `GET /v1/format-runs/:id/result`, `GET /v1/format-runs/:id/events`, `POST /v1/format-runs/:id/cancel`, `POST /v1/format-runs/:id/webhook/redeliver`, `GET /v1/format-run-queues/:id` | `GET /v1/formats/:id`, `POST /v1/formats/:id/runs`, `POST /v1/formats/:id/bulk-runs` (opaque `skl_…` twins) | Run a saved recipe (one run, or a bulk queue). Get the receipt by webhook or poll. Read the media and the schema-shaped JSON. Refer to [Format API](/formats), [Create a run](/formats/call), [Runs and results](/formats/runs), [Errors and spend](/formats/errors), and [Bulk runs](/formats/bulk-runs). | | Health | `GET /v1/health` | — | Service readiness checks. | | Catalog | `GET /v1/catalog` | — | Find capabilities, models, runtime readiness, and price metadata. | | Account | `GET /v1/me` | — | Make sure that the API key works, and get the resolved workspace context. | | Balance and usage | `GET /v1/balance`, `GET /v1/usage` | — | Read the USD balance and the usage ledger entries. | | Jobs | `GET /v1/jobs`, `GET /v1/jobs/:id`, `GET /v1/jobs/:id/status`, `GET /v1/jobs/:id/result`, `POST /v1/jobs/:id/cancel`, `GET /v1/jobs/:id/events` | — | List, inspect, poll, cancel, recover, and audit jobs. | | Avatar 1.0 | `POST /v1/avatar-1.0/generate`, `POST /v1/avatar-1.0/talking-video`, `GET /v1/avatar-1.0/avatars`, `GET /v1/avatar-1.0/avatars/:id` | `POST /v1/models/sume/avatar/v1.0/runs`, `POST /v1/models/sume/avatar-1.0/generate/runs`, `POST /v1/models/sume/avatar-1.0/talking-video/runs`, `GET /v1/avatars`, `GET /v1/avatars/:id` | Create and read avatars / talking videos. | | Avatar Video 1.0 | `GET /v1/avatar-videos`, `GET /v1/avatar-videos/:id` | `POST /v1/models/sume/avatar-video/v1.0/runs` | Product/scene avatar-video runs and resource reads. | | Avatar Video Previews | `POST /v1/avatar-video-previews`, `GET /v1/avatar-video-previews/:id`, `POST /v1/avatar-video-previews/:id/regenerate`, `POST /v1/avatar-video-previews/:id/generate-video` | — | Preview → generate-video flow. | | Avatar catalog | `POST /v1/avatar-catalog/search` | — | Search reusable catalog avatars. | | Avatar Face Swap (Beta) | — | `POST /v1/models/sume/avatar-face-swap/v1.0/runs` | Face-swap model runs. | | Image 1.0 | `POST /v1/image-1.0/generate` | `POST /v1/models/sume/image-1.0/runs` | Image generation. | | Video 1.0 | `POST /v1/video-1.0/generate` | `POST /v1/models/sume/video-1.0/runs` | Video generation. | | Music Router | `POST /v1/music-router/generate` | — | Music generation. `sume/music-auto` selects the engine. | | Music 1.0 | `POST /v1/music-1.0/generate` | `POST /v1/models/sume/music-1.0/runs` | Sume will retire it. It resolves through Music Router. | | Video captions | `POST /v1/video-captions`, `GET /v1/video-captions/:id` | — | Caption jobs and resource reads. | | Trending videos | `POST /v1/trending-videos/search` | — | Search TikTok trending video metadata. | | Actions | `GET /v1/actions`, `GET /v1/actions/:id`, `POST /v1/actions/:id/runs`, `GET /v1/action-runs/:id`, `POST /v1/action-runs/:id/cancel` | — | Recurring schedules. Refer to [Scheduled](/agents/actions). | | Agent Completions | `POST /v1/agent/completions`, `GET /v1/agent-runs`, `GET /v1/agent-runs/:id`, `POST /v1/agent-runs/:id/cancel` | — | Run the Agent on an ad-hoc task. Refer to [Agent Completions](/agents/completions). | For the full method/path tables and the OpenAPI hide-list notes, refer to the [API reference](/api/reference). Older `www.sume.so` consumer-product routes are not part of the current `sume.com` developer platform API. Examples of these routes are `/credits`, `/uploads/presign`, `/brand`, `/ads/videos`, `/face-swap`, and `/reference-analysis`. Internal voice capabilities, raw provider model IDs, and provider task URLs are not public API surfaces, unless they are in `/v1/catalog` and the OpenAPI schema. #### Job-first workflow Model-run and compatibility submit endpoints accept work, create a job, and return a response that you can poll or recover later. ```text submit generation (canonical or /models/.../runs) -> receive job id -> poll /v1/jobs/:id/status -> fetch /v1/jobs/:id/result when completed -> use media.sume.com artifact URLs ``` We recommend that most integrations store the job ID and poll with backoff. Use webhooks if you have a public HTTPS callback endpoint. With webhooks, Sume sends terminal job events to your server. Paid generation uses queue-first admission. Workspace concurrency limits are applicable to jobs that are in the `processing` state. While the queue has capacity, Sume can still accept valid submissions as `queued`. For tier limits, `generation_limits`, and queue-full behavior, refer to [Generation admission](/workflows/generation-admission). #### OpenAPI The local docs preview serves a snapshot at: ```text /api/openapi.json ``` Production serves the live schema at: ```text https://api.sume.com/reference/json ``` Swagger UI is available from the API service at: ```text https://api.sume.com/reference ``` Use the live schema as the accurate source of truth for requests and responses. A sync refreshes the docs repo snapshot from that endpoint (refer to README / `pnpm openapi:sync`). ### Authentication Source: https://docs.sume.com/authentication.md How Sume Developer API keys work and how to send them safely. Sume uses workspace-scoped API keys to authenticate Developer API requests. You create these keys in the [API Keys dashboard](https://www.sume.com/dashboard/api-keys). #### Send an API key Use a server-side environment variable: ```bash export SUME_API_KEY="sume_live_..." ``` Then send either Bearer auth: ```bash curl https://api.sume.com/v1/me \ -H "Authorization: Bearer $SUME_API_KEY" ``` or the API key header: ```bash curl https://api.sume.com/v1/me \ -H "x-api-key: $SUME_API_KEY" ``` The current API accepts both forms. In each integration, use one form consistently. The Sume CLI uses `x-api-key` by default. If you configure it, the CLI can use Bearer mode. **Send exactly one**. If a request has both `Authorization: Bearer` and `x-api-key`, the API rejects it with `401 unauthorized` and the message `Send only one API key credential.` Neither header wins. The second header does not silently hide the first. This problem occurs with gateways and `fetch` wrappers that add their own `Authorization` header to a client that already sends `x-api-key`. Remove one of the two headers. Do not expect a precedence rule, because that rule does not exist. #### Scope Sume gets the workspace, owner, and API key metadata from the key. Do not put `workspace_id`, `owner_user_id`, or `user_id` in public API request bodies. Responses show key metadata such as id, name, prefix, scopes, and last-used time. Responses do not show the full secret. When you create a key, Sume fixes its scopes. You cannot add scopes later. A key that you created before a scope existed does not have that scope. This rule is important for `actions:read` and `actions:write`. Calls to [Scheduled via the Actions API](/agents/actions/api-trigger) must have these scopes. Sume mints them only on keys that you created after the Actions API-call trigger shipped. An older key returns `403 insufficient_scope` on each Action run request. Create a new key, then rotate to it. The same problem applies to `formats:read` / `formats:write`. If a pre-Formats key calls a Format, the result is `403 insufficient_scope`, not `404 format_not_found`. There is no API to add scopes to a key that already exists. #### Server-side proxy pattern Browser and mobile clients must call your backend. Your backend must attach the Sume API key. ```ts export async function POST(request: Request) { const body = await request.json(); const response = await fetch("https://api.sume.com/v1/avatar-1.0/generate", { method: "POST", headers: { Authorization: `Bearer ${process.env.SUME_API_KEY}`, "Content-Type": "application/json", "Idempotency-Key": crypto.randomUUID(), }, body: JSON.stringify(body), }); return new Response(await response.text(), { status: response.status, headers: { "Content-Type": "application/json" }, }); } ``` Validate the user input before you forward requests to Sume. Also enforce your own authorization before you forward the requests. #### Rotation Create a replacement key. Deploy the new key to your server. Use `GET /v1/me` to verify the new key. Then revoke the old key from the dashboard. If a key shows in logs or chat history, rotate it. #### Safety rules - Keep API keys on trusted servers, CI secret stores, or local developer machines. - Do not put API keys in frontend JavaScript, mobile apps, support tickets, or screenshots. - Signed upload and download URLs are temporary secrets. Protect them. - If a key is exposed, rotate keys from the dashboard. - Give agents read-only commands first. Make explicit confirmation mandatory before write or paid generation commands. #### Rate limits Each API key gets a request budget per minute for all of `/v1`. The subscription plan of the workspace that owns the key sets this budget. **Reads and writes have separate budgets**. Thus, a tight status-poll loop cannot cause a 429 on your own submits. | Plan | Writes per minute | Reads per minute | |---|---:|---:| | Free | 120 | 4800 | | Pro | 300 | 12000 | | Startup | 600 | 24000 | | Scale | 1200 | 48000 | | Enterprise | Contact sales | Contact sales | A **read** is any `GET` or `HEAD`, for example a poll of `status_url`, `events_url`, or `result_url`, or a list of Formats or runs. The two POSTs that submit nothing are also reads: `/v1/generation/admission-preview` and the MCP endpoint itself. All other requests are **writes**: run creation, cancellation, and uploads. The plan number is the write number. Reads get **forty times** that number in their own bucket. We sized that multiple for agents, not for a person who monitors one run. An agent harvest can hold twenty-odd jobs open and poll each of them. That is thousands of reads a minute for work that costs nothing. Thus, reads are deliberately cheap. The write budget is the tier that a plan actually buys, and we do not change it. Enterprise is not self-serve. Until Sume provisions a contracted number, an Enterprise key uses the Scale row above. An MCP tool call spends the write budget one time, for the run that it creates. It does not spend the write budget for the JSON-RPC request that carried it. A `jobs_status` poll over MCP does not spend any write budget. Do not count requests yourself. Read `ratelimit-remaining`. On `retry-after`, wait before you send more requests. The headers describe the budget that the current request spent from. A `429` names that budget in `error.details.scope` (`read` or `write`). Request rate is not the same as generation capacity. The concurrency limit of your plan controls the number of generations that run at the same time. The `generation_limits` object reports this limit. If you increase your request rate, the concurrency limit does not increase. Each response shows the current state: | Header | Meaning | |---|---| | `ratelimit-limit` | Requests permitted in the current window. | | `ratelimit-remaining` | Remaining requests in the current window. | | `ratelimit-reset` | Seconds until the window resets. | | `retry-after` | Seconds to wait, sent on `429`. | Sume limits unauthenticated requests per client IP at the Free rate. Their read bucket stays at **four times** the write rate, not forty. The agent-sized read budget is for callers who own the jobs that they poll. A larger anonymous bucket only makes the abuse surface larger. The read multiple is a deployment configuration value (`SUME_COM_API_RATE_LIMIT_READ_MULTIPLIER`). Thus, a self-hosted or preview deployment can be different. `ratelimit-limit` on the response is always the authority for the deployment that you send requests to. The table above shows the shipped default. #### Common failures | Status | Common cause | Next step | |---|---|---| | `401` | Missing, malformed, or revoked key. | Examine the header. If necessary, create a new key. | | `403` | The key is valid, but it is not permitted to access the requested surface. | Examine the workspace membership and the key scope. | | `429` | Your requests went above the requests-per-minute budget of your plan. | Wait for the `retry-after` time, then try again. | ### Core concepts Source: https://docs.sume.com/workflows/core-workflow.md The main objects and lifecycle behind the Sume developer platform. This page is the **job and admission** map: catalog, jobs, artifacts, usage. For a product map of Agents, Formats, Models, and clients, refer to [Sume basics](/the-basics). Sume generation workflows use a small set of public objects: catalog items, media inputs, jobs, artifacts, usage entries, and dashboard resources. #### Platform domains | Domain / URL | Role | |---|---| | `www.sume.com/dashboard` | Dashboard home and operator surfaces. | | `www.sume.com/dashboard/api-keys` | API key management. | | `www.sume.com/dashboard/jobs` | Job inspection. | | `www.sume.com/dashboard/usage` | Usage and balance summary. | | `www.sume.com/dashboard/subscription` | Billing & subscription / credit top-ups. | | `www.sume.com/playground` | Avatar playground. | | `api.sume.com` | Public Developer API (`/v1`) and OpenAPI (`/reference/json`). | | `media.sume.com` | First-party generated media artifacts. | #### Catalog `GET /v1/catalog` is the best first call for programmatic discovery. It lists available capabilities, model ids, endpoint paths, runtime readiness, and pricing metadata. #### Jobs Generation requests create durable jobs. A job tracks request metadata, status, public provider model, result, public error, events, and webhook delivery state. The supported statuses are: ```text queued -> processing -> completed queued -> processing -> failed queued -> canceled ``` #### Generation lifecycle ```text client submit -> Sume validates API key and request -> optional usage reservation -> provider-backed execution -> artifact mirroring to media.sume.com -> usage capture or refund -> terminal webhook delivery when configured -> result available from /v1/jobs/:id/result ``` If a client disconnects or times out locally, keep the job id. Then recover with the jobs API. Do not submit duplicate paid work. #### Artifacts Completed jobs can include public artifacts under `https://media.sume.com`. Sume-owned artifact URLs are the public contract. Raw provider URLs are not. #### Media inputs Launch generation requests accept public HTTPS media URLs in schema-defined fields, for example `input.image_url`, `product_image`, `scene.image_url`, and `video_url` (face-swap / captions). Normal integrations do not need a separate asset upload step. The API returns generated outputs as Sume-hosted artifacts under `media.sume.com`. Refer to [Media inputs](/workflows/asset-library). #### Usage Provider-backed generation can reserve estimated usage in USD. It can capture the actual cost on success. It can refund on failure, or on cancellation before capture. Use `/v1/balance` and `/v1/usage` to examine the current balance and the ledger entries. #### Dashboard The dashboard mirrors the API surfaces for humans: - [API keys](https://www.sume.com/dashboard/api-keys) - [Jobs](https://www.sume.com/dashboard/jobs) - [Usage](https://www.sume.com/dashboard/usage) - [Billing & subscription](https://www.sume.com/dashboard/subscription) - [Playground](https://www.sume.com/playground) ### API reference Source: https://docs.sume.com/api/reference.md Route map, authentication notes, and response patterns for the current Sume Developer API. This page is a **human-readable route map** for the Sume Developer API. This page is not a second schema. For exact JSON request/response shapes, enums, and field requirements, use the live OpenAPI document. The Markdown tables on this page can be older than the live document. #### OpenAPI schema Local docs snapshot: ```text /api/openapi.json ``` Live schema (source of truth for this docs snapshot): ```text https://api.sume.com/reference/json ``` Download the schema: ```bash curl https://api.sume.com/reference/json \ -o sume-openapi.json ``` Refresh the checked-in snapshot locally: ```bash pnpm openapi:sync # or: node scripts/sync-openapi.mjs ``` #### Authentication You must send a Sume API key to all `/v1` API endpoints, except these public routes: `GET /v1/health`, `GET /v1/catalog`, `GET /v1/openapi.json`, `GET /v1/bgm/catalog`, `GET /v1/bgm/categories`, and `POST /v1/bgm/pick`. ```bash curl https://api.sume.com/v1/me \ -H "Authorization: Bearer $SUME_API_KEY" ``` The API also accepts `x-api-key: $SUME_API_KEY`. API responses do not return the full secret key. #### Account, catalog, and usage | Method | Path | Notes | |---|---|---| | `GET` | `/v1/health` | Versioned API health. | | `GET` | `/v1/catalog` | Available capabilities, endpoints, runtime readiness, models, and pricing metadata. | | `GET` | `/v1/me` | Current API key, owner, and workspace context. | | `GET` | `/v1/balance` | The available balance in USD. | | `GET` | `/v1/usage` | Usage ledger entries such as reservations, captures, refunds, and top-ups. | #### Jobs | Method | Path | Notes | |---|---|---| | `GET` | `/v1/jobs` | List workspace jobs. You can filter by status/type. | | `GET` | `/v1/jobs/:id` | Read the public job envelope. | | `GET` | `/v1/jobs/:id/status` | Lightweight status read for polling. | | `GET` | `/v1/jobs/:id/result` | Completed result payload. It includes public artifact URLs when they are available. | | `POST` | `/v1/jobs/:id/cancel` | Cancel a job before generation starts. After generation starts, the endpoint returns `409 job_generation_already_started`. Idempotent on an already-canceled job. | | `GET` | `/v1/jobs/:id/events` | Public timeline events. Use them to debug and to recover jobs. | Terminal job statuses are `completed`, `failed`, and `canceled`. Non-terminal statuses are `queued` and `processing`. #### Media inputs Generation requests accept fetchable public HTTPS media URLs directly in the fields that the live OpenAPI schema documents (for example Avatar photo `input.image_url`, Avatar Video `product_image` / `scene.image_url`). The API rejects localhost, private-network, non-HTTPS, and non-image responses before it submits the generation. First-party upload helpers are **not** the default public API path for normal integrations (refer to [Hidden from public OpenAPI](#hidden-from-public-openapi)). #### Canonical generation paths For new integrations, these product-style endpoints are the best choice. | Method | Path | Family | Notes | |---|---|---|---| | `POST` | `/v1/avatar-1.0/generate` | Avatar 1.0 | Canonical avatar create. | | `POST` | `/v1/avatar-1.0/talking-video` | Avatar 1.0 | Canonical talking-video create. | | `GET` | `/v1/avatar-1.0/avatars` | Avatar 1.0 | List Avatar 1.0 avatars. | | `GET` | `/v1/avatar-1.0/avatars/:id` | Avatar 1.0 | Read one Avatar 1.0 avatar. | | `POST` | `/v1/image-1.0/generate` | Image 1.0 | Canonical image generate. | | `POST` | `/v1/video-1.0/generate` | Video 1.0 | Canonical video generate. | | `POST` | `/v1/music-1.0/generate` | Music 1.0 | Canonical music generate. | #### Compatibility model-run aliases These `/v1/models/sume/.../runs` paths stay in the public OpenAPI and continue to work. When both exist, the canonical paths above are the better choice. | Method | Path | Public model | Notes | |---|---|---|---| | `POST` | `/v1/models/sume/avatar/v1.0/runs` | `sume/avatar/v1.0` | Legacy Avatar 1.0 create. | | `POST` | `/v1/models/sume/avatar-1.0/generate/runs` | `sume/avatar-1.0/generate` | Alias of Avatar 1.0 generate. | | `POST` | `/v1/models/sume/avatar-1.0/talking-video/runs` | `sume/avatar-1.0/talking-video` | Alias of Avatar 1.0 talking video. | | `POST` | `/v1/models/sume/avatar-video/v1.0/runs` | `sume/avatar-video/v1.0` | Legacy Avatar Video 1.0 create. | | `POST` | `/v1/models/sume/avatar-face-swap/v1.0/runs` | `sume/avatar-face-swap/v1.0` | Avatar Face Swap 1.0 Beta. | | `POST` | `/v1/models/sume/image-1.0/runs` | `sume/image-1.0` | Alias of Image 1.0 generate. | | `POST` | `/v1/models/sume/video-1.0/runs` | `sume/video-1.0` | Alias of Video 1.0 generate. | | `POST` | `/v1/models/sume/music-1.0/runs` | `sume/music-1.0` | Alias of Music 1.0 generate. | Where the OpenAPI schema documents them, submit endpoints support the common communication fields `mode`, `webhook_url`, and `wait_timeout_seconds`. Avatar creation uses a top-level `avatar_handle` and an `input` union: `prompt`, `props`, or `photo`. Avatar Video uses a top-level `avatar_handle` and exactly one of `script` or `video_inputs`. Avatar Video accepts `quality: "standard" | "plus" | "max"`. The default is **`plus`**. Deep guides: [Avatar overview](/models), [previews](/models/avatar-video-previews), [face swap](/models/face-swap), [captions](/models/video-captions), [trending](/models/trending-videos). #### Avatar resources, previews, catalog, captions, trending | Method | Path | Notes | |---|---|---| | `GET` | `/v1/avatars` | List avatar resources (compatibility list path). | | `GET` | `/v1/avatars/:id` | Read one avatar resource. | | `GET` | `/v1/avatar-videos` | List avatar-video resources. | | `GET` | `/v1/avatar-videos/:id` | Read one avatar-video resource. | | `POST` | `/v1/avatar-catalog/search` | Search the avatar catalog. | | `POST` | `/v1/avatar-video-previews` | Create an avatar-video preview. | | `GET` | `/v1/avatar-video-previews/:id` | Read a preview. | | `POST` | `/v1/avatar-video-previews/:id/regenerate` | Regenerate a preview. | | `POST` | `/v1/avatar-video-previews/:id/generate-video` | Generate a video from a preview. | | `POST` | `/v1/video-captions` | Submit a video caption job. | | `GET` | `/v1/video-captions/:id` | Read a video caption resource. | | `POST` | `/v1/trending-videos/search` | Search TikTok trending video metadata. | For terminal event payloads, signature headers, and retry behavior, refer to [Webhooks](/workflows/webhooks). #### Actions Agents Actions have their own run resource and status vocabulary. Actions are not jobs, and they do not show in `/v1/jobs`. The key must have `actions:read` for all routes, except the two writes. For the two writes, the key must have `actions:write`. | Method | Path | Notes | |---|---|---| | `GET` | `/v1/actions` | List Actions. Filters: `limit`, `status`, `trigger_type`. | | `GET` | `/v1/actions/:action_id` | Read one Action. | | `GET` | `/v1/actions/:action_id/runs` | List runs for an Action. | | `POST` | `/v1/actions/:action_id/runs` | Start a run through the API-call trigger. `actions:write` is necessary. | | `GET` | `/v1/actions/:action_id/runs/:run_id` | Read one run. This path is an alias under its Action. | | `GET` | `/v1/action-runs/:run_id` | Read a run receipt. | | `GET` | `/v1/action-runs/:run_id/status` | Trimmed status payload for polling. | | `GET` | `/v1/action-runs/:run_id/result` | Terminal receipt. Returns `409 run_not_completed` while the run is in progress. | | `POST` | `/v1/action-runs/:run_id/cancel` | Idempotent cancel. `actions:write` is necessary. | There is no public endpoint to create, edit, or delete an Action. There is also no `/v1/action-runs/:run_id/events` endpoint. The `events_url` field on run receipts is always `null`. For the request body, idempotency rules, and the full error table, refer to [Advanced: run a schedule via API](/agents/actions/api-trigger). #### Formats Formats have their own run resource, parallel to Actions. The key must have `formats:read` for all routes, except the writes (create a run, create a bulk-run queue, cancel). For the writes, the key must have `formats:write`. | Method | Path | Notes | |---|---|---| | `GET` | `/v1/formats` | List the Formats that the key can see: your own Formats and the first-party catalog. `limit` filter. | | `GET` | `/v1/formats/:format_id` | Read one Format. The endpoint does not return the `SKILL.md` body. | | `GET` | `/v1/formats/:format_id/runs` | List runs for a Format, newest first. | | `POST` | `/v1/formats/:format_id/runs` | Start a run through the API-call trigger. `formats:write` is necessary. | | `POST` | `/v1/formats/:format_id/bulk-runs` | Queue a maximum of 100 runs with a `concurrency` window (1–16). `formats:write` is necessary. `202` queue receipt. | | `GET` | `/v1/formats/:handle/:slug` | Read a Format that you own, by handle and slug. | | `GET` | `/v1/formats/:handle/:slug/runs` | List runs, addressed by handle and slug. | | `POST` | `/v1/formats/:handle/:slug/runs` | Start a run, addressed by handle and slug. `formats:write` is necessary. | | `POST` | `/v1/formats/:handle/:slug/bulk-runs` | Same bulk queue, addressed by handle and slug. `formats:write` is necessary. | | `GET` | `/v1/format-run-queues/:queue_id` | Bulk-queue progress (`counts` + per-item status). `formats:read` is necessary. | | `GET` | `/v1/format-runs/:run_id` | Read a run receipt. | | `GET` | `/v1/format-runs/:run_id/status` | Trimmed status payload for polling. | | `GET` | `/v1/format-runs/:run_id/result` | Terminal receipt. Returns `409 run_not_completed` while the run is in progress. | | `GET` | `/v1/format-runs/:run_id/events` | Phase timeline for one run, oldest first. | | `POST` | `/v1/format-runs/:run_id/cancel` | Idempotent cancel. `formats:write` is necessary. | | `POST` | `/v1/format-runs/:run_id/webhook/redeliver` | Re-POST the terminal `format.run.terminal` receipt to the webhook URL of the run. `formats:write` is necessary. | To author a Format, use the Agents dashboard. You can also use the [Contents API](/formats/contents) to create and edit the package of a Format. Any valid key can call the curated Formats by Sume directly at `sume/{slug}`. Sume bills the run to that key. For the request body and the full error table, refer to [Calling a Format](/formats/call). For the queue contract, refer to [Bulk runs](/formats/bulk-runs). For polling and receipts, refer to [Runs and results](/formats/runs). For `format.run.terminal` delivery, refer to [Run webhooks](/agents/run-webhooks) (not generation-job [Webhooks](/workflows/webhooks)). For schema rules, refer to [Structured output](/formats/structured-output). For ready-made Formats, refer to the [Format catalog](/formats/catalog). #### Agent Completions An Agent Completion runs the Agent on an ad-hoc prompt with nothing saved. For reads, the key must have `agent_completions:read`. For the two writes, the key must have `agent_completions:write`. | Method | Path | Notes | |---|---|---| | `POST` | `/v1/agent/completions` | Start a completion. Async only. Returns `202` and a receipt, not `choices[]`. | | `GET` | `/v1/agent-runs` | List completions, newest first. | | `GET` | `/v1/agent-runs/:run_id` | Read a run receipt. | | `GET` | `/v1/agent-runs/:run_id/status` | Trimmed status payload for polling. | | `GET` | `/v1/agent-runs/:run_id/result` | Terminal receipt. Returns `409 run_not_completed` while the run is in progress. | | `POST` | `/v1/agent-runs/:run_id/cancel` | Idempotent cancel. | For the request shape and the differences from an OpenAI chat completion, refer to [Agent Completions](/agents/completions). #### Hidden from public OpenAPI The API **implements** some routes, but intentionally does **not** include them in the public OpenAPI document (`hidePreLaunchCompatibilityOpenApiPaths`). Do not use these routes as a documented public contract until they are in `https://api.sume.com/reference/json` again. | Hidden path family | Status | |---|---| | `/v1/assets`, `/v1/assets/upload-url`, `/v1/assets/:id`, `/v1/assets/:id/complete`, `/v1/assets/:id/download-url` | Implemented. Hidden from public OpenAPI. In generation requests, it is better to use public HTTPS media URLs. | | `/v1/generation/admission-preview` | Implemented. Hidden from public OpenAPI. [Generation admission](/workflows/generation-admission) describes the admission behavior for paid jobs. | | `POST /v1/avatars`, `POST /v1/avatar-videos` | Create through POST on these resource paths is hidden. Use the canonical / model-run submit endpoints. | | `/health` (unversioned) | Hidden. Use `GET /v1/health`. | | `/v1/models/{model_owner}/{model_name}/{model_version}/runs` | The generic template path is hidden. Use the concrete model paths that the tables above list. | Related asset-library workflow notes can still describe URL-first inputs when OpenAPI does not list the upload helpers. #### Result and artifact shape Completed jobs can include public artifacts: ```json { "id": "job_...", "status": "completed", "result": { "artifacts": [ { "id": "artf_...", "url": "https://media.sume.com/artifacts/...", "type": "image", "content_type": "image/png" } ] } } ``` Public results must use `media.sume.com` URLs. Raw provider URLs and provider task URLs are not part of the public result contract. #### Error envelope Errors use a consistent envelope. The request id is inside `error`. ```json { "error": { "code": "invalid_request", "message": "Invalid request body, parameters, or headers.", "request_id": "req_..." } } ``` Keep the request id for support. Redact API keys, signed URLs, private media URLs, user ids, workspace ids, and raw provider identifiers from logs. ### Generation admission Source: https://docs.sume.com/workflows/generation-admission.md How Sume admits paid generation jobs, applies workspace concurrency, queues work, and reports queue status. Sume generation APIs use queue-first admission for paid work that runs for a long time. A submit request creates a durable job when three conditions are true. The request is valid, Sume can reserve the balance, and the workspace still has capacity for accepted jobs. The job can start immediately, or it can wait in `queued` until a workspace concurrency slot opens. ```text submit -> queued -> processing -> completed | failed | canceled ``` Concurrency is a dispatch limit, not a submit limit. If your workspace is already at its generation concurrency limit, Sume can still accept more jobs as `queued`. This is true while queue capacity remains. Later, workers move queued jobs to `processing` under the concurrency guard of each workspace. #### Limits at a glance Sume separates four controls that are easy to confuse: | Control | Applies to | What happens when full | |---|---|---| | Generation concurrency | Paid generation jobs with status `processing`. | Sume can still accept new valid jobs as `queued` if queue capacity remains. | | Queue capacity | Paid generation jobs that Sume accepted and that did not start yet. | New paid generation submissions fail with `429 queue_full`. | | Submit rate limits | Request volume for public API submit endpoints. | Requests fail with `429 rate_limited`. Retry with backoff and an idempotency key. | | Balance and reservation | Spendable USD balance for the authenticated workspace. | Generation submit fails with `402 insufficient_credits` before provider work starts. | Read/status/list endpoints can also have rate limits. Treat those limits as poll backpressure, not generation concurrency. #### Processing concurrency Generation concurrency is **plan-only**. Prepaid top-ups do **not** increase the processing concurrency limit. Admin overrides can increase the effective `concurrency_limit` (`limit_source: admin_override`). The default queue capacity is `max(3, concurrency_limit × 5)`. | Plan | Processing concurrency | Queue capacity (default) | Accepted job capacity | |---|---:|---:|---:| | Free | 1 | 5 | 6 | | Pro | 4 | 20 | 24 | | Startup | 8 | 40 | 48 | | Scale | 20 | 100 | 120 | | Enterprise | 20 | 100 | 120 | The dashboard **Concurrency** tab is the source of truth for the configured processing cap of the workspace. The API shows this cap as `generation_limits.concurrency_limit`. Org workspaces have a floor of 10. Enterprise has a default of 20 and uses admin overrides for higher contract limits. Always prefer the effective field to this static table. `accepted job capacity` is `concurrency_limit + queued_jobs_limit`. This value is the maximum number of paid generation jobs that can be `processing` or `queued` for the workspace at the same time. #### Queue-first behavior If your workspace has `concurrency_limit: 1`, you can submit several valid jobs at the same time. Sume can return all of them as `queued` while balance and queue capacity are available. Only one generation job of the same workspace can move to `processing` at a time. This is the intended behavior: ```text Job A: queued -> processing -> completed Job B: queued -------------> processing -> completed Job C: queued ---------------------------> processing -> completed ``` Do not treat `queued` as a failure. Store the `job_id`. Poll the status with backoff. Fetch the result only when the job reports `result_ready: true` or `status: completed`. #### Immediate rejection Sume rejects a request immediately only when it cannot safely accept the request. | Status | Code | Why it happens | Client behavior | |---|---|---|---| | `400` | `invalid_request` | The request body, model id shape, mode, webhook options, or headers are not valid. | Correct the request before you retry. | | `401` | `unauthorized` | The API key is missing, malformed, revoked, or not valid. | Correct the authentication. | | `402` | `insufficient_credits` | Sume cannot reserve the estimated generation cost from the workspace balance. | Upgrade the plan / wait for included Gen$, or submit a less expensive request. Do not invent prepaid top-ups. | | `404` | `model_not_found` or `not_found` | The public model or resource does not exist in this workspace. | Use `/v1/catalog`, or make sure that the ids are correct. | | `409` | `idempotency_conflict` | A client used the same idempotency key again for a different operation or payload. | Use a key again only for an exact retry. | | `429` | `queue_full` | The workspace has no remaining accepted generation capacity. | Wait for jobs to finish, or cancel queued jobs. Then retry with the same idempotency key. | | `429` | `rate_limited` | API request volume exceeded an abuse-protection limit. | Use `retry-after` for the backoff when it is present. | | `503` | `provider_capacity_exceeded` or runtime configuration errors | Sume cannot start or dispatch generation work safely. | Retry later with the same idempotency key, unless the error tells you not to retry. | Full concurrency alone is not an error. It becomes a submit error only when the queue is also full. #### `generation_limits` Generation submit responses include `generation_limits` when Sume can compute the workspace admission snapshot. ```json { "generation_limits": { "plan_id": "pro", "limit_source": "plan", "plan_concurrency_limit": 4, "concurrency_limit": 4, "queued_jobs_limit": 20, "accepted_generation_jobs_limit": 24, "active_generation_jobs": 0, "queued_generation_jobs": 0, "queue_capacity_remaining": 24, "wave_size_hint": 18 } } ``` The fields have these meanings: | Field | Meaning | |---|---| | `plan_id` | Subscription plan that sets the default concurrency map. | | `limit_source` | `plan` or `admin_override` for the effective concurrency. | | `plan_concurrency_limit` | Plan-default processing concurrency (when an override applies, do not use it to calculate the wave size). | | `concurrency_limit` | **Effective** maximum number of paid generation jobs in the same workspace that can be `processing`. | | `queued_jobs_limit` | More paid generation jobs in the same workspace that can wait in `queued`. | | `accepted_generation_jobs_limit` | `concurrency_limit + queued_jobs_limit`. | | `active_generation_jobs` | Current generation jobs in the same workspace with status `processing`. | | `queued_generation_jobs` | Current generation jobs in the same workspace with status `queued`. | | `queue_capacity_remaining` | The remaining queued-job budget plus the idle processing seats, before `queue_full`. | | `wave_size_hint` | Submission-wave hint only: `max(1, floor(queue_capacity_remaining * 0.75))`. Not a concurrency limit, override, or processing width. Never use it to calculate the size of in-flight work. | The counts are a snapshot. They can change immediately after the response, when workers claim jobs or other clients submit work. #### Sizing in-flight work Use `max(0, concurrency_limit - active_generation_jobs - queued_generation_jobs)` as the budget for new in-flight work. Limit that budget to `queue_capacity_remaining`. Count each newly submitted job against that budget until the next live snapshot. At zero headroom, wait and refresh the preview before you submit more. If the counts are not available, refresh before you select a width. This pace control on the client keeps open work in the processing cap. The API can still accept queued work under its separate queue-first admission policy. For example, `concurrency_limit: 100`, `queued_jobs_limit: 500`, and no active or queued jobs give `queue_capacity_remaining: 600` and `wave_size_hint: 450`. The workspace is set to **100**, with a maximum of 100 new in-flight jobs in this snapshot. **450 is only a submission-wave hint**, and it includes queue slots. With 30 processing jobs and 10 queued jobs, the new in-flight budget is 60. Never show the hint as concurrency. Do not use the pre-override `plan_concurrency_limit` / `purchased_concurrency_limit` fields in its place. A full queue still means that you must wait, although the minimum value of the hint is 1. #### Before bulk submissions For launch integrations, use `GET /v1/balance` and the `generation_limits` from generation submit responses to make conservative queue decisions. Later, the live OpenAPI can show a read-only admission preview endpoint for your environment. If it does, treat that endpoint only as an optional preflight. It must not create a job, reserve credits, capture usage, refund usage, or call generation providers. #### Submit and poll pattern For production integrations, prefer async submit with an idempotency key. ```bash curl -X POST https://api.sume.com/v1/avatar-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: avatar-batch-001-item-001" \ -d '{ "avatar_handle": "studio_presenter", "input": { "type": "prompt", "prompt": "Friendly studio presenter" }, "mode": "async" }' ``` Then poll the status: ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" ``` Fetch the result after completion: ```bash curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Recommended client behavior: - Treat `queued` and `processing` as normal non-terminal states. - Use exponential backoff for polls. Do not use tight loops across many jobs. - Continue to poll until `terminal: true`, or until your own application deadline. - Use `Idempotency-Key` for each paid submit that a client can retry. - Do not submit a paid request again only because your local worker timed out. - Store `status_url`, `result_url`, `events_url`, and `cancel_url` when they are present. - Examine `generation_limits`. When queue capacity is low, do not add more work. The Sume CLI uses the same model: ```bash sume avatars create --confirm-paid --avatar-handle studio_presenter --type prompt --prompt "Friendly studio presenter" --json sume jobs status job_123 --agent --json sume jobs result job_123 --agent --json sume jobs events job_123 --agent --json ``` #### Queue-full handling `queue_full` means that the workspace used all of its accepted generation capacity: ```json { "error": { "code": "queue_full", "message": "Workspace generation queue is full. Wait for running jobs to complete before submitting more generation work.", "request_id": "req_..." } } ``` The error details can include a `generation_limits` snapshot and job metadata for the failed admission attempt. When applicable, Sume releases or refunds the reservation for the failed admission. When you receive `queue_full`: - Do not add more generation work for that workspace. - Poll the current jobs until at least one job reaches a terminal state. - Cancel the queued jobs that you no longer need. - Retry with the same idempotency key after capacity opens. - Use `retry-after` when it is present. #### Cancellation and billing Paid generation uses public Sume USD estimates. At submit time, Sume reserves the estimated amount when it accepts the request. A successful completion captures the reserved usage. Where applicable, failed jobs and failed queue admission release or refund the reservation. Cancellation succeeds only before generation work starts: ```bash curl -X POST https://api.sume.com/v1/jobs/job_123/cancel \ -H "Authorization: Bearer $SUME_API_KEY" ``` Cancel the queued jobs that you no longer need, before they start to process. After generation starts, cancel returns `409 job_generation_already_started` (`details.cancelable: false`). The job then completes or fails normally. A cancel of a job that is already `canceled` is idempotent. #### Edge cases and current boundaries - Sume currently shows queue counts and the remaining accepted capacity. It does not show a precise queue position or ETA for each job. - `sync` and `subscribe` modes can wait for a maximum of 30 seconds. When the wait budget ends, continue to poll the job id. - Queue expiration and an explicit client-supplied fail-fast queue length are not currently public API options. If they are not in the live OpenAPI schema, treat them as future contract additions. - Public API responses are provider-neutral. They do not show hidden provider names, raw provider task ids, raw provider URLs, or internal workflow names. They also do not show storage object keys, API keys, or private workspace/user metadata. ### Jobs and results Source: https://docs.sume.com/workflows/jobs-and-results.md How to inspect Sume jobs, poll status, fetch results, read events, and request cancellation. Sume generation endpoints create durable jobs. Store the job id from submit responses, so that your integration can recover work after process restarts. For paid generation, `queued` is a normal accepted state. Workspace concurrency limits apply when workers move jobs into `processing`, not when the API accepts valid jobs. Refer to [Generation admission](/workflows/generation-admission) for queue capacity, tier limits, and queue-full errors. #### Statuses | Status | Meaning | Terminal | |---|---|---| | `queued` | Sume accepted the request, and the job waits to run. | No | | `processing` | The job runs, or Sume finalizes it. | No | | `completed` | The result is ready. | Yes | | `failed` | The job reached a terminal failure with a public error. | Yes | | `canceled` | A cancellation request arrived, and the job is terminal. | Yes | #### Who can read a job A job belongs to its workspace and to the member whose key or Agent turn created it. Sume records its usage against that member. Only that member can cancel it. The reads (`GET /v1/jobs/:id`, `/status`, `/result`, `/events`, and `GET /v1/jobs`) obey one rule: - An API key reads the jobs that its own member created in the workspace of the key. - A Studio Agent turn reads each job in the thread that it runs on, whoever created the job. This includes the jobs that an API-fired Format run made in that thread. Thus, a teammate who continues the thread can read the results of the run. An example is the voice and model of each narration job. An interactive turn can also read the jobs of its own member in sibling threads. An unattended run stays on its thread. - All other reads get `404 not_found`: jobs in other workspaces, the jobs of other members in other threads, and, for an API key, any job that its member did not create. A `thread_id` filter makes a list smaller. It never increases what a key can read. When a turn reads a job that another member created, `webhook_delivery.url` and `webhook_delivery.last_error` are `null`. The turn sees the delivery state, but not the callback endpoint of the owner. #### Poll status ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" ``` Use exponential backoff. Stop the polls on `completed`, `failed`, or `canceled`. Do not submit the original paid request again only because a local process timed out. #### MCP `jobs_wait` (single or batch) Remote MCP `jobs_wait` accepts one of these: - `job_id` — single-job wait. The response shape does not change (`object: "job_wait"`). - `job_ids` — 1–20 ids with optional `wait_for: "all" | "any"` (default `all`). The response is `object: "job_wait_batch"`, with a status snapshot for each requested id. After parallel fan-outs, prefer one batch wait to N single waits. Each remote HTTP `POST /mcp` caller, and also Studio Agent threads, holds a maximum of 55s for each call (omitted default 50). On `wait_slice_expired`, retry `jobs_wait` with the same ids. Never submit the paid create again. There is no longer a thread-header 600s server-side hold. An HTTP request that stays open for that long dies at the edge (`502` / `Transport send error`) before it can answer. `wait_for: "any"` still reports each id. The remaining jobs continue and still bill. Unknown or foreign-workspace ids cause the full call to fail. If you pass `include_results: true`, each completed id comes back with its `jobs_result` answer in `results[]`. These are the same entries that a batch `jobs_result` returns. Thus, a wave needs no separate result read. `results_omitted.job_ids` names the results that do not fit in one answer. Read those results with one batch `jobs_result`. `outcome: "operator_stopped"` means that Sume operations stopped at least one id (`error_code` `ops_*`, `public_reason` `job_stopped_by_operations`). Those jobs are terminal and produce no output. Sume refunded their holds. The wait answers immediately when it sees one stopped id. `operator_stopped.pending_job_ids` names the ids that still run. If you issue the wait again on the stopped ids, the answer cannot change. ##### Slices are bounded, and the server enforces it On remote MCP, the default of `timeout_seconds` is **50**, and its cap is **55**. (The API accepts values up to 600 and clamps them.) A wait returns at the moment that its job is terminal. An image that still runs at 50s is usually stuck, not slow. A wait is one HTTP request that stays open for the full slice, and no data moves during that time. Each edge closes such a request at some time. The caller then gets no tool result, while the job continues to run and to bill. The API clamps a larger `timeout_seconds` and does not reject it. The response tells you about the clamp in `wait_slice_clamped`. To wait for a ten-minute render, do the wait again. Do not ask for a longer wait. A `524` (or `522` / `523` / `525`) on `jobs_wait` is a transport failure, never a job outcome. Issue `jobs_wait` again on the same ids, or read `jobs_status` one time. Do not submit the paid create again. Do not report the job as blocked. #### Fetch a result ```bash curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Results are available only after completion. If the job is not complete yet, the API returns a conflict response. It does not return an empty result. ##### Reading a whole wave (MCP) Over MCP, `jobs_result` also takes `job_ids`, with the same 1–20 ceiling as `jobs_wait`. Thus, if you waited on a wave in one call, you can read it back in one call, not in N calls: ```json { "job_ids": ["job_a", "job_b", "job_c"] } ``` The response is a `job_result_batch`: `results[]` in request order, with one entry for each id. Each entry has `ok` and either `value` or a typed `error`. Partial success is normal and intentional. An id that still runs comes back as `job_not_completed`, and each finished id still returns its result. `partial_failure.failed_job_ids` names exactly the ids that need a second read. Read `ok` for each entry. A failure on one id tells you nothing about the other ids. #### Read events ```bash curl https://api.sume.com/v1/jobs/job_123/events \ -H "Authorization: Bearer $SUME_API_KEY" ``` Events give a public timeline that you can use to debug and to recover work: - `job.created` - `job.queued` - `job.started` - `generation.submitted` - `job.completed` - `job.failed` - `job.canceled` - `webhook.delivery` Public events do not show raw provider task ids or raw provider URLs. #### Request cancellation ```bash curl -X POST https://api.sume.com/v1/jobs/job_123/cancel \ -H "Authorization: Bearer $SUME_API_KEY" ``` Cancellation succeeds only before generation work starts. After generation starts, the API returns `409 job_generation_already_started` with `details.cancelable: false`, and the job runs to completion. A cancel of a job that is already `canceled` is idempotent. It returns the same canceled job. #### Communication modes Each submit endpoint accepts a `mode`. The mode decides **how you learn the outcome**. It never changes whether Sume creates a job, what the job costs, or how long the job takes to run. | Mode | HTTP returns | Job id in the first response | Server blocks | What the client does next | |---|---|---|---|---| | `async` (default) | `202` with the job envelope and poll URLs | Yes | No | Poll `status_url` until `terminal` is true. Then `GET result_url` when `result_ready` is true. | | `sync` | The same envelope, after a wait of up to `wait_timeout_seconds` (max **30**) for a terminal transition | Yes | Yes, at most 30s. Less when waiter capacity is not available. | Terminal? Read the job from the response. Not terminal? **Poll. Do not resubmit.** | | `subscribe` | The same as `sync`: the same bounded wait | Yes | Same as `sync` | Same as `sync`. For a fal-style long wait, use the client subscribe recipe below with `async`. | | `webhook` | `202` with the job envelope and poll URLs. Sume stores the callback. | Yes | No | Wait for the terminal callback, and verify its signature. Continue to poll as a backup. | If you omit `mode`, you get `async`. If you send `webhook_url` (or its alias `callback_url`) without a `mode`, you get `webhook`. Sume accepts a submit at the moment that it has a durable job id. Thus, **every mode returns the job id in its first response.** A `2xx` means that the job exists and paid work is in flight. It does not mean that the job finished. To know the difference, read `terminal` and `result_ready` from the envelope. ##### `sync` and `subscribe` are aliases They run the same bounded waiter and return the same envelope. Sume keeps `subscribe` because clients ported from other queue APIs use it. It is not a long-lived subscription, an event stream, or a longer wait. There is no SSE or WebSocket transport on the Developer API today. `GET /v1/jobs/:id/events` is a pull snapshot, not a stream. ##### "Subscribe" means three different things The word occurs on three unrelated surfaces with three different waits. None of them is a push stream. Read the table before you set a timeout. | Where you see it | What it is | How long it waits | |---|---|---| | Job `mode: "subscribe"` (this page) | An alias of `sync`. One bounded HTTP wait on the submit call. | At most `wait_timeout_seconds`, with a cap of **30s**. | | SDK `subscribeFormatRun()` ([TypeScript SDK](/sdk)) | Create the run, then poll it client-side to a terminal receipt. | Minutes. The SDK uses its own timeout, not an HTTP hold. | | Format / Action / Agent `communication.mode` | Delivery selection for a **run**, not a job. The values are `async` and `webhook`. `subscribe` is not one of them. | Nothing blocks. | Two results are important: - If you send `mode: "subscribe"`, you do **not** get progress events. You get the same 30-second wait that `sync` gets. For progress, submit `async` and read `GET /v1/jobs/:id/events`, or use a [webhook](/workflows/webhooks). - `communication.mode` has no `subscribe` value, and its two values operate identically. Only a supplied `webhook_url` makes delivery active. For new integrations, we recommend **`async`** (poll, or read events) or **`webhook`** (Sume tells you). `sync` and `subscribe` stay supported, and Sume will not remove them. But they are the wrong tool for work that can last longer than 30 seconds. Most video work is in that group. ##### 30 seconds is a wait budget, not a job duration The API clamps `wait_timeout_seconds` to `0..30`. It sets a limit on how long the **HTTP request** blocks, not on how long the **job** can take. Image jobs often finish in that time. Video, avatar-video, and face-swap jobs usually do not. When the budget ends, or when the API process has no waiter capacity left and skips the wait fully: 1. The response is still `2xx` and still carries the job id. The end of the wait budget is not an admission failure. 2. The envelope carries `status_url`, `result_url`, `events_url`, `cancel_url`, and a `sync` object. `sync.timed_out` is true when the wait returned before a terminal state. `sync.capacity_exhausted` is true when Sume skipped the wait because the waiter budget of the process was full. 3. You **must** continue with `GET status_url`. When `next_poll_after_seconds` is present, obey it. If it is not present, do a backoff. 4. You **must not** submit a new paid job for the same intent. You can retry the submit itself. Use the **same** `Idempotency-Key` again, so that the retry returns the original job and does not bill a second job. `sync` is null on `async` and `webhook` responses. ##### Client subscribe: poll a job to a terminal state This is the official equivalent of a client-side `subscribe()`. It is the correct answer for work that can last longer than 30 seconds. Submit with `async`, poll, then read the result. The wait lives in **your** client. Thus, its timeout can be minutes, and no HTTP request stays open. ```text job = POST /v1/{product}/generate { mode: "async", ... } with Idempotency-Key loop: s = GET /v1/jobs/{job.id}/status onStatus(s) # optional progress callback if s.terminal: break sleep(s.next_poll_after_seconds or exponential backoff) if s.sume_status == "completed": return GET /v1/jobs/{job.id}/result else: # failed or canceled raise from (GET /v1/jobs/{job.id}).job.error ``` Poll on the booleans (`terminal`, `result_ready`) or on `sume_status`. The status endpoint also returns a queue-shaped `status` field (`IN_QUEUE` / `IN_PROGRESS` / `COMPLETED` / `FAILED` / `CANCELED`) for clients ported from other queue APIs. This field maps one-to-one onto `sume_status`. The two fields always agree, but do not mix them. `GET /v1/jobs/:id/result` is only for completed jobs. For other jobs, it answers `409 job_not_completed`. Thus, read the failure from the job record. Submit (step 1): ```bash curl -X POST https://api.sume.com/v1/image-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: hero-shot-2026-08-03-001" \ -d '{"prompt":"Product hero shot of a matte black bottle on marble","mode":"async"}' ``` The response carries `request_id` (the job id), `status_url`, `result_url`, and `next_poll_after_seconds`. Poll (step 2), until `"terminal": true`: ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" ``` Fetch (step 3), when `"result_ready": true`: ```bash curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Any HTTP client can run this loop. In Kotlin with OkHttp or Ktor, the shape is: ```kotlin // Illustrative sketch — Sume does not publish a Kotlin package. // 1. POST the generate route with mode = "async" and an Idempotency-Key. // 2. GET /v1/jobs/{id}/status on a loop; honor next_poll_after_seconds when // present, exponential backoff otherwise. // 3. Stop on terminal = true; GET /v1/jobs/{id}/result when result_ready = true. // The overall deadline is CLIENT-side — 20 minutes for video is reasonable. // It is not wait_timeout_seconds, which never exceeds 30 seconds. ``` A client-side timeout does not cancel the job. The job continues to run and still bills. You only stopped the wait. Store the job id. Get the job again from `status_url`, or cancel it explicitly. In TypeScript, [`waitForJob`](/sdk/runs) from `@sume-com/sdk` is this loop. ##### Job webhooks are terminal-only `mode: "webhook"` delivers exactly three events to a public HTTPS `webhook_url`: `job.completed`, `job.failed`, and `job.canceled`. There are no progress or partial webhooks. Sume signs the raw body with HMAC SHA-256 over `{timestamp}.{raw_body}`. It sends `x-sume-webhook-timestamp` and `x-sume-webhook-signature: sume-v1=…`. Refer to [Webhooks](/workflows/webhooks) for the payload and a verifier. A webhook is a delivery optimization, not your only recovery path. Keep the `status_url` polls available for missed or retried deliveries. Action, Format, and Agent Completion **runs** are a different surface with their own `*.run.terminal` events. Refer to [Run webhooks](/agents/run-webhooks). #### Idempotency When you retry after client-side timeouts or network failures, send `Idempotency-Key` on submit requests. ```bash curl -X POST https://api.sume.com/v1/avatar-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: avatar-request-2026-06-29-001" \ -d '{"avatar_handle":"studio_presenter","input":{"type":"prompt","prompt":"Friendly studio presenter"}}' ``` Use the same key again only for the same operation and payload. #### Result shape Completed jobs can include artifacts: ```json { "id": "job_...", "status": "completed", "result": { "artifacts": [ { "id": "artf_...", "url": "https://media.sume.com/artifacts/...", "type": "image", "content_type": "image/png" } ] } } ``` Use Sume media URLs from the result. Raw provider URLs are not public API outputs. A completed text-to-speech job (`text_to_speech`) also records how Sume made its audio: `model_id` (the engine), `voice` (`{ "mode": "id", "id": "…" }`), `language`, `output_format`, and the synthesis settings, `generation_config` and `speed`. Each setting is `null` when the request did not send it. Read these values from the job to make the next line sound the same. ### Webhooks Source: https://docs.sume.com/workflows/webhooks.md Receive signed terminal job events from Sume. Use webhooks if your server needs a notification when a job reaches a terminal state. **This page is about generation jobs.** Sume has two webhook surfaces: | You called | You get | Documented at | |---|---|---| | `POST /v1/models/...` or a model endpoint like `/v1/avatar-1.0/generate` | `job.completed` / `job.failed` / `job.canceled` | This page | | An Action, Format, or Agent Completion run endpoint | `action.run.terminal` / `format.run.terminal` / `agent.run.terminal` | [Run webhooks](/agents/run-webhooks) | The event sets do not overlap, and the payloads are different. A run webhook carries the full run receipt, not a job result. The signature scheme is identical. Thus, one verifier covers both. #### Submit with a webhook URL Send `mode: "webhook"` with `webhook_url`. Webhook URLs must be public HTTPS URLs. The API rejects localhost, private-network, and non-HTTPS URLs. Webhook is one of four communication modes. [Communication modes](/workflows/jobs-and-results) compares it to `async`, `sync`, and `subscribe`. That page also describes the poll fallback to keep in place with it. #### Events Sume sends **terminal job events only**. There are no progress or partial deliveries: | Event | When it is sent | |---|---| | `job.completed` | The job completed and a public result is available. | | `job.failed` | The job failed with a public error. | | `job.canceled` | The job reached the canceled state. | #### Payload ```json { "event": "job.completed", "request_id": "job_...", "job_id": "job_...", "status": "OK", "payload": { "artifacts": [ { "id": "artf_...", "url": "https://media.sume.com/artifacts/...", "type": "image", "content_type": "image/png" } ] } } ``` Failed and canceled webhooks use `status: "ERROR"` and include an `error` object. #### Signature headers When webhook signing is configured, Sume signs the raw JSON body with HMAC SHA 256 over: ```text . ``` Headers: ```text x-sume-webhook-timestamp: 1780000000 x-sume-webhook-signature: sume-v1= ``` During a signing-secret rotation, the signature header carries one entry for each live secret. The newest entry is first, and commas separate the entries (`sume-v1=,sume-v1=`). Accept the delivery when any `sume-v1=` entry matches. Reject callbacks when the timestamp is outside your replay tolerance window. Five minutes is a reasonable default. Read your signing secret on the **Webhooks** tab of the dashboard (`/dashboard/webhooks` — Reveal, then copy). You can also read it from `GET /v1/webhooks/signing-secret` with an API key that has `account:read`. Sume derives the secret for your workspace. Thus, it is your secret, not a shared platform value. Store it as `SUME_COM_WEBHOOK_SIGNING_SECRET`, the same name that the delivery worker uses. Job webhooks and [run webhooks](/agents/run-webhooks) share that one secret. Thus, a single verifier covers both. Each delivery also carries `x-sume-webhook-secret-fingerprint`. `webhook_delivery.signing_secret_fingerprint` on the receipt gives the same value. If a signature does not verify, compare that fingerprint with the fingerprint next to the secret in the dashboard. Neither side ever needs to send the secret itself. #### Verify in TypeScript ```ts import crypto from "node:crypto"; export function verifySumeWebhook({ rawBody, timestamp, signatureHeader, secret, toleranceSeconds = 300, }: { rawBody: string; timestamp: string; signatureHeader: string; secret: string; toleranceSeconds?: number; }) { const ts = Number(timestamp); if (!Number.isFinite(ts)) return false; if (Math.abs(Math.floor(Date.now() / 1000) - ts) > toleranceSeconds) { return false; } const digest = crypto .createHmac("sha256", secret) .update(`${ts}.${rawBody}`) .digest("hex"); const expectedBuffer = Buffer.from(`sume-v1=${digest}`); // During a secret rotation the header carries one signature per live // secret, comma-separated. Accept the delivery if any `sume-v1=` entry // matches, and compare every entry so timing does not reveal which one did. let matched = false; for (const entry of signatureHeader.split(",")) { const candidate = entry.trim(); if (!candidate.startsWith("sume-v1=")) continue; const actualBuffer = Buffer.from(candidate); if ( actualBuffer.length === expectedBuffer.length && crypto.timingSafeEqual(actualBuffer, expectedBuffer) ) { matched = true; } } return matched; } ``` #### Delivery behavior After you store the event durably, return any `2xx` response. Sume retries network errors and non-2xx responses until it uses all attempts. Use `job_id` as the idempotency key on your side. | | | |---|---| | Retries | Up to **10** attempts total. | | Spacing | A fixed delay between attempts (30s by default), not exponential backoff. | | Timeout | 10s for each attempt. A slow endpoint uses the budget, and Sume retries it. | After ten refused attempts, you have a failed *delivery* and a job that still reached its real terminal state. Delivery is an optimization, never the only recovery path. Keep the `status_url` polls available for the events that never arrive. #### Send test and Redeliver These are two different actions. Do not use one in place of the other. **Send test** is one control on `/dashboard/webhooks` (or `POST /v1/webhooks/test-deliveries` with `account:write`). It POSTs a dummy signed `webhook.test` payload to a URL that *you type*. It never replays a real job. The dummy body has no `job_id` / `arun_`, and Sume does not append it to Requests. ```json { "event": "webhook.test", "request_id": "req_wh_test_…", "payload": { "ok": true, "message": "Sume webhook test. Not a job or Format run." } } ``` **Redeliver** is per call. Use it on each delivery row, or use `POST /v1/jobs/{job_id}/webhook/redeliver` (`jobs:write`). Sume then re-POSTs the real terminal event of that job (`job.completed` / `job.failed` / `job.canceled`) with a **fresh** timestamp and signature. This still works after Sume used all automatic attempts. It does not use one of the automatic 10. Receivers must treat `job_id` as the idempotency key. Redeliver does not change the destination URL. A new URL is a new job. [Run webhooks](/agents/run-webhooks) documents Format run redeliver. [Run webhooks](/agents/run-webhooks) use the same 10-attempt cap on a different schedule. That page is the authority for `*.run.terminal` deliveries. When it is available, the webhook delivery status, with the attempt count, shows on the job object and in job events. #### Next - [Run webhooks](/agents/run-webhooks) — the same signature scheme for Action, Format, and Agent Completion runs ### Media inputs Source: https://docs.sume.com/workflows/asset-library.md Use public HTTPS media inputs and consume first-party Sume artifacts. Launch generation requests accept media as public HTTPS URLs, in the exact fields that the live OpenAPI schema shows. Before you submit normal Avatar 1.0, Avatar Video, face-swap, or caption requests, you do not need to create a separate Sume asset. #### Input URL fields | Workflow | Field | Use for | |---|---|---| | Avatar 1.0 photo input | `input.image_url` | A reference photo for `input.type: "photo"`. | | Avatar Video product branch | `product_image` | An optional product/reference image. | | Avatar Video scene photo branch | `scene.image_url` | An optional scene reference when `scene.type: "photo"`. | | Avatar Video scene background image | `video_inputs[].background.url` | An image background for each scene when `background.type: "image"`. | | Face swap (Beta) | `video_url` | Public HTTPS source video for face-swap. | | Video captions | `video_url` | Public HTTPS source video to caption. | Input image/video URLs must be fetchable public HTTPS URLs. Before generation submission, the API rejects localhost, private-network URLs, non-HTTPS URLs, signed/private URLs, and mismatched content types. #### Avatar photo example Use `input.type: "photo"` with a public HTTPS `image_url` (edit the fields below if necessary): #### Avatar Video product example To use the OpenAPI default **`plus`**, omit `quality`. #### Face-swap source video example #### Generated artifacts Completed jobs can include artifact objects: ```json { "id": "artf_...", "url": "https://media.sume.com/artifacts/...", "type": "image", "content_type": "image/png" } ``` Sume mirrors generated outputs into Sume-owned media URLs before it shows them in public results. Integrations must store the Sume URL, not raw provider URLs. Signed upload/download URLs and private object keys are not part of the launch public API contract. [Trending videos](/models/trending-videos) returns public watch URLs for research. In the MVP, it does not mirror downloadable source files. ### Errors and rate limits Source: https://docs.sume.com/workflows/errors-and-credits.md Public error envelopes, request ids, rate limits, and backpressure behavior. Sume returns structured public errors. Error bodies include a request id that is safe to share with Sume support. ```json { "error": { "code": "invalid_request", "message": "Invalid request body, parameters, or headers.", "request_id": "req_...", "details": [] } } ``` #### Common API errors | Status | Code | Meaning | |---|---|---| | `400` | `invalid_request` or `bad_request` | The request body, query, path, or headers are not valid. | | `401` | `unauthorized` | The API key is missing or not valid. | | `402` | `insufficient_credits` | The balance is not sufficient for the requested generation. | | `404` | `not_found` | The resource does not exist in the current workspace. | | `409` | `job_not_completed`, `job_not_cancelable`, or `job_generation_already_started` | The requested job operation is not valid for the current status. | | `413` | `payload_too_large` | The request body is larger than the configured API limit. | | `415` | `unsupported_media_type` | The request body was not `application/json` (`details.received_content_type`). | | `429` | `rate_limited` | Too many requests in the current window. | | `429` | `queue_full` | Workspace generation concurrency and queue capacity are both full. | | `503` | `provider_not_configured`, `provider_capacity_exceeded`, or storage configuration errors | A runtime dependency is not available or is at capacity. | #### Request id The API shows the Sume request id in the response body and in the response headers. When you report an issue, include the request id. Do not include API keys, signed URLs, raw media URLs, or private workspace/user ids. #### Rate-limit headers Public API responses can include: ```text ratelimit-limit ratelimit-remaining ratelimit-reset retry-after ``` When you receive `429`, do a backoff. If `retry-after` is present, use it. Do not retry unsafe submit requests without an `Idempotency-Key`. `queue_full` is different from the usual request rate limit. It means that Sume cannot accept another paid generation job for the workspace until a current queued or processing job finishes or is canceled. Full concurrency alone is not an error. While queue capacity remains, Sume accepts valid jobs as `queued`. Refer to [Generation admission](/workflows/generation-admission). #### Provider and worker backpressure Generation can also return capacity or runtime errors before Sume accepts the provider work. | Code | Meaning | Client behavior | |---|---|---| | `provider_capacity_exceeded` | Sume's provider dispatch queue is full. | Retry later with the same idempotency key. | | `provider_not_configured` | Provider execution is not available in this runtime. | Do not retry aggressively. Examine the catalog/runtime status. | | `job_ledger_not_configured` | Job persistence is not available. | Treat it as service unavailable. | | `image_not_fetchable`, `input_media_unreachable`, or storage configuration errors | Sume could not fetch or mirror media safely. | Make sure that the input media is a public HTTPS image URL. Then retry, or contact support with the request id. | #### Job errors Failed jobs show public error metadata, for example category, stage, retryability, retry-after seconds, public reason, and next action. Internal provider payloads are not public API fields. Common job error categories include: | Category | Typical next action | |---|---| | `validation` | Correct the input. | | `auth` | Examine the API key and the workspace access. | | `quota` | Add funds, or decrease the request cost. | | `queue` | Retry later with the same idempotency key. | | `generation_unavailable` | Retry later. | | `generation_rejected` | Examine the events, and correct the unsupported input. | | `generation_timeout` | Poll the status, or retry later. | | `runtime_unavailable` | Retry later. Do not retry aggressively. | | `worker_timeout` | Poll the status, or retry later. | | `internal` | Examine the events, and contact support with the request/job id. | #### Status vocabulary | Object | Values | |---|---| | Job status | `queued`, `processing`, `completed`, `failed`, `canceled` | | Resource status | `processing`, `ready`, `failed`, `canceled`, `archived` | | Webhook delivery status | `pending`, `delivering`, `delivered`, `retrying`, `failed`, `exhausted` | ### Recipes Source: https://docs.sume.com/api/cookbook.md Short current-platform recipes for common Sume API workflows. This page is a short index of recipes. For Avatar, the canonical `/v1/avatar-1.0/*` routes are the best choice. The deep guides are in [Models](/models). A Format that you run for your own customers has a longer recipe of its own. Refer to [Embed a Format in your product](/cookbooks/embed-a-format). #### Verify a key ```bash curl https://api.sume.com/v1/me \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### List current capabilities ```bash curl https://api.sume.com/v1/catalog \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Create an avatar #### Create an avatar video The estimated length of an avatar-video script must be 4-60 seconds inclusive. For the default **`plus`** path, do not send `quality`. #### Preview first frames, then generate After the preview job completes: #### Face swap (Beta) #### Caption an existing video #### Search trending TikTok videos #### Generate an image Deep guide: [Image 1.0](/models/image). #### Generate a video Deep guide: [Video 1.0](/models/video). #### Generate music Deep guide: [Music 1.0](/models/music). Do not send `duration` on Music 1.0. #### Poll and fetch result ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Use media inputs The API is URL-first. For normal flows, a separate asset upload is not necessary. Use public HTTPS URLs directly in these launch request fields: - Avatar photo input: `input.image_url` - Avatar Video product reference: `product_image` - Avatar Video scene photo reference: `scene.image_url` - Face swap / captions source: `video_url` Refer to [Media inputs](/workflows/asset-library). #### Read usage ```bash curl https://api.sume.com/v1/balance \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/usage \ -H "Authorization: Bearer $SUME_API_KEY" ``` ## SDK ### Overview Source: https://docs.sume.com/sdk.md @sume-com/sdk — install, create a client, authenticate, and run a Format end to end from Node, Bun, Deno, or Workers. `@sume-com/sdk` is the official TypeScript client for `api.sume.com`. It covers each operation in the public OpenAPI schema. It also has the helpers that a partner otherwise writes by hand: [`subscribeFormatRun`](/sdk/runs), [`waitForRun`](/sdk/runs), [`waitForJob`](/sdk/runs), [`uploadFile`](/sdk), and [`verifyWebhook`](/sdk/webhooks). ```bash npm install @sume-com/sdk ``` The published package is [`@sume-com/sdk@0.2.0`](https://www.npmjs.com/package/@sume-com/sdk). It has an MIT license and **no runtime dependencies**. It needs `fetch` and WebCrypto: Node 18+, Bun, Deno, or Cloudflare Workers. **This page is not a method reference.** The request and response fields come from the [API reference](/api/reference) and the live OpenAPI (`https://api.sume.com/reference/json`). These two stay the source of truth. This page gives the client part: the factory, auth, and the helpers that have no REST equivalent. #### Create a client ```ts import { createSumeClient } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY!, }); ``` | Option | Default | Notes | |---|---|---| | `apiKey` | — | Required. A Developer API key from [API Keys](https://www.sume.com/dashboard/api-keys). | | `baseUrl` | `https://api.sume.com` | Use `https://api.dev.sume.com` for development. | | `fetch` | the runtime's `globalThis.fetch` | Supply your own function for instrumentation, retries, or tests. | **Pass `client` on every call.** Operations accept a default client at the module level. But that default targets `https://api.sume.com` with no key. It is there only so that the generated code compiles. It does not let you skip the factory. Pass the client that you made: ```ts import { createSumeClient, listFormats } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const { data } = await listFormats({ client }); ``` #### Authentication The client sends **`x-api-key` only**. It does not set `Authorization`. Do not add that header yourself. The API accepts each header alone (refer to [Authentication](/authentication)). But it rejects **both at once**. A request that has `Authorization: Bearer` and `x-api-key` together fails with `401 unauthorized` and `Send only one API key credential.` There is no precedence rule. Neither header wins. Thus, an `Authorization` header in addition to the header of the client causes the request to fail, even when `x-api-key` is correct. That extra header can come from a session token, the credential of a gateway, or an interceptor that you forgot. If you wrap `fetch` through the `fetch` option, make sure that your wrapper does not add one. Your key needs `formats:read` and `formats:write` to run Formats. The scopes of a key are fixed when you create the key. You cannot add scopes later. An older key returns `403 insufficient_scope` on every run. Create a new key and rotate the old key. **Team Formats need a team (workspace) key.** If a personal key calls a team Format, the call fails with `403 workspace_key_required`. Refer to [Calling a Format](/formats/call#team-formats-need-a-team-key). **Server-side only.** A Sume API key spends your credits. There is no browser-safe variant. Never put a key in client JavaScript, a mobile bundle, or a `NEXT_PUBLIC_*` variable. Put your own endpoint in front, and make the Sume request from that endpoint. [Embed a Format in your product](/cookbooks/embed-a-format#1-key-custody) gives the full custody rules. #### A Format run, end to end For Formats, prefer **`subscribeFormatRun`**. One call creates the run and waits for the terminal receipt. If you can skip the wait fully, use it together with [run webhooks](/agents/run-webhooks). ```ts import { createSumeClient, subscribeFormatRun } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const run = await subscribeFormatRun({ client, path: { handle: "acme", slug: "product-promo" }, idempotencyKey: "order-8823-promo-v1", body: { input: { product_url: "https://shop.example.com/p/8823" }, generation_spend_cap_usd: 3, }, onStatus: (status, snapshot) => console.log(status, snapshot.next_action), }); if (run.status === "completed") { console.log(run.primary_output_url); } ``` What to expect today: - **`subscribeFormatRun` polls.** There is no SSE stream. Thus, `onStatus` shows the result of status polls, not a push feed. For progress detail while you wait, read `events_url`. It is a phase timeline, not a log feed. Refer to [Watch a run progress](/formats/runs#watch-a-run-progress). - **Prefer a webhook** when delivery is available for your environment. Pass `communication.webhook_url` on create (or skip the wait and handle the push). Refer to [Verifying webhooks](/sdk/webhooks). - **Default timeout is 20 minutes** (video Formats usually run 10–20). It resolves for each terminal status. It throws `SumeRunRequestError` when the API refuses the create call or when a status read has a non-transient failure. It throws `SumeRunTimeoutError` when the timeout occurs first. - **Generated operations do not throw on an API error.** They resolve with `{ data, error, response }`. `subscribeFormatRun` / `waitForRun` throw, because a poll loop has no place to put a non-result. When you already have a run id, use [`waitForRun`](/sdk/runs). The run id can come from Action / Agent Completion, or from a create that you made yourself. For **generation jobs**, use [`waitForJob`](/sdk/runs). These jobs come from `/v1/image-1.0/generate`, `/v1/video-1.0/generate`, and the Avatar routes. They live at `/v1/jobs/:id` and are not runs. #### Where to go next | You want | Read | |---|---| | Create + wait (or poll an existing run or job) | [Waiting for runs and jobs](/sdk/runs) | | Verify a signed webhook delivery | [Verifying webhooks](/sdk/webhooks) | | Exact request and response fields | [API reference](/api/reference) | | The whole partner integration | [Embed a Format in your product](/cookbooks/embed-a-format) | #### Scope of this section These pages document the hand-written surface: the client factory and the helpers. A generator makes all other exports of the package from the same OpenAPI schema that the [API reference](/api/reference) describes. There is one function for each operation. The name of each function is the operation id: `listFormats`, `createFormatRun`, `getFormatRunStatus`, `cancelFormatRun`, and others. The autocomplete of your editor is a better catalog than a copy of the schema. Thus, this page does not have a copy. ### Waiting for runs and jobs Source: https://docs.sume.com/sdk/runs.md subscribeFormatRun creates a Format run and waits for the receipt; waitForRun polls any Format, Action, or Agent run; waitForJob polls a generation job. Prefer webhooks in production — there is no SSE stream yet. Each Sume run is asynchronous. `POST .../runs` gives back a receipt with a `status_url`, and you learn the result later. A bulk queue is a server-side list of those runs. Poll `GET /v1/format-run-queues/{queue_id}` for counts. The helpers on this page still wait on one run id. Refer to [Bulk runs](/formats/bulk-runs). In `@sume-com/sdk@0.2.0`, the Format path for partners is **`subscribeFormatRun`** (create + wait). When you already have a run id, use **`waitForRun`**. For generation jobs, use **`waitForJob`**. Generation jobs are a separate surface. When you can skip the wait fully, prefer [run webhooks](/agents/run-webhooks). There is **no SSE event stream** today. Thus, "subscribe" means create, then poll. `onStatus` shows the result of status polls, not a live log feed. It also gives a snapshot with more data, for example `next_action`, timestamps, and `cancelable`. On a Format run, you can also poll the **phase timeline**. Refer to [Watching phases while you wait](#watching-phases-while-you-wait). #### Format runs: `subscribeFormatRun` ```ts import { createSumeClient, subscribeFormatRun } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const run = await subscribeFormatRun({ client, path: { handle: "acme", slug: "product-promo" }, idempotencyKey: "order-8823-promo-v1", body: { input: { product_url: "https://shop.example.com/p/8823" }, generation_spend_cap_usd: 3, // Or skip the wait and take the push: // communication: { webhook_url: "https://partner.example/hooks/sume" }, }, onStatus: (status, snapshot) => console.log(status, snapshot.next_action), }); if (run.status === "completed") { console.log(run.primary_output_url); } else { console.error(run.status, run.error); } ``` | Option | Default | Notes | |---|---|---| | `path` | — | Vanity `{ handle, slug }` or `{ format_id }`. | | `body` | — | The same body as `createFormatRun*` (input, caps, schema, attachments, …). | | `idempotencyKey` | auto UUID | The helper sends it as `Idempotency-Key`. A replay of a finished run returns immediately. To send no key, pass `null`. | | `timeout` | **20 minutes** | Longer than the 10 minutes of `waitForRun`. Video Formats usually run 10–20. | | `pollInterval` | 2 seconds | The time between status reads. | | `signal` | — | Aborts the wait and the in-flight request. | | `timeline` | `false` | Also reads the phase timeline on each poll and gives it to `onStatus` as `snapshot.timeline`. | | `onStatus` | — | `(status, snapshot)` on each status read, and also on the terminal read. | | `onCreated` | — | The helper calls it one time with the accepted run, before the polls start. | It resolves for **any** terminal status. It **throws** `SumeRunRequestError` when the API refuses the create call. An example is `403 workspace_key_required` when a personal key calls a team Format. In that condition, there is no run to wait for. It throws the same error when a status read fails and the failure is not transient. It throws `SumeRunTimeoutError` when `timeout` occurs first. [Runs and results](/formats/runs) documents each field of the receipt. #### Already have a run id: `waitForRun` Use this helper for Action / Agent Completion runs. Also use it when you created the Format run yourself and need only the poll loop. ```ts import { createSumeClient, waitForRun } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const run = await waitForRun(runId, { client, family: "format", timeout: 15 * 60_000, pollInterval: 2_000, signal: AbortSignal.timeout(20 * 60_000), onStatus: (status, snapshot) => console.log(status, snapshot.next_action), }); ``` | Option | Default | Notes | |---|---|---| | `family` | — | **Required.** `"format"`, `"action"`, or `"agent"`. | | `client` | module default | The client from `createSumeClient()`. Pass it. The module default has no base URL or key. | | `timeout` | 10 minutes | If the wait is longer, the helper throws `SumeRunTimeoutError`. | | `pollInterval` | 2 seconds | The time between status reads. | | `signal` | — | Aborts the wait and the in-flight request. Rejects with the reason of the signal. | | `timeline` | `false` | Format runs only. Also reads the phase timeline on each poll. | | `onStatus` | — | `(status, snapshot)` on each status read, and also on the terminal read. | **`family` is required. The helper cannot infer it.** A run id does not show its surface. The three families live behind three different URL prefixes (`/v1/format-runs/…`, `/v1/action-runs/…`, `/v1/agent-runs/…`). This option also sets the type of the return value: `family: "format"` resolves as `PublicFormatRun`. The helper checks the deadline **before** it sleeps, not after. If a caller asks for a 5-second timeout, the caller gets the timeout in 5 seconds. The caller does not wait 5 seconds plus one full poll interval. #### Watching phases while you wait `status` tells you that a Format run is `processing`. It does not tell you *what* the run does. For a fifteen-minute video run, that is most of the information that you want. The **phase timeline** gives this information: `preparing`, `running`, `finalizing`, each with a timestamp and a status. Pass `timeline: true`. The timeline then comes on the `onStatus` snapshot: ```ts const run = await subscribeFormatRun({ client, path: { handle: "acme", slug: "product-promo" }, body: { input: { product_url: "https://shop.example.com/p/8823" } }, timeline: true, onStatus: (status, snapshot) => { const phase = snapshot.timeline?.at(-1); console.log(status, phase?.phase, phase?.status); }, }); ``` For a run that you do not wait on, you can also read it directly: ```ts import { getFormatRunTimeline } from "@sume-com/sdk"; const timeline = await getFormatRunTimeline(runId, { client }); // [{ at: "2026-08-05T…", phase: "running", status: "running", duration_ms: null }, …] ``` Know these three things: - **It is a phase timeline, not a log stream.** Sume never publishes agent output, tool calls, or sandbox internals. Refer to [Watch a run progress](/formats/runs#watch-a-run-progress). - **`timeline: true` doubles the request rate of the wait.** Reads have their own rate-limit budget. Thus, this option cannot cause a 429 on your run creates. But it still sends twice the requests for the same run. Thus, it is off by default. - **A timeline read that fails does not end the wait.** The run still executes and still spends. The helper reports the failure through `onTransientError` and keeps the previous timeline. This option is for Format runs only. The helper ignores `timeline: true` for `family: "action"` and `"agent"`. The receipts of those runs report `events_url: null`, because they have no events route. #### Generation jobs: `waitForJob` Runs and jobs are different surfaces. `POST /v1/image-1.0/generate`, `/v1/video-1.0/generate`, the Avatar routes, and the `/v1/models/sume/…/runs` aliases all create **jobs** at `/v1/jobs/:id`, not runs. Thus, they need `waitForJob`, not `waitForRun`. You cannot use a job id as a run id, or a run id as a job id. ```ts import { createSumeClient, generateVideoV1, waitForJob } from "@sume-com/sdk"; const client = createSumeClient({ apiKey: process.env.SUME_API_KEY! }); const { data: submitted, error } = await generateVideoV1({ client, headers: { "idempotency-key": crypto.randomUUID() }, body: { prompt: "Slow push-in on a ceramic mug", mode: "async" }, }); if (error) throw new Error(JSON.stringify(error)); const job = await waitForJob(submitted!.data.request_id, { client, onStatus: (status, snapshot) => console.log(status, snapshot.next_action), }); console.log(job.status, job.result.artifacts); ``` | Option | Default | Notes | |---|---|---| | `client` | module default | The client from `createSumeClient()`. Pass it — the module default has no base URL or key. | | `timeout` | **20 minutes** | Longer than the 10 minutes of `waitForRun`. Video and avatar-video jobs usually run for minutes. If the wait is longer, the helper throws `SumeJobTimeoutError`. | | `pollInterval` | 2 seconds | A **floor**. When `next_poll_after_seconds` in the status payload asks for a longer time, that value wins. | | `signal` | — | Aborts the wait and the in-flight request. | | `onStatus` | — | `(status, snapshot)` on every status read, including the terminal one. | Submit with **`mode: "async"`** (or omit `mode`). The server-side `sync` and `subscribe` modes are the same bounded wait, with a cap of 30 seconds. That cap is a budget for the HTTP request, not for the job. Refer to [Communication modes](/workflows/jobs-and-results). `waitForJob` is the client-side wait, and it can wait longer than that cap. It resolves with the job record from `/v1/jobs/:id`, not from `/v1/jobs/:id/result`. For failed and canceled jobs, `/result` answers `409 job_not_completed`, and then the helper has nothing to give back. Read `status`, `result`, and `error` from the record. The errors are the same as the errors of the run helpers: `SumeJobTimeoutError` and `SumeJobRequestError`. Both errors carry `jobId`. A timeout does not cancel the job. The job continues to run, and it still bills. Store the job id. Read the job again later with `getApiJob`, or cancel it with `cancelApiJob`. #### Terminal is not the same as successful Both run helpers resolve for **any** terminal status: `completed`, `failed`, `canceled`, `skipped`. A failed run is a result that you asked for, not an exception. Thus, read `status` and `error` from the receipt, the same as a webhook handler does: ```ts const run = await subscribeFormatRun({ client, path: { handle: "acme", slug: "product-promo" }, body: { input: { product_url: "https://example.com/p" } }, }); switch (run.status) { case "completed": return attachOutputs(run); case "failed": // `unattended_blocked` means the run hit a gate it could not clear // without a human. The message is written to be shown. return showError(run.error); case "canceled": case "skipped": return noop(); } ``` Give `skipped` its own branch. It means that a run was already in flight and you passed `on_active_run: "skip"`. (Format runs allow concurrency by default.) Refer to [Runs and results](/formats/runs) for the full lifecycle. #### Errors they throw The generated operations resolve with `{ data, error }`. These helpers are different: they throw, because a poll loop has no place to put a non-result. | Error | When | |---|---| | `SumeRunTimeoutError` | `timeout` occurred first. Carries `runId` and `lastStatus`. | | `SumeRunRequestError` | The API refused the create, or a read failed and the failure was not transient. Carries `runId` (or `"(not created)"`) and all the fields of `SumeApiError`. | | the `signal`'s reason | You aborted. | `SumeRunRequestError` extends `SumeApiError`. Thus, the envelope is available as typed fields. You do not have to get it out of `body`: ```ts import { SumeRunRequestError } from "@sume-com/sdk"; try { const run = await subscribeFormatRun({ client, path, body }); } catch (error) { if ( error instanceof SumeRunRequestError && (error.status === 402 || error.code === "insufficient_credits") ) { return topUpAndAlert(error.requestId); // next_action: "add_funds" } if (error instanceof SumeRunRequestError) { log.error({ code: error.code, requestId: error.requestId, retryable: error.retryable }); } throw error; } ``` `SumeApiError` subclasses: `SumeAuthenticationError` (401), `SumeInsufficientCreditsError` (402), `SumePermissionError` (403), `SumeNotFoundError` (404), `SumeConflictError` (409), `SumeRateLimitError` (429), `SumeServerError` (5xx). Each one carries `code`, `requestId`, `retryable`, `retryAfterSeconds`, `nextAction`, `details`, and the raw `body`. The run helpers always throw `SumeRunRequestError` itself, never one of these subclasses. Thus, branch on its `status` or `code`. #### Transient failures do not end the wait A `429` or a `5xx` on a status read means that the *read* failed, not the run. The run still executes and still spends. Thus, the helpers do a backoff and poll again, and they do not throw. If you lose the handle to a live run, the cost is much higher than one more second of wait. Two layers do this. Both are on by default: | Layer | Default | Option | |---|---|---| | `createSumeClient` retries `408`/`429`/`5xx` and transport failures | 2 retries, exponential backoff + jitter, honors `retry-after` | `maxRetries`, `timeout` | | `waitForRun` accepts consecutive transient read failures | 6 | `maxTransientFailures`, `onTransientError` | The client retries a `POST` only when it carries an `Idempotency-Key`. Without a key, a replay starts and bills a second run. `subscribeFormatRun` generates a key for you, unless you pass your own key (or `null`). The helpers add jitter to polls. Reads have their own rate-limit budget. But without jitter, several clients that start together stay in phase and hit that ceiling as a group. ```ts import { SumeRunTimeoutError, waitForRun } from "@sume-com/sdk"; try { const run = await waitForRun(runId, { client, family: "format" }); } catch (error) { if (error instanceof SumeRunTimeoutError) { // The run is still going. Nothing was lost — read it later from `result_url`. await markPending(error.runId, error.lastStatus); } else { throw error; } } ``` A timeout does **not** cancel the run. The run continues. You only stopped the wait. Store the run id. Get the run again later with `getFormatRun`, or cancel it explicitly with `cancelFormatRun`. #### Prefer a webhook where you can A poll loop needs one timer and one open request for each run in flight, and runs usually take minutes. [Run webhooks](/agents/run-webhooks) deliver the same receipt without these costs. [`verifyWebhook`](/sdk/webhooks) is the half on the receiver side. Polls are the correct tool in three conditions. You are in a job that can block without a problem. You make a prototype. Or webhook delivery is still off for your environment. Build the receiver now. Keep `subscribeFormatRun` / `waitForRun` as the fallback. #### Next - [Verifying webhooks](/sdk/webhooks) — the receiver check for the push path - [Runs and results](/formats/runs) — each field of the receipt - [Bulk runs](/formats/bulk-runs) — queue many runs, then poll `GET /v1/format-run-queues/{id}` - [Embed a Format in your product](/cookbooks/embed-a-format) — the whole partner integration ### Verifying webhooks Source: https://docs.sume.com/sdk/webhooks.md verifyWebhook checks the sume-v1 signature on a Sume delivery — raw bodies, header shapes, the replay window, and why it is async. Sume signs each webhook delivery with HMAC-SHA256 over `.` and sends the signature as `sume-v1=`. `verifyWebhook` does that check. Thus, you do not write it. ```ts import { verifyWebhook } from "@sume-com/sdk"; export async function POST(request: Request) { const body = await request.text(); // raw, before any JSON.parse const ok = await verifyWebhook({ body, headers: request.headers, // Same name the Sume delivery worker uses. Self-serve: Dashboard → Webhooks, // or GET /v1/webhooks/signing-secret. secret: process.env.SUME_COM_WEBHOOK_SIGNING_SECRET!, }); if (!ok) return new Response("bad signature", { status: 401 }); const event = JSON.parse(body); await recordTerminalRun(event.request_id, event); // dedupe on request_id return new Response(null, { status: 204 }); // fast 2xx, then work } ``` Read your signing secret on the **Webhooks** tab of the dashboard (`/dashboard/webhooks` — Reveal, then copy). You can also read it from `GET /v1/webhooks/signing-secret` with an API key that has `account:read`. Sume derives the secret for your workspace. Thus, a valid signature proves that Sume signed the delivery for you, not for any holder of a shared platform secret. Store the secret the same way that you store the API key. Use the env name `SUME_COM_WEBHOOK_SIGNING_SECRET`, so that your local samples use the same name as the Sume worker that signs. The secret is not the API key, and the client has no part in the check. `verifyWebhook` takes no `client` and makes no request. The dashboard also shows a **fingerprint** of the secret. Each delivery carries the same value as `x-sume-webhook-secret-fingerprint`. When a signature does not verify, compare the fingerprints. The fingerprint is the only part of this data that is safe to paste into a ticket. #### Rotating the secret If it is possible that your secret leaked, rotate it. Use **Webhooks → Rotate secret**, or `POST /v1/webhooks/signing-secret/rotate` with a key that has `account:write`. A rotation is not a cutover. For **24 hours** after the rotation, Sume signs each delivery with both secrets. Sume sends the two signatures in `x-sume-webhook-signature`, with a comma between them and the newest first: ```http x-sume-webhook-signature: sume-v1=,sume-v1= ``` A receiver with one of the two secrets can verify the delivery. Thus, you can redeploy on your own schedule, not at the same instant that you push the button. After the window, a check with the old secret fails. The dashboard shows the deadline while the window is open. `rotation.previous_valid_until` gives the deadline on both API responses. CAUTION: Upgrade the receiver *before* you rotate. A hand-rolled verifier that compares the header for equality fails on each delivery during the window. `verifyWebhook` in **`@sume-com/sdk` 0.2.0** (the current release) already handles the multi-signature header. If you never rotate, nothing changes, because Sume sends only one signature outside a window. `x-sume-webhook-secret-fingerprint` names the **new** secret from the moment that you rotate, and also during the window. It tells you the secret to move to. It does not tell you which secrets Sume still accepts. If you rotate two times in one window, Sume immediately retires the secret from two rotations before. This is how you make a leak really stop. Delivery is live on **`api.dev.sume.com` and `api.sume.com`**. Polls or [`subscribeFormatRun`](/sdk/runs) stay a valid backup. Refer to [Run webhooks](/agents/run-webhooks#availability). #### Input | Field | Notes | |---|---| | `body` | The **raw** body: `string`, `ArrayBuffer`, or a typed array. | | `headers` | A `Headers`, a `Map`, or a plain object (Node's `req.headers`). Case-insensitive. | | `secret` | Your Sume webhook signing secret. | | `toleranceSeconds` | Replay window. Default `300`. `0` skips the timestamp check. | The helper reads these two headers: ```text x-sume-webhook-timestamp: 1785000000 x-sume-webhook-signature: sume-v1= ``` #### Four rules that decide whether this works - **Pass the raw body.** A parsed-and-reserialized object does not verify, because the key order and the whitespace are part of the signed data. A framework that parses JSON for you destroyed the bytes before your code gets them. In Express, mount `express.raw({ type: "application/json" })` on the webhook route only. In Next.js App Router, `await request.text()` before all other steps. - **It is `async`.** The implementation uses WebCrypto, not `node:crypto`. Thus, you can import the package from Workers, Deno, and bundlers that refuse `node:` specifiers. `await` it. - **It returns `false` and does not throw** on a malformed delivery. A missing header, a bad timestamp, and a wrong signature all give only a failed verification. You branch on one value, with no `try`/`catch`. - **Comparison is constant-time.** The helper enforces the replay window before it computes the HMAC. #### One verifier, two surfaces Run webhooks (`*.run.terminal`, with `run_id`) and generation-job webhooks (`job.*`, with `job_id`) use exactly the same `sume-v1` scheme. The payloads are different, but the signature is the same. Thus, one verifier covers both. **Route on `event`**, and never expect that a body has `run_id`: ```ts const event = JSON.parse(body); switch (event.event) { case "format.run.terminal": return handleFormatRun(event); case "job.completed": case "job.failed": return handleJob(event); default: return new Response(null, { status: 204 }); // unknown event, not a 500 } ``` When you treat an unrecognized event as `204`, a new event type does not cause a 500 and a retry storm. #### Constants, if you need them `verifyWebhook` covers the usual case. The package also exports the pieces, for when you do not control the receiver. An example is a gateway that verifies before your code runs: | Export | Value | |---|---| | `SUME_WEBHOOK_SIGNATURE_VERSION` | `"sume-v1"` | | `SUME_WEBHOOK_SIGNATURE_HEADER` | `"x-sume-webhook-signature"` | | `SUME_WEBHOOK_TIMESTAMP_HEADER` | `"x-sume-webhook-timestamp"` | | `DEFAULT_SUME_WEBHOOK_TOLERANCE_SECONDS` | `300` | #### Next - [Run webhooks](/agents/run-webhooks) — delivery, events, payloads, retries, and the raw scheme for non-JavaScript receivers - [Webhooks](/workflows/webhooks) — generation-job webhooks, the other surface - [Waiting for runs](/sdk/runs) — the poll path (an optional fallback, because delivery is live on production) ## Agents ### Overview Source: https://docs.sume.com/agents.md How Sume developer tools fit into agent workflows. Sume developer tools use one boundary: the public API is the source of truth, and wrappers stay thin. #### Tooling layers 1. Developer API: workspace-scoped HTTP endpoints. 2. CLI: shell-friendly wrapper around the Developer API. 3. Hosted MCP: remote MCP at `https://mcp.sume.com/mcp` for Cursor, Claude, and similar clients. 4. Local `sume mcp`: future CLI surface (`coming_soon` in current releases). Use hosted MCP or direct CLI commands instead. 5. Dashboard tools: app-integrated tools with server-resolved workspace context. Agents must use catalog, jobs, media schema, and usage reads to make a plan. They must do this before they ask for confirmation for write or paid actions. Refer to [Safe automation](/agents/safe-automation). #### Where to start | Goal | Start here | |---|---| | Call a saved authoring recipe from your backend | [Format API](/formats) · [Format catalog](/formats/catalog) | | Interactive Cursor / Claude connector | [MCP quickstart](/mcp/quickstart) | | Local shell agent with CLI | [Quick start](/) · [CLI overview](/cli) | | Backend / server integration | [Quick start](/) (API) · [Public API](/public-api) | | Run the Agent on an ad-hoc task from your backend | [Agent Completions](/agents/completions) | | Run an agent on a recurring schedule | [Scheduled](/agents/actions) | | Create an API key | [API Keys dashboard](https://www.sume.com/dashboard/api-keys) | | Agents product UI | [Agents](https://www.sume.com/agents) | | Avatar or Image first job | [Create new avatar](/models/avatar) · [Image 1.0](/models/image) | Hosted MCP paid tools include `generate_image`, `generate_video`, `music_create`, `tts_create`, and Avatar/Stage P creates. Sume Image 1.0 and Video 1.0 (`images_create` / `videos_create`) stay REST-only. For those two products, call the Developer API. Details: [MCP overview](/mcp). #### Workspace context Agent tools must not ask the user for a workspace id. Workspace context comes from the API key or the authenticated app session. ### Scheduled Source: https://docs.sume.com/agents/actions.md Run a saved Agents automation on a recurring schedule, and how its runs differ from one-shot generation jobs. A schedule is a saved Agents automation that runs on a cadence. It has instructions, a model, a cron expression, and a spend cap. When it fires, Sume runs it as an Agent in a fresh thread. Sume then returns a structured run receipt. **A cadence is the whole point.** Do you want to call Sume from your own backend when *your* user does something? If so, use the [Format API](/formats) instead. The Format API has the same run engine and the same receipt. You address it on each call, and you give your inputs. Use a schedule when only the clock starts the work. You create schedules in the Agents dashboard. You can also ask the Agent in chat to make one for you. The Developer API can list them, read them, start runs, and monitor runs. It cannot create or edit them. The rest of Scheduled is on three pages that this page links to: [Create a schedule](/agents/actions/create), [Runs and results](/agents/actions/runs), and [Advanced: run a schedule via API](/agents/actions/api-trigger). > **Scheduled and the Actions API.** The name of the product is Scheduled. The > HTTP API namespace is still `/v1/actions`, ids are `aut_…`, and the API returns > objects as `object: "action"`. Those names are stable and will not change. On > this page, "Action" is the wire spelling of a schedule. #### Schedule vs. generation job A schedule run is not a job. It does not appear in `/v1/jobs`. It does not use the job lifecycle that [Jobs and results](/workflows/jobs-and-results) describes. | | Schedule run | Generation job | |---|---|---| | Started by | A cron schedule, or `POST /v1/actions/{action_id}/runs` | `POST /v1/{family}-1.0/...` | | Unit of work | Saved instructions that an Agent executes in a new thread | One model invocation | | Read back from | `/v1/action-runs/{run_id}` | `/v1/jobs/{id}` | | Statuses | `queued`, `processing`, `completed`, `failed`, `canceled`, `skipped` | See [Jobs and results](/workflows/jobs-and-results) | | Result shape | `output` projected onto an output schema, plus `artifacts` | Job `result` | | Overlap policy | `on_active_run` (`skip` or `reject`) | None | Use a generation job when you want one model invocation. Use a schedule when you want an Agent to do saved instructions on a cadence. This work can use several generations. If the task changes on each call and you have nothing to save, use [Agent Completions](/agents/completions) instead. It uses the same agent, with no saved object. You supply the instruction on each request. #### Anatomy of a schedule `GET /v1/actions` and `GET /v1/actions/{action_id}` return this shape. ```json { "id": "aut_...", "object": "action", "title": "Weekly product teaser", "status": "active", "trigger_type": "api", "api_trigger_enabled": true, "cron": null, "model": "...", "output_schema": null, "primary_output_key": null, "generation_spend_cap_usd_micros": 1000000, "last_run_at": "2026-07-30T09:00:00.000Z", "created_at": "2026-07-20T12:00:00.000Z", "updated_at": "2026-07-30T09:00:00.000Z", "invoke_url": "https://api.sume.com/v1/actions/aut_.../runs" } ``` | Field | Notes | |---|---| | `status` | `active` or `inactive`. An `inactive` schedule rejects API runs. | | `trigger_type` | `cron` or `api`. Fixed at create time. | | `api_trigger_enabled` | When `true`, `POST /v1/actions/{action_id}/runs` is permitted. A `cron` schedule can also enable it. | | `cron` | `{ "expr", "timezone", "next_run_at" }`, or `null` for the API-only case. | | `output_schema` | Default structured output binding (`{ "name", "strict" }`), or `null` for the built-in default. Bind one in the dashboard. Refer to [Structured output](/formats/structured-output). | | `generation_spend_cap_usd_micros` | The generation cap for each run, in USD micros. `null` means that the $1.00 default applies. | | `invoke_url` | The invoke endpoint of the schedule. | The public shape does not include the `instructions` text. This is intentional. Use the dashboard to read and edit instructions. #### Triggers You select the trigger type at create time. After that, you cannot change it: - **Scheduled** (`cron`) — the default. The schedule runs on a 5-field cron expression in an IANA timezone. - **API call** (`api`) — an advanced option with no cadence. The schedule runs only when your service calls `POST /v1/actions/{action_id}/runs`. Refer to [Advanced: run a schedule via API](/agents/actions/api-trigger). A cron schedule can also set `api_trigger_enabled` to accept API runs in addition to its cadence. An API-only schedule never has a cadence. #### Where schedules live Create and monitor schedules at `https://www.sume.com/agents/scheduled`. For the dashboard flow, refer to [Create a schedule](/agents/actions/create). #### Limits | Limit | Value | |---|---| | `input` properties | 64 | | `input` size | 2097152 UTF-8 bytes (2 MiB) | | Default generation spend cap | $1.00 (`1000000` USD micros) if not set | | Per-run spend cap override | The API clamps a number > 0 to `min(request, schedule cap)`. It can lower the cap, but it can never raise it. `null` runs without the automation ceiling. The API rejects `0` | | `Idempotency-Key` length | 1–255 characters | | `limit` on list endpoints | 1–100, default 50 | #### What schedules do not support yet Before you make a design that uses schedules, know these gaps: - **Signing secrets are not self-serve yet.** The API accepts, validates, stores, and delivers `communication.webhook_url` on `api.dev.sume.com` and `api.sume.com`. [Run webhooks](/agents/run-webhooks) documents the contract. - **No events endpoint.** `events_url` on a run receipt is always `null`. The API does not expose run lifecycle events. Use `status_url` and `result_url`. - **No MCP tool and no CLI command.** [MCP](/mcp) and the [CLI](/cli) do not expose schedules. - **No write endpoints.** The Developer API cannot create, edit, or delete a schedule. #### Next - [Create a schedule](/agents/actions/create) - [Runs and results](/agents/actions/runs) - [Safe automation](/agents/safe-automation) - [Advanced: run a schedule via API](/agents/actions/api-trigger) - [Run webhooks](/agents/run-webhooks) ### Agent Completions Source: https://docs.sume.com/agents/completions.md Call the Sume Studio Agent like a model — messages or an instruction in, an async agent.run receipt out, with tools, media generation, and a sandbox. An Agent Completion runs the Sume Agent on an ad-hoc prompt. It uses the same runtime as the Agents chat UI: full sandbox, tools, MCP bridge, and media generation. Your own backend can use an API key to start it, and no person has to monitor it. A [schedule](/agents/actions) stores *what to do*. A [Format](/formats) stores *how to do it*. An Agent Completion stores nothing. You send the task on each call. #### Which one do I want? | Surface | Use it when | Start with | |---|---|---| | **[Format](/formats)** | You have a saved packaged workflow and only the inputs change. | `POST /v1/formats/{handle}/{slug}/runs` | | **[Scheduled](/agents/actions)** | You want a schedule or a trigger on a saved automation. | `POST /v1/actions/{handle}/{slug}/runs` | | **Agent Completions** | The task changes on each call. You only want the Agent to do this one task. | `POST /v1/agent/completions` | All three run the same agent and return the same receipt shape. The only differences are the source of the instruction and what Sume saved for you. #### Async, not a chat drop-in An Agent Completion is **not** a synchronous chat completion. A real agent turn opens a sandbox, calls tools, and can generate media. This work takes too long to keep one HTTP request open. Thus, the create call returns `202` and a receipt, and you poll the receipt. The request uses the OpenAI `messages[]` shape, so your current integration code fits. But the response is a run receipt, not `choices[]`. Streaming and a sync OpenAI-compatible wire are not available at this time. #### Scopes | Scope | Needed for | |---|---| | `agent_completions:read` | Read and list runs. | | `agent_completions:write` | Create a completion, cancel a run. | **Keys that you created before Agent Completions shipped do not have these scopes**. Each request with an older key fails with `403 insufficient_scope`. You cannot add scopes to a key that already exists. Create a new key at [API Keys](https://www.sume.com/dashboard/api-keys). Then rotate to the new key. Refer to [Authentication](/authentication). Service-account keys cannot create Agent Completions. A request with a service-account key fails with `403 insufficient_scope` and `details.reason` of `service_account_agent_completions_unsupported`. #### Create a completion To rewrite the example, put your values in the required fields. `generation_spend_cap_usd` has no default. If you do not send it, the request fails. If Sume accepts the completion, it returns `202` and a receipt: ```json { "data": { "id": "agrun_...", "object": "agent.run", "model": "sume-agent", "thread_id": "thr_...", "status": "queued", "status_url": "https://api.sume.com/v1/agent-runs/agrun_.../status", "cancel_url": "https://api.sume.com/v1/agent-runs/agrun_.../cancel", "created_at": "2026-08-01T16:00:00.000Z", "output": null, "artifacts": [], "usage": { "generation_spend_cap_usd_micros": 5000000 } } } ``` ##### Request fields | Field | Required | Notes | |---|---|---| | `instruction` | one of | The task, as a plain string. | | `messages` | one of | `system` and `user` turns. Send exactly one of `instruction` or `messages`. Do not send both. | | `generation_spend_cap_usd` | **yes** | The maximum generation spend on this run. Refer to the section below. | | `model` | no | Only `sume-agent`. If you do not send it, you get the same agent. | | `input` | no | Caller data. Sume writes all of it to `/workspace/inputs/sume-action-input.json`. The prompt has a bounded pointer to the file. Sume uses this value only as data, not as instructions. | | `attachments` | no | A maximum of 30 images that the agent can see and use. Refer to [Attachments](#attachments). | | `output_schema` | no | Bind the `output` of the run to your own schema. The contract is the same as for Action runs. | | `primary_output_key` | no | The `output` key that holds the headline result. | | `communication.webhook_url` | no | Public HTTPS URL that Sume notifies when the run gets to a terminal status. Refer to [Run webhooks](/agents/run-webhooks). On `api.dev.sume.com` and `api.sume.com`, Sume accepts, stores, and delivers it. | `Idempotency-Key` works the same as on Action runs. If you send a key again, the API returns the original receipt with `idempotency_hit: true`. If you use the key again with a different payload, the API returns `409 idempotency_conflict`. ##### `messages[]` ```json { "messages": [ { "role": "system", "content": "Be terse." }, { "role": "user", "content": "Summarize https://example.com/p" } ], "generation_spend_cap_usd": 2 } ``` `content` accepts a string or an OpenAI-style `[{ "type": "text", "text": "..." }]` array. Sume joins the turns, in order, into one prompt. `content` also accepts `{ "type": "input_text", "text": "..." }` as an alias for `text`, and `{ "type": "input_image", ... }` parts. Refer to [Attachments](#attachments). The API **rejects** `assistant` turns. It does not ignore them. Acceptance of these turns implies that Sume replays a prior conversation. This endpoint does not do that at this time. Each completion runs in a new thread. The `thread_id` in the receipt identifies that thread. #### Attachments You can send images that the agent can actually look at. Send them at the top level or as `input_image` content parts. Both forms use the same item shape. Sume merges the two sources into one list. ```bash curl -sS -X POST "https://api.sume.com/v1/agent/completions" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "messages": [ { "role": "user", "content": [ { "type": "input_text", "text": "Describe this product shot." }, { "type": "input_image", "image_url": "https://cdn.example.com/shot.jpg" } ] } ], "output_schema": { "name": "caption", "schema": { "type": "object", "properties": { "caption": { "type": "string" } }, "required": ["caption"], "additionalProperties": false } }, "generation_spend_cap_usd": 2 }' ``` A turn with only images is permitted. If you do not send the text part, Sume tells the agent to use the attached files. You can use attachments and `output_schema` together. The images go to the agent. After the run completes, Sume still parses the `output` of the run against your schema. For the item shape, limits, upload path, and error codes, refer to [Format runs](/formats#attachments). These items are the same on both surfaces. #### The spend cap is required `generation_spend_cap_usd` has no default. If you do not send it, the request fails with `400 invalid_request`. This rule is intentional. An Agent Completion is an unattended agent. It has tools and access to your generation wallet. In the chat UI, an interactive spend-approval prompt protects you. A backend caller does not get this prompt. The cap replaces the prompt. Set the cap for each run to the maximum spend that you accept for that run. To find the correct cap, read the metered rates that the run will use on the [API pricing page](https://www.sume.com/pricing/api). #### Poll for the result ```bash curl -sS "https://api.sume.com/v1/agent-runs/$RUN_ID" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The statuses are the same as for Action runs: `queued`, `processing`, `completed`, `failed`, `canceled`. Poll `status_url` until the value of `next_action` is not `poll_status`. A completed run fills `output`. By default, this output has the `sume/action-run-output/v1` shape. The last text of the Agent is in `output.text`. Any generated media is in `output.images`, `output.videos`, `output.audio`, and `output.files`. The run also fills `artifacts`, and it records the spend in `usage`. Media URLs are durable `media.sume.com` HTTPS URLs. To stop a run that is in progress: ```bash curl -sS -X POST "https://api.sume.com/v1/agent-runs/$RUN_ID/cancel" \ -H "Authorization: Bearer $SUME_API_KEY" ``` `GET /v1/agent-runs` lists your completions, newest first. #### Errors | Status | Code | Cause | |---|---|---| | `400` | `invalid_request` | Missing `generation_spend_cap_usd`, neither or both of `instruction`/`messages`, an `assistant` turn, a malformed `input`, or a `model` other than `sume-agent`. | | `400` | `invalid_attachment` | Bad attachment item: wrong `type`, missing or non-HTTPS URL, both `image_url` and `asset_id`, or a source that is not a permitted image. | | `400` | `attachment_not_found` | `asset_id` is unknown in this workspace. | | `413` | `attachment_too_large` | An image is more than 30 MB, or the total of the set is more than 500 MB. | | `502` | `attachment_fetch_failed` | Sume could not fetch the image. Causes: unreachable host, hotlink protection, or a non-2xx response. | | `403` | `insufficient_scope` | The key does not have `agent_completions:*`, or it is a service-account key. | | `404` | `agent_run_not_found` | Unknown run id, or a run that another account owns. An Action or Format run id will not resolve here. | | `409` | `idempotency_conflict` | You used the `Idempotency-Key` again with a different payload. | #### Not available yet - Non-image attachments. At this time, `input_image` is the only `type`. PDFs and other files will come later. - Streaming, and a synchronous OpenAI-compatible `choices[]` response. - Continuation of a prior thread with `thread_id`, and `assistant` turns in `messages[]`. - Team-owned threads. Completions are user-owned. ### Run webhooks Source: https://docs.sume.com/agents/run-webhooks.md Receive one signed POST when an Action, Format, or Agent Completion run completes or fails. Each run surface accepts a `communication.webhook_url`. When the run **completes** or **fails**, Sume sends **one** signed POST to that URL. The POST contains the same receipt that the poll endpoints return. The delivery uses the same shape as the fal webhook. The envelope `status` is `OK` or `ERROR` for those outcomes, but **not** for cancel. A webhook is the alternative to a poll loop. You still get `status_url` and `result_url`, and you can still poll. A webhook only means that you do not have to run a loop for each run. #### Availability | Environment | Delivery | | ---------------------------------- | ----------------------------------- | | Development — `api.dev.sume.com` | **Live.** Sume calls your endpoint. | | Production — `api.sume.com` | **Live.** Sume calls your endpoint. | If you supply `communication.webhook_url`, Sume arms delivery in both environments. You can still poll `status_url` / `result_url` as a backup. **Any run whose caller already supplied a `webhook_url` can receive a POST**. This includes older runs that get to a terminal state after enablement. Register only the endpoints where you still want traffic. This page is about **run** webhooks. Generation-**job** webhooks (`job.completed` and related events, from `POST /v1/models/...`) are a separate surface with a separate event set. Refer to [Webhooks](/workflows/webhooks). The signature scheme is the same, so one verifier works for both. #### Ask for a webhook When you start the run, send `communication.webhook_url`. It works the same on all three surfaces. | Field | Notes | |---|---| | `communication.webhook_url` | Public HTTPS URL, maximum 2048 characters. Sume rejects localhost, private-network, and non-HTTPS URLs with `400 invalid_request`. | | `communication.callback_url` | Accepted alias for `webhook_url`. The behavior is the same. Send one or the other. | | `communication.mode` | `async` (default) or `webhook`. The URL arms delivery. `mode` is only descriptive. | | Top-level `webhook_url` / `callback_url` / `mode` | fal-shaped aliases. Sume normalizes them into `communication.*`. The same value on both layers is permitted. Values that do not agree return `400 invalid_request`. | Sume validates the URL as a public HTTPS URL again at delivery time, not only when you submit. Sume does not follow redirects. A `3xx` is not a delivery. #### Events Each run family has one terminal event. The outcome is in `status` and `payload.status`, not in the event name. | Surface | Event | Receipt `object` | |---|---|---| | Action runs | `action.run.terminal` | `action.run` | | Format runs | `format.run.terminal` | `format.run` | | Agent Completions | `agent.run.terminal` | `agent.run` | A partner that embeds only Formats can route on `event === "format.run.terminal"` and does not have to examine the body. ##### One per turn, not one per artifact A run is one **agent turn**. Thus, its terminal event fires exactly one time. The number of clips, images, or intermediate files that the turn produced does not change this. There is no per-artifact run event, and we do not plan one. This fact is important for [continued runs](/formats/runs#continue-a-run). When you continue a run, Sume starts a **new** run. The new run delivers its own single terminal webhook with the new run id. The webhook of the original run already fired and will not fire again. If you want progress *inside* a turn, use the generation-**job** layer ([Webhooks](/workflows/webhooks)). That layer fires one event for each job when the job completes. Job events name a job, not a step of your recipe. #### Payload ```json { "event": "format.run.terminal", "request_id": "arun_e43e6c5cb2b74052", "run_id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "OK", "outcome": "ok", "created_at": "2026-08-05T09:00:00.000Z", "payload": { "id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "completed", "format": { "id": "skl_...", "slug": "product-promo", "title": "Product promo", "version": 3 }, "output": { "text": "...", "videos": [] }, "primary_output_url": "https://media.sume.com/artifacts/artf_.../out.mp4", "artifacts": [], "usage": { "currency": "USD", "billable_amount_usd_micros": 240000, "generation_spend_cap_usd_micros": 1000000 }, "status_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/status", "result_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/result", "cancel_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/cancel", "request_id": "arun_e43e6c5cb2b74052" }, "error": null } ``` | Field | Notes | |---|---| | `event` | Refer to the table above. | | `request_id` | Equals `run_id`. Stable across retries. Use it to dedupe. | | `run_id` | The run that this delivery is about. | | `object` | The `object` of the receipt. | | `status` | `OK` when the run completed, `ERROR` when it failed. Binary. Refer to `outcome`. | | `outcome` | `ok`, `degraded`, or `error`. **Branch on this** when the question is "did I get usable output". | | `created_at` | The time when Sume built this delivery body. Use it to put deliveries in order. `request_id` cannot do this, because it is stable across retries. | | `payload` | The run receipt. `null` only on overflow. Refer to the section below. | | `error` | `null` when `status` is `OK`. In other cases, `{ code, message }`. | ##### `OK` does not always mean you got output A run can complete, bill you, and produce real media in `artifacts[]`, but still fail to project that media into your `output_schema`. On that path, `status` is `OK`, because the run really completed. But `output` is `null`, and `output_error` gives the reason. `outcome: "degraded"` is the name for this condition. ```ts switch (event.outcome) { case "ok": return ship(event.payload.output); case "degraded": // Real artifacts, no structured output. Usually a schema that asks for a // field the Format never produces. return reviewManually(event.payload.artifacts, event.payload.output_error); case "error": return retryOrAlert(event.payload?.error ?? event.error); } ``` A handler that uses only `status` continues to work. But it cannot tell the difference between `ok` and `degraded`. ##### Two different `request_id`s The `request_id` of the envelope is the **run id**. It is the dedupe key, and it is deliberately stable across retries. The receipt value at `payload.request_id` is a *correlation* id for the call that produced the receipt. In the webhook, this value is the run id. When you read the same receipt from `GET /v1/format-runs/{run_id}`, this value is an HTTP `req_…` id. Dedupe on the `request_id` of the envelope (or `run_id`, which equals it). Ignore the nested value. `usage.billable_amount_usd_micros` contains the generation spend that Sume attributes to the run. If Sume cannot read that spend, `usage` is `null`. Refer to [Runs and results](/agents/actions/runs#the-run-receipt). `GET /v1/usage` stays the authoritative billing record. ##### `payload` is the receipt `payload` is byte-identical to the `data` object of `GET /v1/{family}-runs/{run_id}` for the same run. The poll response wraps it in `{ "data": ... }`. The webhook does not wrap it. ```ts // One handler, two transports. handleRun(webhookBody.payload); handleRun((await fetchRun(runId)).data); ``` The same code path that the poll endpoint calls builds the payload. Thus, the two cannot become different. ##### Failure, cancelation, and skips A **failed** run comes with `status: "ERROR"` and a populated `error`. `payload` is still the full receipt. The receipt of a failed run contains `artifacts` and `output_error`, and usually you want them. ```json { "event": "format.run.terminal", "request_id": "arun_e43e6c5cb2b74052", "run_id": "arun_e43e6c5cb2b74052", "object": "format.run", "status": "ERROR", "payload": { "id": "arun_e43e6c5cb2b74052", "status": "failed", "error": { "code": "format_run_failed", "message": "..." } }, "error": { "code": "format_run_failed", "message": "..." } } ``` `error.code` is a copy of `payload.error.code` when the receipt has one. Usually it is a specific reason such as `output_schema_unsatisfied`. If not, it is the generic code of the family: `action_run_failed` / `format_run_failed` / `agent_run_failed`. **A `canceled` run does not deliver a webhook**. Cancel is a separate API path (the same idea as the fal queue cancel, with no cancel webhook status). After you `POST …/cancel`, accept the cancel response as correct. Then poll `status_url` until `payload.status` is `canceled`. Do not wait for a POST. **A `skipped` run does not deliver a webhook**. `on_active_run: "skip"` records a terminal run immediately and does not start work. Thus, there is no completion to tell you about. The create response already told you. Read `status` on the response that you got. Do not wait for a POST that will not arrive. The defaults are different for each surface. The **Format** default is `allow` (concurrency). The **Action** default is `skip`. For Formats, refer to [Calling a Format](/formats/call#request-body). For Actions, refer to [Action API trigger](/agents/actions/api-trigger#request-body). ##### Oversized receipts Sume cannot deliver a receipt of more than **1 MiB** inline. Sume sends the envelope with `payload: null` and an error that tells you where to fetch the receipt: ```json { "status": "OK", "outcome": "ok", "payload": null, "error": { "code": "payload_too_large", "message": "Run receipt exceeded the 1048576-byte webhook body limit. Fetch the receipt from result_url instead.", "result_url": "https://api.sume.com/v1/format-runs/arun_e43e6c5cb2b74052/result" } } ``` `status` still reports the real outcome of the run. If a run succeeded but was too large to ship, the run did not fail. #### Signature Sume uses HMAC-SHA256 over `.` to sign the raw JSON body. ```text content-type: application/json x-sume-webhook-timestamp: 1785000000 x-sume-webhook-signature: sume-v1= ``` Verify the signature against the **raw** request body, before any JSON parse or re-serialize. Reject a timestamp that is outside your replay window. Five minutes is a good default. You can read your signing secret on the **Webhooks** tab of the dashboard (`/dashboard/webhooks`). You can also get it from `GET /v1/webhooks/signing-secret` with any API key that has `account:read`. Sume derives the secret for your workspace. The secret of a different workspace cannot verify a delivery that Sume signed for you. Store the secret as `SUME_COM_WEBHOOK_SIGNING_SECRET` (the same name that the delivery worker uses when it signs). Each delivery has `x-sume-webhook-secret-fingerprint`. `webhook_delivery.signing_secret_fingerprint` on the run receipt repeats this value. Compare it with the fingerprint next to the secret in the dashboard. Then you know that both sides have the same secret, and you do not send the secret anywhere. In TypeScript, [`@sume-com/sdk`](/sdk) ships this check (refer to [Verifying webhooks](/sdk/webhooks)): ```ts import { verifyWebhook } from "@sume-com/sdk"; const ok = await verifyWebhook({ body: rawBody, headers: request.headers, secret: process.env.SUME_COM_WEBHOOK_SIGNING_SECRET!, }); ``` This is the full scheme. Use it in a receiver that cannot use the SDK, for example another language or a gateway in front of your app: ```ts import crypto from "node:crypto"; export function verifySumeWebhook({ rawBody, timestamp, signatureHeader, secret, toleranceSeconds = 300, }: { rawBody: string; timestamp: string; signatureHeader: string; secret: string; toleranceSeconds?: number; }) { const ts = Number(timestamp); if (!Number.isFinite(ts)) return false; if (Math.abs(Math.floor(Date.now() / 1000) - ts) > toleranceSeconds) { return false; } const digest = crypto .createHmac("sha256", secret) .update(`${ts}.${rawBody}`) .digest("hex"); const expected = `sume-v1=${digest}`; const actual = Buffer.from(signatureHeader); const expectedBuffer = Buffer.from(expected); if (actual.length !== expectedBuffer.length) return false; return crypto.timingSafeEqual(actual, expectedBuffer); } ``` This is the same verifier that validates generation-job webhooks. Write it once. #### Delivery behavior | Property | Value | |---|---| | When | One time for each run, when the run completes or fails. | | Success | Any `2xx`. | | Retries | A maximum of 10 attempts in total. Then `webhook_delivery.status` is `exhausted`. | | Backoff | `min(max(30s × 2^(attempt−1) with jitter, Retry-After), 1h)`. Obey `Retry-After` on 429/503. | | Timeout | 10s per attempt. | | Redirects | Sume does not follow them. A `3xx` is a failed attempt. | First, record the event durably. Then return `2xx` quickly. Process the event after you return the response. A slow endpoint uses all of the 10-second attempt budget, and Sume retries the delivery. Dedupe on `request_id`. It is the same value on each retry of the same run. A delivery outcome does not change the run. If an endpoint rejects all ten attempts, you have a failed *delivery* and a run that is still `completed`. Fetch the run from `result_url`. #### Send test and Redeliver There is one Send test control, on `/dashboard/webhooks` (also `POST /v1/webhooks/test-deliveries`, `account:write`). It sends a dummy `webhook.test` payload. It is **not** a replay of a real Format run. To replay a real terminal call, use **Redeliver** on that delivery row, or: ```http POST /v1/format-runs/{run_id}/webhook/redeliver ``` The key must have `formats:write`. The body is empty. The endpoint re-POSTs the current `format.run.terminal` receipt with a new timestamp and signature. Sume signs it with the **same secret** as the original delivery. Thus, `x-sume-webhook-secret-fingerprint` (and `webhook_delivery.signing_secret_fingerprint` on the row) is the same 12 characters, and your verifier does not change. This still works after the automatic attempts are exhausted. It does not use one of the automatic 10. The response has `redelivery` (`delivered`, `status_code`, `error`) for the POST that you triggered. `webhook_delivery` describes the row. If a redeliver fails for a call that Sume delivered before, the row keeps that earlier 2xx. Sume only counts the attempt in `manual_redeliveries`. If the run had no `webhook_url`, the endpoint returns `409 webhook_not_configured`. If the run is still in progress, it returns `409 run_not_terminal`. A run that you cannot see returns `404 format_run_not_found`. If the key does not have `formats:write`, the result is `403 insufficient_scope`, not 404. Dedupe on `request_id` / `run_id`. Redeliver does not send to a different URL. [Webhooks](/workflows/webhooks) documents job redeliver. #### Next - [Verifying webhooks](/sdk/webhooks) — `verifyWebhook` and `SUME_COM_WEBHOOK_SIGNING_SECRET` - [Runs and results](/agents/actions/runs) — the receipt, field by field - [Advanced: run a schedule via API](/agents/actions/api-trigger) - [Calling a Format](/formats/call) - [Embed a Format in your product](/cookbooks/embed-a-format) — the whole partner integration, end to end - [Webhooks](/workflows/webhooks) — generation-job webhooks, the other surface ### Safe automation Source: https://docs.sume.com/agents/safe-automation.md Guidelines for using Sume APIs and tools from agents without leaking secrets or crossing workspace boundaries. Use the public API, CLI, and MCP in a way that keeps each workspace isolated and keeps billing clear. Related: [MCP tools and gates](/mcp/tools-and-gates), [CLI security](/cli/security). #### Workspace isolation The API key or the app session selects the workspace. Tools must not accept a workspace id from the user. The only exception is a product that explicitly supports a change between workspaces. #### Credit-spending actions Generation and analysis creation can spend credits. Agent tools must make those actions explicit. Agent tools must also keep read-only operations separate. On hosted MCP, the best choice for read-only exploration is OAuth `mcp:read`. Before you use paid tools such as `generate_image` / `avatars_create`, grant `mcp:write` (or use an API key). Write calls and paid calls must send `idempotency_key`. The legacy `allow_write` / `allow_paid` arguments are optional. #### Logging Safe logs include: - request ids, - job ids when necessary, - high-level status, - sanitized media metadata. Unsafe logs include: - API keys, - signed URLs, - raw private media URLs, - too much user content or too many transcripts. ## Models ### Overview Source: https://docs.sume.com/models.md Sume generation families — Avatar, Image, Video, Music, and related utilities. Sume makes job-backed generation families available on the public Developer API. Each family has a primary product URL. Most families also have a model-run alias. Submit a request with an `Idempotency-Key`. Poll the job, then read the result. To discover the current production capabilities, use [catalog](/api/reference) (`GET /v1/catalog`) and the OpenAPI snapshot. The detailed guides below agree with the **production** OpenAPI request schemas (`api.sume.com` / docs snapshot). #### Avatar 1.0 Avatar 1.0 is a two-step workflow: 1. Create a reusable avatar. 2. Use this avatar to generate talking videos from scripts or multi-scene inputs. For new integrations, the canonical product routes are preferred: | Step | Canonical route | |---|---| | Create avatar | `POST /v1/avatar-1.0/generate` | | List / read avatars | `GET /v1/avatar-1.0/avatars`, `GET /v1/avatar-1.0/avatars/:id` | | Create talking video | `POST /v1/avatar-1.0/talking-video` | | List / read videos | `GET /v1/avatar-videos`, `GET /v1/avatar-videos/:id` | Legacy model-run aliases, for example `POST /v1/models/sume/avatar/v1.0/runs` and `POST /v1/models/sume/avatar-video/v1.0/runs`, are still supported for compatibility. Each guide page gives the details. First, create an avatar from a prompt, a reference photo, or a supported avatar input. Avatar creation is job-backed. Thus, Sume returns a job first. When the job completes, the avatar becomes a reusable resource in your workspace. When possible, use a stable avatar handle. The handle gives your app or agent a simple name to use again later. Thus, your app or agent does not depend only on a generated id. After the avatar is ready, send a script (or `video_inputs`) and the avatar handle to create an avatar video. Each video is also job-backed. Submit the request, and poll or wait for completion. Then read the result URL. Avatar Video supports `quality: "standard" | "plus" | "max"`. To use the default **`plus`** execution path, omit it. Use `standard` for the fastest path. Use `max` when quality is more important than turnaround. ##### Related Avatar utilities | Guide | Use for | |---|---| | [Create your avatar](/models/avatar) | Avatar creation request. | | [Generate avatar video](/models/avatar-videos) | Talking video from a ready avatar. | | [Avatar video previews](/models/avatar-video-previews) | First-frame stills before a full render. Then `generate-video`. | | [Face swap (Beta)](/models/face-swap) | Swap a ready avatar face onto a public source video. | | [Video captions](/models/video-captions) | Burn captions onto a public video URL that you already have. | | [Video inspect](/models/video-inspect) | Probe + stills + optional STT of one `media.sume.com` clip. Default clip inspection on dest and prod. | | [Reference ingest](/models/reference-ingest) | One reference clip → `ReferenceVideoManifest` (frame-exact shots, source-resolution OCR text tracks, audio facts, labeled strip). Sume bills it by its Modal compute (+ STT when it transcribes). Dest first. | | [hypit Understand](/models/hypit-understand) | One reference clip → Hypit's understand artifacts (probe, word-timed transcript with scores, cut candidates, time- and word-labeled grids, labeled excerpts, notes). Each is a durable artifact with an id. `align` binds a script segment to the words of a generated take. `hypit_compose` renders word-anchored SVML on HyperFrames. Sume bills it by its Modal compute, WhisperX included (transcribe with `engine: scribe_v2` bills STT, notes / align are free, and the final bake bills as a render). Dest only. | | [Video frames](/models/video-frames) | Exact stills at `at[]` / `fps` from one hosted clip. Source-size `artf_` images. Sume bills it by its Modal compute. Always `202`. | | [Video trim](/models/video-trim) | `[start, end)` of one hosted clip → new MP4. `$0.02` flat. Material for timeline, not placement. | | [Audio detach](/models/audio-detach) | Audio track of one hosted video → durable wav / mp3. `$0.01` flat. | | [Video filter](/models/video-filter) | Dim / crop / allowlisted pixel graph on one hosted clip → new MP4. `$0.02` encode. `/check` is free. | | [Timeline 1.0](/models/timeline) | Audio spine + ordered `video[]` → one MP4. `$0.10` / ceil(output minute). The assembly surface. | | [Timeline compose](/models/timeline-compose) | Still + video in one frame (반배너 / overlay) → one MP4 shot. `$0.02` flat. | | [Timeline audio](/models/timeline-audio) | Concat / split Sume-hosted audio → durable files. `$0.01` flat. | | [Video analyses](/models/video-analyses) | Legacy `vana_` resource. Dest create is `410`. Prod still accepts it until #5953 PR-C2. | | [Trending videos](/models/trending-videos) | Discover TikTok trending video metadata for research. | #### Image, Video, Music, Fabric These are managed product models. Sume selects the providers. Callers do not send provider queue ids. VEED Fabric 1.0 is public as `veed/fabric-1.0`. | Family | Primary URL | Guide | |---|---|---| | Image 1.0 | `POST /v1/image-1.0/generate` | [Image 1.0](/models/image) | | Image API | `POST /v1/images` | [Image API](/models/images) | | Video 1.0 | `POST /v1/video-1.0/generate` | [Video 1.0](/models/video) | | Video generation | `POST /v1/videos` | [Video generation](/models/videos) | | Music Router | `POST /v1/music-router/generate` | [Music Router](/models/music-router) | | Music 1.0 (retiring, resolves through Music Router) | `POST /v1/music-1.0/generate` | [Music 1.0](/models/music) | | VEED Fabric 1.0 | `POST /v1/veed/fabric-1.0` | Talking still + audio clips (`veed/fabric-1.0`) | | MiniMax H3 Max Lip Sync | `POST /v1/minimax/h3-max/lip-sync` | Same still + audio body as Fabric, audio 5–14.8 s, list × 1.25 (`minimax/h3-max/lip-sync`) | ##### Fabric alias migration **Deprecated compatibility aliases that Sume will retire — still supported:** | Deprecated route | Correct call | |---|---| | `POST /v1/avatar-1.0/image-to-video` | `POST /v1/veed/fabric-1.0` | | `POST /v1/models/sume/avatar-1.0/image-to-video/runs` | `POST /v1/models/veed/fabric-1.0/runs` (or `POST /v1/veed/fabric-1.0`) | The public model id is `veed/fabric-1.0`. Both aliases still accept the same body. Thus, current Mobidoo / live-commerce integrations can migrate if they change only the URL. Send `audio_url`, measured `duration_seconds`, and only one visual source. The preferred source is the `image_url` of the generated, inspected posed still. Use `avatar_handle` only when the user named that avatar. You cannot send the two sources together. With MCP, the same body is inside `payload` on the supported `avatar-image-to-video_create` tool. Each shot where a person speaks on camera is Fabric with an accepted still + TTS. This rule is applicable to short UGC and presenter ads, testimonials, Recreate beats where a person speaks, and LC / long-form host talk. Video models do not lip-sync to generated TTS or to a later voice-over. Thus, a face that talks is never a video-model clip with narration under it. Wordless beats, B-roll and product motion use Auto image → inspect → Auto video, without Fabric or Stage P. Model-run aliases use `POST /v1/models/sume/-1.0/runs` (same request body) for Image 1.0 / Video 1.0 / Music 1.0. If you do not need an explicit catalog model id, Image 1.0 + `routing_preset` is the preferred method. Video 1.0 (`POST /v1/video-1.0/generate`) is deprecated. It maps to `sume/auto` and ignores `routing_preset`. Thus, send `model: "sume/auto"` or an explicit catalog id (for example Seedance). Pin that id on [Video generation](/models/videos) (`POST /v1/videos`) or [Image API](/models/images) (`POST /v1/images`). Legacy `/v1/video-router/*` stays registered as a Sume-envelope alias (refer to [Video Router](/models/video-router)). #### Shared job lifecycle These families return the same job envelope pattern: 1. Submit → store `job.id`, `status_url`, `result_url`. 2. Poll the status (or wait with `mode: sync` / `subscribe` for a maximum of 30s). 3. When `result_ready` is true, fetch `/result`. 4. Read Sume-hosted `artifacts` for media jobs. Refer to [Jobs and results](/workflows/jobs-and-results), [Generation admission](/workflows/generation-admission), [Media inputs](/workflows/asset-library), and [API recipes](/api/cookbook). #### Ahead of production OpenAPI Some more generators (for example STT) can appear on `api.dev.sume.com` before the production OpenAPI snapshot lists them. Until these paths land in the docs OpenAPI snapshot and the `api.sume.com` reference, they are not production Developer API surface. ##### Experimental: `POST /v1/avatar-1.0/fabric` `sume/avatar-1.0/fabric` is a **temporary, test-only** route. Sume uses it to compare a different talking-clip backend against `POST /v1/veed/fabric-1.0` (legacy `POST /v1/avatar-1.0/image-to-video`) on identical inputs. It takes the same request body. Thus, a comparison script changes only the path. Differences from `image-to-video`: | | `image-to-video` | `fabric` (experimental) | |---|---|---| | `duration_seconds` | 1–300 | 1–15 (rounded up, and the route rejects requests over 15) | | `speed_tier` | selects a provider speed tier | accepted and ignored | | Price | $0.1875/s @720p | $0.3024/s @720p | Do not build production integrations on this route: - The name `fabric` is a placeholder test name and **will change before general availability**. - Sume can change the route or remove it completely after the comparison is complete. - VEED Fabric 1.0 (`veed/fabric-1.0`) stays the supported path for talking clips. This includes Live Commerce and the MCP tools. ### Create new avatar Source: https://docs.sume.com/models/avatar.md Create a reusable avatar with Prompt, Profile, or Image inputs. There are three ways to make an avatar: 1. **Prompt**: describe the avatar that you want. 2. **Profile**: provide structured traits for the avatar. 3. **Image**: use a reference image. Each request creates a job. Poll the job until it completes. Then use the returned avatar handle or resource id to generate avatar videos. The canonical Avatar 1.0 route is preferred: ```text POST /v1/avatar-1.0/generate ``` Avatar creation uses a top-level `avatar_handle` and an `input` union. The handle can start with `@`. Sume normalizes the handle and stores it without `@`. #### 1. Prompt Use this method when you want to create an avatar from text only. #### 2. Profile Use this method when your app already has profile details for the avatar. In the API, this method uses the `props` input type. ```bash curl -X POST https://api.sume.com/v1/avatar-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: avatar-profile-001" \ -d '{ "avatar_handle": "product_host", "input": { "type": "props", "ethnicity": "Asian", "sex": "female", "age": 28 } }' ``` #### 3. Image Use this method when you have a reference image. In the API, this method uses the `photo` input type. ```bash curl -X POST https://api.sume.com/v1/avatar-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: avatar-image-001" \ -d '{ "avatar_handle": "reference_presenter", "input": { "type": "photo", "image_url": "https://example.com/reference.png" } }' ``` `image_url` must be a fetchable public HTTPS image URL. Before Sume submits the generation, it rejects localhost, private-network URLs, non-HTTPS URLs, and non-image responses. Refer to [Media inputs](/workflows/asset-library). #### Poll the job ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" ``` When the job is `completed`, fetch the result. ```bash curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Read avatar resources The Avatar 1.0 resource routes are preferred: ```bash curl https://api.sume.com/v1/avatar-1.0/avatars \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/avatar-1.0/avatars/avatar_123 \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Compatibility aliases These older paths are still supported and use the same request body: | Alias | Notes | |---|---| | `POST /v1/models/sume/avatar-1.0/generate/runs` | Canonical model-run alias. For new integrations, `/v1/avatar-1.0/generate` is preferred. | | `POST /v1/models/sume/avatar/v1.0/runs` | Legacy launch alias. | | `GET /v1/avatars`, `GET /v1/avatars/:id` | Compatibility list/read routes. The response shape is the same as `/v1/avatar-1.0/avatars`. | #### Next Use the returned avatar handle on [Generate avatar video](/models/avatar-videos). For first-frame review before a full render, refer to [Avatar video previews](/models/avatar-video-previews). ### Generate avatar video Source: https://docs.sume.com/models/avatar-videos.md Generate talking avatar videos from a ready avatar, script or multi-scene inputs, and optional product or scene references. Avatar videos change a ready avatar into a script-driven talking video. The canonical Avatar 1.0 route is preferred: ```text POST /v1/avatar-1.0/talking-video ``` Launch requests use the top-level `avatar_handle` (or per-scene character fields inside `video_inputs`) to identify a ready avatar. Provide one of `script` or `video_inputs`, not both. Sume accepts scripts and multi-scene plans when it estimates the target video duration at 4-60 seconds inclusive. Make longer scripts shorter, or split them into multiple jobs. #### Create an avatar video ##### Optional product and scene inputs - Omit `product_image` for a productless avatar video. - Use `scene: { "type": "prompt", "prompt": "..." }` for scene direction. - Use `scene: { "type": "photo", "image_url": "https://..." }` for a photo scene reference. - Media fields must be fetchable public HTTPS URLs. Refer to [Media inputs](/workflows/asset-library). ##### Quality Avatar Video accepts `quality: "standard" | "plus" | "max"`. | Value | Behavior | |---|---| | `plus` | **Default** when omitted. Balanced quality path. | | `standard` | Fastest Sume execution path. | | `max` | Highest quality tier. Slower turnaround. | ##### Aspect ratio and resolution - `aspect_ratio` supports `1:1`, `3:4`, `9:16`, `4:3`, and `16:9`. Default: `9:16`. - `resolution` is `720p` at this time. #### Multi-scene `video_inputs` For scene hooks, demos, or silence beats in one composed video, use ordered `video_inputs`, not a single `script`. The total planned duration must still be in the 4-60 second window. ```bash curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: avatar-video-multi-001" \ -d '{ "avatar_handle": "sume_clawra", "aspect_ratio": "9:16", "quality": "plus", "video_inputs": [ { "id": "hook", "voice": { "type": "text", "script": "Wait, this turned one selfie into a whole video?", "duration": 3 }, "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }, { "id": "demo", "voice": { "type": "silence", "duration": 4 }, "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } }, { "id": "cta", "voice": { "type": "text", "input_text": "You pick a template, drop in your photo, and it builds the clip around you.", "duration": 5 }, "background": { "type": "prompt", "prompt": "Casual bedroom framing, native UGC lighting" } } ] }' ``` `voice.type: "silence"` is a beat with no speech. `duration` is required, and `script` / `input_text` are not permitted. Spoken scenes still use `type: "text"` with one of `script` or `input_text`, not both. The current execution supports one resolved avatar for each final video. It also expects that the scene backgrounds resolve to one shared scene. #### Inline captions After generation, the optional `captions` burns styles into the clean final MP4. It uses the spoken script / `video_inputs` text. Sume never adds captions to preview stills. ```json { "captions": { "enabled": true, "style": "slam", "language": "auto" } } ``` - `captions` takes the same four settings as standalone [Video captions](/models/video-captions): `style`, optional `font`, a `language` hint and `script_text`. - Styles: `slam` (default), `punch`, `tiktok-green`, `korean-ad` (Hangul karaoke for Korean speech), and the Hangul identities `weight-shift`, `black-outline`, `highlight`, `pill-karaoke`, `clip-wipe` and `editorial-emphasis`. - If a Korean script uses `style: "slam"` (or `punch` / `tiktok-green`), Sume rejects it with `400 caption_hangul_text_latin_style`. Sume does not change the style. Those font faces render Hangul as tofu. For Korean speech, select a Hangul style. - For inline captions, Sume rejects an estimated duration of more than 60 seconds. - A failure in the caption stage is a soft failure. The avatar job can still succeed with a clean primary `video_url` and `captions.status=failed`. - Inline captions do **not** create a separate billed video-caption job. To add captions to a public video URL that you already have, use [Video captions](/models/video-captions). #### Poll and recover ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/events \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Completed results can include public `media.sume.com` video artifacts. They can also include public-safe preview fields, for example `preview_image_url` and `scene_previews`. #### Read avatar-video resources ```bash curl https://api.sume.com/v1/avatar-videos \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/avatar-videos/avatar_video_123 \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Preview first, then generate To review first-frame stills before you pay for a full render, create an [Avatar video preview](/models/avatar-video-previews). Then call `generate-video` on the preview id. #### Compatibility aliases | Alias | Notes | |---|---| | `POST /v1/models/sume/avatar-1.0/talking-video/runs` | Canonical model-run alias. For new integrations, `/v1/avatar-1.0/talking-video` is preferred. | | `POST /v1/models/sume/avatar-video/v1.0/runs` | Legacy launch alias. Same body contract. | #### Related - [Create new avatar](/models/avatar) - [Avatar video previews](/models/avatar-video-previews) - [Face swap (Beta)](/models/face-swap) - [Video captions](/models/video-captions) - [Jobs and results](/workflows/jobs-and-results) ### Avatar video previews Source: https://docs.sume.com/models/avatar-video-previews.md Create first-frame Avatar Video previews, regenerate stills, then generate the final video. Avatar video previews generate the first-frame still stage. They do not start the full talking-video render. Use them when you want to approve the composition before you pay for a full Avatar Video generation. ```text POST /v1/avatar-video-previews GET /v1/avatar-video-previews/:id POST /v1/avatar-video-previews/:id/regenerate POST /v1/avatar-video-previews/:id/generate-video ``` #### When to use - Review the scene composition / first frames before a full render. - Multi-scene `video_inputs` where you want one still per scene. - Store caption intent on create. Then apply captions only at `generate-video` time (Sume never burns captions into preview stills). For a direct full render without the preview stage, use [Generate avatar video](/models/avatar-videos). #### Create a preview The create body has the same fields as Avatar Video. Provide one of `script` or `video_inputs`, not both. You can also provide the optional `product_image`, `scene`, `quality`, `aspect_ratio`, `title`, and `captions`. If you omit `quality`, the default is **`plus`** (`standard` | `plus` | `max`). The response includes the URLs to poll the job and an `avatar_video_preview_id`. Poll the job as for any other generation: ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Read the preview resource ```bash curl https://api.sume.com/v1/avatar-video-previews/avp_123 \ -H "Authorization: Bearer $SUME_API_KEY" ``` When the preview is ready, the public-safe fields include: - `preview_image_url` — primary still (scene 0 for multi-scene). - `scene_previews[]` — one still per input scene when available. - `resource_status` / `job_status` — these are preferred to the legacy `status` field for readiness vs job status. For multi-scene previews with a shared scene, later scene stills are pose-anchored continuations of the first frame. #### Regenerate stills Reuse the stored preview request (avatar, script/`video_inputs`, scene, quality, aspect ratio), and refresh only the first-frame stills: The call returns the same `avatar_video_preview_id` and a new preview-only job. #### Generate the final video When the preview is correct, start the normal Avatar Video workflow from the preview id. When the preview first frame is available, Sume uses it again. The captions that you stored on preview create apply at this step. An empty body (or `{}`) keeps the quality that you selected at preview create. The optional `quality` overrides only the **final render** tier. Preview stills are tier-independent, and Sume always uses them again. Thus, if you approve and then downgrade/upgrade, you do not need a new preview. Admission, pre-spend, ledger reservation, provider submit, and readback all use the effective (overridden) tier. Changes to structural fields (`script`, `video_inputs`, `avatar_handle`, `scene`, `aspect_ratio`) still need a new preview. Poll the returned job. Then read the avatar-video resource: ```bash curl https://api.sume.com/v1/avatar-videos/avatar_video_123 \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Constraints - Same duration window as Avatar Video: estimated 4-60 seconds inclusive. - Media inputs are URL-first public HTTPS fields (`product_image`, `scene.image_url`, and any scene background image URLs). - Sume stores inline captions from preview create for `generate-video`. Sume does not burn them into preview stills. - Preview stills are tier-independent. `generate-video` `quality` changes only the provider tier of the final video. - Exact request/response schemas: live [OpenAPI](https://api.sume.com/reference/json). #### Related - [Generate avatar video](/models/avatar-videos) - [Video captions](/models/video-captions) (standalone caption jobs) - [Jobs and results](/workflows/jobs-and-results) ### Face swap (Beta) Source: https://docs.sume.com/models/face-swap.md Beta Avatar Face Swap — apply a ready avatar face onto a public source video. Avatar Face Swap 1.0 is a **Beta** model-run endpoint. It creates a job-backed face-swap resource from a ready avatar handle and a public HTTPS source video. ```text POST /v1/models/sume/avatar-face-swap/v1.0/runs ``` This endpoint is not the old consumer-product `/face-swap` route. Use only the Developer API path above. #### When to use - You already have a ready Avatar 1.0 identity. - You have a short public source video, and you want to apply the avatar face to that video. - You do **not** need script-driven talking-video generation. For that task, use [Avatar videos](/models/avatar-videos). #### Create a face-swap job The required fields are `avatar_handle`, `video_url`, and `quality`. In Beta, `quality` is a required field (`standard` | `plus` | `max`). If you omit the field, this endpoint does not use a default value. ##### Hard constraints - `video_url` must be a fetchable public HTTPS video URL. - The endpoint rejects localhost, private-network, non-HTTPS, signed/private URLs, and provider task URLs. - In Beta, the worker validation is for source videos that are applicable to the face-swap process. The current plan is approximately **4-15 seconds** with usable audio. - By design, the endpoint does not support prompts, transcripts, duration knobs, aspect ratio, avatar ids in the body, or provider fields. #### Poll and recover ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` When a completed resource is ready, it exposes a public-safe `video_url` / artifacts under `media.sume.com`. Use `resource_status` as the primary signal for readiness. Use `job_status` when you poll. #### Communication modes This endpoint has the same options as other generation submits that OpenAPI documents: `async` (default-style immediate return), `sync` / `subscribe` with `wait_timeout_seconds`, and `webhook` with a public HTTPS `webhook_url`. #### Related - [Create new avatar](/models/avatar) - [Generate avatar video](/models/avatar-videos) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) ### Video captions Source: https://docs.sume.com/models/video-captions.md Burn captions onto an existing public video URL with style, language, and optional script alignment. Standalone video captions use a public HTTPS video URL as input. The API runs a job and returns a captioned video. We recommend this API when you already have a finished clip. For Avatar Video, you can also set inline `captions` on [talking-video](/models/avatar-videos). Or you can store the caption intent on [previews](/models/avatar-video-previews). ```text POST /v1/video-captions GET /v1/video-captions/:id ``` #### Create a caption job You must send `video_url`. These fields are optional: `style`, `font`, `language`, `script_text`, `words`, `cues` / `segments` (authored overlay text, without speech-to-text), and the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. Speech-to-captions (`script_text`, or STT if you do not send it) works only when the clip has audible speech. If the clip is silent, the job fails with `caption_no_speech` (`next_action: use_overlay_captions`). This error is not a generic policy rejection. To burn authored overlay text without ASR, send `cues` (or `segments`) with `text` + `start` + `end`. ##### Style, font and language | Field | Values / default | |---|---| | `style` | If you do not send it, the text of the captions sets the style: `slam` for Latin, `black-outline` for Korean. Or select `slam`, `punch`, `tiktok-green`, `korean-ad` (CapCut-style Hangul karaoke for Korean speech: one short phrase at a time, in the lower third, and the spoken word changes to a heavy weight. Use it with `language: "ko"`. Korean `script_text` is supported), or one of the Hangul identities below | | `font` | Optional Hangul face for the style that you select. If you do not send it, the style keeps its own face. Only for Hangul styles. Refer to [Fonts](#fonts). | | `design` | Optional overrides for one request. They change the colors, typography, placement, phrasing, and motion of the style. Refer to [Design overrides](#design-overrides). | | `language` | Speech-to-text hint (`ko`, `en`, …). If you do not send it, speech-to-text finds the language automatically. | `style` selects the look and the motion. `design` changes that look. `font` selects the face of the text. `language` only tells speech-to-text which language to expect. **`language` never selects the style or the font.** If you do not send `style`, the text of the captions sets the style. Korean text resolves to `black-outline`, and not to the Latin default, which burns Korean text as tofu. But if you select a style, Sume renders that style. ##### Design overrides A style is a set of design tokens. `design` overrides these tokens for one request. Each field is optional. Sume merges each field over the value of the style. Thus, one key changes one thing: ```json { "video_url": "https://example.com/clean.mp4", "style": "black-outline", "design": { "colors": { "active": "#22D3EE" } } } ``` This request burns `black-outline` with a cyan emphasis. All the other values of the style stay the same. | Group | Fields | |---|---| | `colors` | `base`, `active`, `stroke`, `accent`, `accent_deep`, `card` (`null` shows no card) | | `typography` | `base_weight`, `active_weight`, `active_scale`, `font_size_ratio`, `safe_width_ratio`, `stroke_width_px` | | `placement` | `anchor_ratio`, `landscape_anchor_ratio`: the center of the line as a fraction of the frame height | | `phrasing` | `max_words`, `max_chars`, `pause_seconds` | | `motion` | `enter_seconds`, `exit_seconds`, `emphasis_in_seconds`, `emphasis_out_seconds` | Colors are hex, `rgb()`/`rgba()`, or `transparent`. Sume rejects other CSS syntax and does not put it into the render document. A number outside its documented range gives a `400`. Thus, an incorrect look fails at request time, and you do not pay for an incorrect render. `punch` and `tiktok-green` do not support `design`. These styles still render on a path that reads none of these tokens. ##### Korean text on a Latin style is rejected `slam`, `punch`, and `tiktok-green` use Latin display faces that have no Hangul glyphs. If you send Korean text to one of these styles, the API returns `400` (`caption_hangul_text_latin_style`). The API does not change the style of the job. A caption style that you did not select is worse than an error. The alternative is a video full of tofu boxes, and that video has the same price as a good video. Latin text on `slam` works as before. This rule is also applicable to `font`. The faces below are Hangul faces. If you send one of these faces with a Latin style, the API returns `400` (`caption_font_requires_hangul_style`). ##### Hangul caption identities For Korean **speech** (talking head, UGC, creator voiceover), use one of these styles and not `slam`. The Latin display face of that style has no Hangul glyphs and renders Korean as tofu. These styles put words into phrase cards and do not show one word at a time. They burn the transcript accurately as written, with no case folding. | Style | Look | |---|---| | `black-outline` | CapCut white fill on a thick black outline, in the middle of the frame. The safe default. | | `weight-shift` | Phrase cards. The spoken word gets the heavy weight, and the other words go back to a light weight. | | `highlight` | An accent block moves in behind the word that the voice speaks. | | `pill-karaoke` | The card is in a dark pill. The color changes with the voice. | | `clip-wipe` | Each word comes in with a wipe from left to right. The most clear style at small phone sizes. | | `editorial-emphasis` | Left-aligned card with two lines. The lead words stay small. The last word of the phrase moves to a second line at approximately two times the size, in a display face. That word moves in from the margin. | `korean-ad` is the ad karaoke look (weight-shift, with an accent color on the spoken word). If you do not send `style`, the style does **not** resolve to this look. It resolves to `black-outline`. Thus, if you want the ad look, select `korean-ad`. If you do not send `style`, the render also has no accent color. In their original design, the two default styles have a gold tint on the spoken word. Sume does not add a tint to a look that you did not select. Thus, a default render keeps the fill color on the spoken word. If you select `black-outline` (or `slam`), that style keeps its own gold tint. `design.colors.active` sets the tint in the two cases. ##### Fonts This field is optional. If you do not send `font`, the style keeps its own face. The face is Pretendard for `korean-ad`, `weight-shift`, `highlight`, `pill-karaoke`, and `editorial-emphasis`. The face is Do Hyeon for `black-outline` and `clip-wipe`. If you set it, the style stays the same and the face changes. `editorial-emphasis` always shows its emphasis line in Black Han Sans, for all values of `font`. The look of this identity is the contrast between its two faces, and not one of the two faces. `font` changes its lead line, as it changes the single line in all the other styles. | `font` | Family | Feel | |---|---|---| | `pretendard` | Pretendard Variable | Clean baseline | | `do-hyeon` | Do Hyeon | Thick rounded, the CapCut classic | | `black-han-sans` | Black Han Sans | Impact | | `jua` | Jua | Soft cute rounded | | `dunggeunmo` | DungGeunMo | Pixel / retro | | `bagel-fat-one` | Bagel Fat One | Fat rounded | | `dongle` | Dongle Bold | Playful rounded display | | `gasoek-one` | Gasoek One | Ultra-thick impact | | `yeon-sung` | Yeon Sung | Brushy | | `single-day` | Single Day | Soft cute handwritten | | `hi-melody` | Hi Melody | Soft rounded cute | | `nanum-pen` | Nanum Pen Script | Handwritten | | `gowun-dodum` | Gowun Dodum | Soft editorial | | `gmarket-sans` | Gmarket Sans | Geometric retail display, the CapCut / YouTube title staple | | `noto-sans-kr` | Noto Sans KR | Workhorse sans | | `noto-serif-kr` | Noto Serif KR | Workhorse serif | | `ibm-plex-sans-kr` | IBM Plex Sans KR | Workhorse sans | | `gothic-a1` | Gothic A1 | Workhorse sans | | `hahmlet` | Hahmlet | Display | | `song-myung` | Song Myung | Display | | `poor-story` | Poor Story | Script | | `gamja-flower` | Gamja Flower | Script | | `stylish` | Stylish | Display | | `sunflower` | Sunflower | Display | | `nanum-gothic` | Nanum Gothic | Workhorse sans | | `nanum-myeongjo` | Nanum Myeongjo | Workhorse serif | | `gaegu` | Gaegu | Script | | `cute-font` | Cute Font | Script | | `east-sea-dokdo` | East Sea Dokdo | Script | All of these fonts use the SIL Open Font License 1.1, and the renderer includes them. The API rejects all other names and does not use a replacement. Thus, a caption never uses a fallback face that you did not select. `weight-shift` and `korean-ad` animate the `wght` axis. Only Pretendard has this axis. On a static face, these styles keep their color and scale emphasis, but they do not have the weight travel. ##### Restyling without transcribing twice To burn the same video again with a different style, send `source_caption_id` and not `video_url`: ```json { "source_caption_id": "…", "style": "black-outline" } ``` Sume uses the source video of that caption again, with the word timings that it already has. Thus, speech-to-text does not run a second time. Send `words` with it only to correct the text. The price does not change, because a restyle is still a render. You can also send `words` (word-level) or `cues` / `segments` (phrase-level overlay cards) with `text`, `start`, `end` in seconds. With these forms, Sume does not run speech-to-text. Sume burns that text accurately at those times. This is the path for silent clips. You can send only one of `script_text`, `words`, `cues`, and `segments`. ##### Optional `script_text` If you send this field, Sume keeps the speech-to-text word timings as the source of truth for time. Sume aligns the burned-in text to your script. The alignment can fail with these typed public job errors: - `script_alignment_mismatch` - `script_alignment_failed` Recommended next action: `simplify_script_text_or_omit`. To burn the STT text, do not send `script_text`. ```bash curl -X POST https://api.sume.com/v1/video-captions \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-caption-script-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/example/clean.mp4", "style": "punch", "script_text": "Say hello to the Sume developer platform." }' ``` #### Pricing note Each accepted standalone caption job reserves and captures **$0.20 USD** of Sume usage. This price is for videos of maximum 60 seconds, under the current fixed estimate. For the live price, refer to `GET /v1/catalog` and OpenAPI. Inline Avatar Video captions are a separate add-on in the avatar-video estimate. They do **not** create a `video_caption` resource. #### Poll and read ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/video-captions/vc_123 \ -H "Authorization: Bearer $SUME_API_KEY" ``` When the job is complete, the resource returns the public-safe status, the style, and the captioned `video_url` / artifacts. Raw transcripts, renderer internals, and signed source URLs are not part of the public contract. #### Constraints - `video_url` must be a public HTTPS video URL that Sume can fetch. - The API rejects localhost, private-network, non-HTTPS, signed/private, and provider task URLs. - The API does not support SRT uploads and provider task IDs. To send phrase-level text, use `cues` / `segments`. #### Related - [Generate avatar video](/models/avatar-videos) (inline captions) - [Avatar video previews](/models/avatar-video-previews) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) ### Video inspect Source: https://docs.sume.com/models/video-inspect.md Probe, sample stills from, and optionally transcribe one Sume-hosted clip. Default clip inspection on dest and prod. > **Current SoT.** For new clip inspection, use this surface and not > [video analyses](/models/video-analyses). > > - **Dest and prod:** `POST /v1/video-inspect` (MCP `video_inspect`). > Sume bills probe + stills by their Modal compute. The default is > `mode: sync`. > - **Not typed scenes.** Inspect returns probe facts, stills, and optional > STT. On dest, use `video_analyze` / `video_segment` for semantic > questions / intervals, but only when those names are in `tools_list`. > - **Legacy `POST /v1/video-analyses`:** dest returns `410 > video_analysis_retired`. On prod, it stays available until #5953 PR-C2. Video inspect 1.0 reads **one** `media.sume.com` clip that the workspace already owns. It never re-encodes the source, and it never makes an MP4. The job ID **is** the resource ID (`sume/video-inspect-1.0`, type `video_inspect`). ```text POST /v1/video-inspect GET /v1/video-inspect/:id ``` Hosted MCP: `video_inspect` (`packages/mcp-server/src/mcp.ts`). A write must have `idempotency_key` (and `mcp:write` under OAuth). There is no GET MCP wrapper. When the submit returns `202`, poll with `jobs_wait` and then `jobs_result`. #### Create an inspect You must send `video_url` (a `media.sume.com` artifact or asset of this workspace). The API does not fetch from the open internet. First, import the clip (`POST /v1/media-imports`). You must send `Idempotency-Key`. Optional fields: `frames`, `transcribe`, and (only with `transcribe: true`) `language_code`, `segmentation`, `duration_seconds`. You can also send the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. The default `mode` is **`sync`**. The handler waits a maximum of **30 seconds**. If the inspect completes in that time, the handler returns `200` with the result. If not, it returns `202` with the queued job, and you poll that job. ```bash curl -X POST https://api.sume.com/v1/video-inspect \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-inspect-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4" }' ``` If the submit is successful, the API returns `request_id` = `video_inspect_id` (the job ID) and a `video_inspect` object. Read the object later: ```bash curl https://api.sume.com/v1/video-inspect/$REQUEST_ID \ -H "Authorization: Bearer $SUME_API_KEY" ``` When the inspect is complete, the resource contains `probe` and `frames[{t,url,width,height}]` as durable `media.sume.com` image artifacts. If the request included a transcript, the resource also contains `transcript` (`text`, `words[]`, optional sentence `segments[]`, `audio_url`). It also contains `warnings[]`. #### Frames program | `frames` | Effect | |---|---| | not sent | **8** mid-bin stills (1 fps if the clip is shorter than 8 s) | | `false` | Only the probe, with no stills | | `{ at: [seconds…] }` | Explicit timestamps, 1–24 values, each ≥ 0 | | `{ fps: n }` | Sample rate, `0 < n ≤ 2`, mid-bin, with a limit of 24 | An object must contain **only one** of `at[]` or `fps`. Optional fields on that object: `format` `jpeg` (default) or `png`, and `max_edge` 64–2160 (default **768**). For a first-frame restage, send the source edge. `seek` on that object selects how Sume finds each still: | `seek` | Effect | |---|---| | `precise` (default) | Decodes to the accurate instant. Programs that you already use do not change. | | `fast` | Moves each still to the keyframe **at or before** its instant, and does not do the decode. The still can be earlier by a maximum of one GOP (approximately 0–5 s on typical sources), but it is never later. The quality, the resolution, and the transcript do not change. | Use `fast` for a **quick look** at a clip. Use `precise` when the timestamp must be accurate (explicit `at[]` instants, the default 8-still midpoints). With `fast`, each grid has `seek: "fast"`. `sample_times` are the instants that the tiles show. `requested_times` are the instants in the program. Limits (from `packages/api-contract/src/index.ts`): source ≤ **1800** s, and **24** stills for each call. For accurate frames at the source size at one `t`, use [video frames](/models/video-frames), not this route. #### Transcript (optional, billed) `transcribe: true` runs Sume STT 1.0 on the audio of the clip. Each inspect reserves its Modal compute ceiling. Sume captures each inspect at its own container seconds × Modal list × 1.25, plus the platform fee. The capture is never more than the hold. The transcript adds its per-minute rate to that reservation. - Public rate: **$0.01 per audio minute** (`STT_PUBLIC_PRICING`). For the live price, refer to `GET /v1/catalog`. - Without `duration_seconds`, Sume reserves **1 minute**. Maximum hint: **600** s. - `language_code` (for example `en` or `ko`) is an STT hint. If you do not send it, STT uses auto-detect. - `segmentation.mode: "sentence"` also returns sentence `segments[]` with no gaps (in the shape of caption lines). You can also set `silence_split_seconds` in the range 0.2–3. - If you send `language_code` / `segmentation` / `duration_seconds` without `transcribe: true`, the API returns `400 video_inspect_transcribe_required`. - A silent clip gives `inspect_source_has_no_audio`. First, examine `probe.has_audio` (a `frames: false` inspect is sufficient). #### Refusals (stable codes) | Code | When | |---|---| | `ffmpeg_fields_rejected` | The client sent `vf` / `filter` / `ffmpeg` / `cmd` / `codec` / `crf` / related fields. The server compiles ffmpeg. | | `video_inspect_frames_program_conflict` | The request has `frames.at[]` and `frames.fps` together. | | `video_inspect_frames_program_required` | `frames` is an object that has no `at[]` and no `fps`. | | `video_inspect_transcribe_required` | The request has STT-only fields without `transcribe: true`. | | `source_not_found` | The `media.sume.com` URL is dead or is not from this workspace. | | `frame_time_out_of_range` | An `at` value is outside `[0, duration)`. The error gives the duration. | | `inspect_source_has_no_audio` | The request has `transcribe: true` for a clip with no audio track. | The API rejects off-host URLs (`https://example.com/…`) at admit. Import the clip first. #### Not this surface | Need | Use | |---|---| | Typed `scenes[]` / TwelveLabs Pegasus | Legacy [video analyses](/models/video-analyses) (prod until PR-C2, dest `410`) | | Dest semantic Q&A / intervals | `video_analyze` / `video_segment`, when they are in the list on `mcp.dev.sume.com` | | The accurate frame at `t`, at the source size | [Video frames](/models/video-frames) | | Burn captions onto a public URL | [Video captions](/models/video-captions) | | A new MP4 cut | [Video trim](/models/video-trim) | | Audio track as wav / mp3 | [Audio detach](/models/audio-detach) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | Put several clips in a sequence | [Timeline 1.0](/models/timeline) | #### Related - [Video analyses](/models/video-analyses) (legacy `vana_` resource) - [Video frames](/models/video-frames) - [Video trim](/models/video-trim) - [Audio detach](/models/audio-detach) - [Video filter](/models/video-filter) - [Timeline 1.0](/models/timeline) - [Video captions](/models/video-captions) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Video frames Source: https://docs.sume.com/models/video-frames.md Extract stills from one Sume-hosted clip at times you name. Durable image artifacts at source size. Billed by its Modal compute. > **Current SoT.** Media L2 exact-frame extract (`video_frames`, #5831). It is > available on dest and prod. This API does **not** inspect a clip. For clip > inspection, use [video inspect](/models/video-inspect) (probe + sampled > stills + optional STT). This API does **not** make a new MP4. For range > cuts, use [video trim](/models/video-trim). Video frames uses **one** workspace `media.sume.com` clip and a program (`at[]` or `fps`) as input. It returns **durable** `media.sume.com` image artifacts. The source does not change. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `submitVideoFramesJob` / `getVideoFrames`, `apps/api/src/schemas.ts` `createVideoFramesSchema`, and `apps/worker/src/video-frames-executor.ts`). ```text POST /v1/video-frames GET /v1/video-frames/:id ``` This family has **no** `/v1/models/sume/…/runs` alias. The resource ID **is** the job ID (`request_id` = `video_frames_id`). A submit always returns **`202`**. `submitVideoFramesJob` pins `communicationMode: "async"`. Do not send `mode: "sync"`. It does not give a `200`. Poll this GET (or `GET /v1/jobs/:id/status`). Hosted MCP: `video_frames_create` / `video_frames_get` (`packages/mcp-server/src/mcp.ts`). A write must have `idempotency_key` (and `mcp:write` under OAuth). Flow: `video_frames_create` → `jobs_wait` → `video_frames_get`. This job does not use an admission seat (same class as a screenshot hop). At submit, it reserves its Modal compute ceiling. Sume bills the job by its own Modal compute (container seconds × Modal list × 1.25, plus the platform fee). The bill is never more than the hold. #### Create an extract You must send `video_url` (a `media.sume.com` artifact or asset of this workspace) **and only one** of `at[]` or `fps`. The API does not fetch from the open internet. First, import the clip (`POST /v1/media-imports`). MCP writes must have `Idempotency-Key`. Also send it on REST, so that a retry does not queue a second extract. Optional fields: `format` `jpeg` (default) or `png`, and `max_edge` **16–2160** (a long-edge clamp). If you do not send the clamp, the frames keep the source frame size. ```bash curl -X POST https://api.sume.com/v1/video-frames \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-frames-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "at": [0, 2.5] }' ``` If the submit is successful, the API returns `202` with `request_id` = `video_frames_id` (the job ID) and a `video_frames` object. Read the object later: ```bash curl https://api.sume.com/v1/video-frames/$REQUEST_ID \ -H "Authorization: Bearer $SUME_API_KEY" ``` When `resource_status` is `ready`, `frames[{t,url,width,height}]` are durable `artf_` images. If the extract fails for one instant, that frame has `url` `null`. This does **not** cause the job to fail. `source_duration_seconds` is the duration that the worker probed. #### Program | Field | Effect | |---|---| | `at[]` | Explicit seconds. **1–24** values, each ≥ 0. Each value must be in the range `0 <= t < duration`. If not, the worker fails with `frame_time_out_of_range` and gives the probed duration. | | `fps` | Sample rate, as an alternative to a list. `0 < fps ≤ 2`. Sume expands it to mid-bin samples (`0.5/fps`, `1.5/fps`, …), with a limit of **24** frames. | | `format` | `jpeg` (default) or `png` (lossless inspection). | | `max_edge` | Optional long-edge clamp, **16–2160**. If you do not send it, the frames keep the source size (the restage path). | Send **only one** of `at[]` or `fps`. Limits: source ≤ **300** s (`VIDEO_ANALYSIS_HARD_MAX_DURATION_SECONDS`), and **24** frames for each call. For evidence on the full clip (probe, eight mid-bin stills, optional STT), use [video inspect](/models/video-inspect). Inspect stills use a default `max_edge` of **768**. This route does not use that clamp. #### Refusals (stable codes) | Code | When | |---|---| | `ffmpeg_fields_rejected` | The client sent `vf` / `filter` / `filter_complex` / `select` / `ffmpeg` / `cmd` / `codec` / `crf` / `preset`. The server compiles the extract. | | `400` (schema) | The request has `at[]` and `fps` together, or has none of them. Or it has more than 24 `at` values, `fps > 2`, or a `video_url` that is not on `media.sume.com`. | | `frame_time_out_of_range` | An `at` value is outside `[0, duration)` (worker, after the probe). | | `duration_out_of_range` | The source is longer than 300 s (worker). | | `invalid_source_url` | The stored job has no `video_url` that the worker can use (worker). | The API rejects off-host URLs (`https://example.com/…`) at admit. Import the clip first. When the source is longer than 90 s (probe reuse), `warnings[]` can include `low_confidence_long_video`. This warning does not cause the job to fail. #### Not this surface | Need | Use | |---|---| | Probe / sampled stills / optional STT | [Video inspect](/models/video-inspect) | | A new MP4 cut | [Video trim](/models/video-trim) | | Audio track as a durable wav / mp3 | [Audio detach](/models/audio-detach) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | Put several clips in a sequence | [Timeline 1.0](/models/timeline) | | Look at an unrendered HyperFrames composition | `hyperframes_check` then `hyperframes_snapshot` | #### Related - [Video inspect](/models/video-inspect) - [Video trim](/models/video-trim) - [Audio detach](/models/audio-detach) - [Video filter](/models/video-filter) - [Timeline 1.0](/models/timeline) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Video trim Source: https://docs.sume.com/models/video-trim.md Cut a [start, end) range out of one Sume-hosted clip into a new MP4. Material preparation, not timeline placement. > **Current SoT.** Media L2 range cut (`sume/video-trim-1.0`, #5953). It is > available on dest and prod. This API does **not** inspect a clip. For clip > inspection, use [video inspect](/models/video-inspect). This API does > **not** assemble clips. The clip sequence, transitions, and the audio spine > stay on [Timeline 1.0](/models/timeline). Video trim 1.0 uses **one** workspace `media.sume.com` clip and a range as input. It returns a **new** MP4 artifact that contains only `[start, end)`. The source does not change. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `createVideoTrimV1` / `submitSumeVideoTrimJob`). ```text POST /v1/video-trim POST /v1/models/sume/video-trim-1.0/runs # same body, no extra `model` field ``` There is **no** `GET /v1/video-trim/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` Hosted MCP: `video_trim` (`packages/mcp-server/src/mcp.ts`). A write must have `idempotency_key` (and `mcp:write` under OAuth). Flow: `video_trim` → `jobs_wait` → `jobs_result`. #### Create a trim You must send `video_url` (a `media.sume.com` artifact or asset of this workspace), `start` (seconds, ≥ 0), and **only one** of `end` or `duration`. The API does not fetch from the open internet. First, import the clip (`POST /v1/media-imports`). You must send `Idempotency-Key`. Optional fields: `precision`, `audio`, `output` (exact only), and the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. The default `mode` is **`async`** (`readCommunicationOptions`). To wait a maximum of **30 seconds** for a `200` with the completed job, send `mode: "sync"`. If the job does not complete in that time, you get `202`, and then you poll the job. ```bash curl -X POST https://api.sume.com/v1/video-trim \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-trim-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 2, "duration": 8 }' ``` If the submit is successful, the API returns a job (`request_id` is the job ID). When the status is `result_ready`, `GET /v1/jobs/:id/result` returns `kind: video_trim`. The result contains `video_url` (a new `artf_`, never the source), `duration_seconds`, `actual_start_seconds`, `precision`, `audio`, `output`, and optional `warnings[]`. Put that MP4 into `timeline_create` `video[]` with `source_in` 0. Public rate: **$0.02 per job** (`VIDEO_TRIM_PUBLIC_PRICING`). For the live price, refer to `GET /v1/catalog`. There is no provider inference. Only the worker ffmpeg runs. #### Program | Field | Effect | |---|---| | `start` | In-point, in seconds from the source start. You must send it. | | `end` | Out-point, in seconds. Send only one of `end` / `duration`. If the value is after the end of the source, it **clamps**, and the result gives the warning `trim_clamped_to_source`. | | `duration` | Length of the cut, in seconds. `0.2`–`900`. Send only one of `end` / `duration`. | | `precision` | `exact` (default): frame-accurate re-encode (`libx264`, `yuv420p`). `keyframe`: stream copy. The cut can start a GOP early. Re-base your times against `actual_start_seconds`. | | `audio` | `keep` (default) or `drop`. Exact precision remuxes the kept audio as AAC. | | `output` | Optional `{ width, height, fps }` conform for the output. **exact only.** Width/height 256–2160. `fps` `24` \| `25` \| `30` \| `60`. If you do not send it, the output keeps the source values. | Limits (from `packages/api-contract/src/index.ts`): source ≤ **1800** s, output ≤ **900** s, and output ≥ **0.2** s. #### Refusals (stable codes) | Code | When | |---|---| | `video_trim_range_required` | The request has no `end` and no `duration`. | | `video_trim_range_conflict` | The request has `end` and `duration` together. | | `video_trim_range_empty` | `end` ≤ `start`, or the range is longer than 900 s. | | `video_trim_output_requires_exact` | The request has `output` with `precision: "keyframe"`. | | `ffmpeg_fields_rejected` | The client sent `vf` / `filter` / `ffmpeg` / `cmd` / `codec` / `crf` / related fields. The server compiles ffmpeg. | | `unsupported_media_source` | `video_url` is not on the Sume media host. | | `source_not_found` | The `media.sume.com` URL is dead or is not from this workspace. | | `unsupported_media_type` | The HEAD result is not a video. | | `source_duration_exceeded` | The source is longer than 1800 s (worker). | The API rejects off-host URLs (`https://example.com/…`) at admit. Import the clip first. #### Not this surface | Need | Use | |---|---| | Probe / stills / optional STT | [Video inspect](/models/video-inspect) | | Audio track as a durable wav / mp3 | [Audio detach](/models/audio-detach) | | Put several clips in a sequence | [Timeline 1.0](/models/timeline) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | The accurate frame at `t`, at the source size | [Video frames](/models/video-frames) | #### Related - [Video inspect](/models/video-inspect) - [Video frames](/models/video-frames) - [Audio detach](/models/audio-detach) - [Video filter](/models/video-filter) - [Timeline 1.0](/models/timeline) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Audio detach Source: https://docs.sume.com/models/audio-detach.md Extract the audio track of one Sume-hosted video into a durable wav or mp3. The video is untouched. > **Current SoT.** Media L2 demux (`sume/audio-detach-1.0`, #5953). Dest > and prod. This is **not** clip inspection. For clip inspection, use > [video inspect](/models/video-inspect). For many ranges, detach **once**. > Then split with [timeline audio](/models/timeline-audio). Audio detach 1.0 takes **one** workspace `media.sume.com` video and returns a **new** audio artifact. The default is sample-exact wav (`pcm_s16le`). `timeline_create` `audio.url`, `POST /v1/timeline-1.0/audio`, and speech-to-text use this format. The video does not change. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `createAudioDetachV1` / `submitSumeAudioDetachJob`). ```text POST /v1/audio-detach POST /v1/models/sume/audio-detach-1.0/runs # same body, no extra `model` field ``` There is **no** `GET /v1/audio-detach/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` Hosted MCP: `audio_detach` (`packages/mcp-server/src/mcp.ts`). Writes need `idempotency_key` (and `mcp:write` under OAuth). Flow: `audio_detach` → `jobs_wait` → `jobs_result`. #### Create a detach Required: `video_url` (this workspace’s `media.sume.com` artifact or asset). The server does not fetch from the open internet. Import the file first (`POST /v1/media-imports`). `Idempotency-Key` is required. Optional: `format`, `range`, `channels`, `sample_rate`, and the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. The default `mode` is **`async`**. To wait for a maximum of **30 seconds** for a `200` completed job, pass `mode: "sync"`. Otherwise, you get `202` and must poll. ```bash curl -X POST https://api.sume.com/v1/audio-detach \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: audio-detach-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4" }' ``` A successful submit returns a job (`request_id` is the job id). When `result_ready` is true, `GET /v1/jobs/:id/result` is `kind: audio_detach` and contains `audio_url` (new `artf_`), `duration_seconds`, `format`, `channels`, `sample_rate` (null when omitted, because the value comes from the source), `source_duration_seconds`, and optional `range` / `warnings[]`. Public rate: **$0.01 per job** (`AUDIO_DETACH_PUBLIC_PRICING`, and confirm the live rate in `GET /v1/catalog`). The job uses only worker ffmpeg, with no provider inference. #### Program | Field | Effect | |---|---| | `format` | `wav` (default, `pcm_s16le` sample-exact) or `mp3` (128 kbps). | | `range` | Optional `{ start, end? }` seconds. Omit it for the whole track. When you omit `end`, the range is open-ended. | | `channels` | `source` (default) or `mono`. | | `sample_rate` | `16000` \| `44100` \| `48000`. Omit it to inherit the source. `16000` + `channels: "mono"` is the STT shape. | Caps (from `packages/api-contract/src/index.ts`): source ≤ **1800** s, output ≤ **900** s. A whole track longer than 900 s needs a `range`. A source with no audio track fails with `detach_source_has_no_audio`. First, use [video inspect](/models/video-inspect) to examine `probe.has_audio` (`frames: false` is sufficient). #### Refusals (stable codes) | Code | When | |---|---| | `audio_detach_range_empty` | `range.end` ≤ `range.start`, or the range is longer than 900 s. | | `detach_source_has_no_audio` | The source has no audio track (worker). | | `detach_start_past_source` | `range.start` is more than the probed duration (worker). | | `ffmpeg_fields_rejected` | The client sent `af` / `filter` / `ffmpeg` / `cmd` / `codec` / similar fields. The server compiles ffmpeg. | | `unsupported_media_source` | `video_url` is not on the Sume media host. | | `source_not_found` | A dead `media.sume.com` URL, or one from a different workspace. | | `unsupported_media_type` | HEAD is not a video. | | `source_duration_exceeded` | The source is longer than 1800 s (worker). | The server rejects off-host URLs (`https://example.com/…`) at admit. Import the file first. #### Not this surface | Need | Use | |---|---| | Probe / stills / optional STT | [Video inspect](/models/video-inspect) | | Exact frame at `t`, source size | [Video frames](/models/video-frames) | | A new MP4 cut | [Video trim](/models/video-trim) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | Many audio ranges from one track | Detach once, then [timeline audio](/models/timeline-audio) `operation: "split"` | | Sequence several clips | [Timeline 1.0](/models/timeline) | #### Related - [Video inspect](/models/video-inspect) - [Video frames](/models/video-frames) - [Video trim](/models/video-trim) - [Video filter](/models/video-filter) - [Timeline 1.0](/models/timeline) - [Timeline audio](/models/timeline-audio) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Video filter Source: https://docs.sume.com/models/video-filter.md Dim, crop, or apply an allowlisted pixel filtergraph to one Sume-hosted clip. Returns a new MP4. Material preparation, not timeline placement. > **Current SoT.** Media L2 pixel pass (`sume/video-filter-1.0`, #5820 / > #5839). It is available on dest and prod. This API does **not** cut a > range. For a range cut, use [video trim](/models/video-trim). This API > does **not** assemble clips. The clip sequence, transitions, and the > audio spine stay on [Timeline 1.0](/models/timeline). Video filter 1.0 uses **one** workspace `media.sume.com` clip and a validated program (`ops[]` dim/crop and/or a filters-only `filtergraph`) as input. It returns a **new** MP4. The source does not change. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `createVideoFilterV1` / `submitSumeVideoFilterJob`, and `packages/timeline-compiler/src/filter.ts` `compileVideoFilterProgram`). ```text POST /v1/video-filter POST /v1/video-filter/check # unbilled contract check POST /v1/models/sume/video-filter-1.0/runs # same body, no extra `model` field ``` There is **no** `GET /v1/video-filter/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` Hosted MCP: `video_filter` (`packages/mcp-server/src/mcp.ts`). A write must have `idempotency_key` (and `mcp:write` under OAuth). Flow: `video_filter` with `check_only: true` (free) → `video_filter` → `jobs_wait` → `jobs_result`. #### Check a program (unbilled) `POST /v1/video-filter/check` runs the same checks as the encode: the schema, the op whitelist, the filtergraph allowlist, and the Sume-host / HEAD source preflight. It returns diagnostics and not a `400`. It does **not** create a job, reserve credits, boot a box, or use the encoder. A program that passes this check can still fail on the box (bad expression, memory, time). In that case, you get a structured job error. `Idempotency-Key` is **not** necessary for the check. A valid response is `object: video_filter_check`. It contains `valid`, `encode: "not_run"`, `diagnostics[]`, and the compiled `program.filters` (names only, no argv). If the program is valid, it also contains an `estimate`. It also contains `next_action` `submit_video_filter` | `fix_program_and_recheck`. ```bash curl -X POST https://api.sume.com/v1/video-filter/check \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "ops": [{ "op": "dim", "amount": 0.45 }] }' ``` #### Encode You must send `video_url` (a `media.sume.com` artifact or asset of this workspace) **and** a program. The program must contain a minimum of one of these: `ops[]` or a non-empty `filtergraph`. The API does not fetch from the open internet. First, import the clip (`POST /v1/media-imports`). You must send `Idempotency-Key`. Optional fields: `ops`, `filtergraph`, `metadata`, and the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. The default `mode` is **`async`** (`readCommunicationOptions`). To wait a maximum of **30 seconds** for a `200` with the completed job, send `mode: "sync"`. If the job does not complete in that time, you get `202`, and then you poll the job. ```bash curl -X POST https://api.sume.com/v1/video-filter \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-filter-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "ops": [{ "op": "dim", "amount": 0.45 }] }' ``` If the submit is successful, the API returns a job (`request_id` is the job ID). When the status is `result_ready`, `GET /v1/jobs/:id/result` returns `kind: video_filter`. The result contains `video_url` (a new `artf_`, never the source), `duration_seconds`, `ops_applied`, `filtergraph`, the compiled `filters[]`, and optional `warnings[]`. Put that MP4 into `timeline_create` `video[]`. Public rate: **$0.02 per encode job** (`VIDEO_FILTER_PUBLIC_PRICING`). For the live price, refer to `GET /v1/catalog`. The check is free. There is no provider inference. Only the worker ffmpeg runs. #### Program | Field | Effect | |---|---| | `ops[]` | Ordered program. Sume applies it **before** `filtergraph`. Maximum **8**. You must send a minimum of one op **or** a filtergraph. | | `ops[].op: "dim"` | Multiplies the luma of the full clip. `amount` is in **(0, 1]**. `0.45` makes the clip less bright, and `1` makes no change. The API refuses `0` and `>1`. Black stays black, and the chroma does not change. | | `ops[].op: "crop"` | Rectangle as **fractions** of the source frame. `x` and `y` are in `[0, 1]`. `width` and `height` are in `[0.05, 1]`. `x+width ≤ 1` and `y+height ≤ 1`. The compiler rounds to even values for yuv420p. | | `filtergraph` | Filters-only ffmpeg graph. Sume applies it after `ops[]`. Maximum **2048** characters and **32** named filters. No inputs, no outputs, and no paths. The server adds `[0:v]…[vout]` around the graph. Internal labels (`split[a][b]`) are permitted. Stream specifiers (`[0:v]`) are not permitted. | Limits (from `packages/api-contract/src/index.ts`): source ≤ **300** s (compose-band clip ceiling). The output keeps the geometry, frame rate, and audio of the source, unless the program changes them. The allowlisted filter names are in `packages/timeline-compiler/src/filter.ts` `VIDEO_FILTER_GRAPH_ALLOWED_FILTERS` (tone, blur, geometry, fade, internal compositing). These items are **not** on that list: `trim` / `setpts` (use [video trim](/models/video-trim)), `drawtext` / `subtitles` / `movie` / `lut3d`, and all filters that read a file or a socket. #### Refusals (stable codes) | Code | When | |---|---| | `video_filter_ops_empty` | The request has no `ops[]`, and `filtergraph` is missing or empty. | | `video_filter_too_many_ops` | More than 8 ops. | | `unsupported_filter_op` | `ops[].op` is not `dim` or `crop`. Put other pixel work in `filtergraph`. | | `unsupported_filter_op_field` | An op has an extra key (dim accepts only `{op, amount}`, and crop accepts only `{op, x, y, width, height}`). | | `video_filter_amount_out_of_range` | The dim `amount` is not in `(0, 1]`. | | `video_filter_crop_out_of_bounds` | The crop fractions are outside the source frame, or a side is less than `0.05`. | | `invalid_filtergraph` | The graph is empty or too long, or it has an unknown filter, an illegal alphabet, or a stream specifier. `unknown_filter` gives the token and the allowlist. | | `ffmpeg_fields_rejected` | The client sent `vf` / `filter` / `ffmpeg` / `cmd` / `codec` / `crf` / `-i` / related fields. The server compiles ffmpeg. | | `unsupported_media_source` | `video_url` is not on the Sume media host. | | `source_not_found` | The `media.sume.com` URL is dead or is not from this workspace. | | `unsupported_media_type` | The HEAD result is not a video. | | `source_too_large` | The source is larger than the timeline download budget. | | `output_duration_exceeded` | The source is longer than 300 s (worker). | The API rejects off-host URLs (`https://example.com/…`) at admit. Import the clip first. #### Not this surface | Need | Use | |---|---| | Probe / stills / optional STT | [Video inspect](/models/video-inspect) | | A new MP4 cut | [Video trim](/models/video-trim) | | Audio track as a durable wav / mp3 | [Audio detach](/models/audio-detach) | | Put several clips in a sequence | [Timeline 1.0](/models/timeline) | | Caption plates | HyperFrames compose / caption assembler | | The accurate frame at `t`, at the source size | [Video frames](/models/video-frames) | #### Related - [Video inspect](/models/video-inspect) - [Video frames](/models/video-frames) - [Video trim](/models/video-trim) - [Audio detach](/models/audio-detach) - [Timeline 1.0](/models/timeline) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Timeline 1.0 Source: https://docs.sume.com/models/timeline.md Assemble one audio spine plus ordered video slots into a long-form MP4. The only public assembly surface. > **Current SoT.** Timeline 1.0 render (`sume/timeline-1.0`). It is > available on dest and prod. This job is **assembly**: the sequence, > transitions, and the audio spine. Material preparation stays on [video trim](/models/video-trim), > [audio detach](/models/audio-detach), [video filter](/models/video-filter), > and [timeline compose](/models/timeline-compose). Timeline 1.0 takes a declarative document (one audio spine + ordered `video[]` slots) and returns **one** MP4. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `renderTimelineV1` / `submitSumeTimelineRenderJob`). Callers never send filtergraphs, codecs, or shell fragments. ```text POST /v1/timeline-1.0/render POST /v1/timeline-1.0/plan # unbilled compile preflight POST /v1/models/sume/timeline-1.0/runs # same body as render, no extra `model` field POST /v1/models/sume/timeline/v1.0/runs # legacy model-run alias ``` This surface has **no** `GET /v1/timeline-1.0/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` The hosted MCP flow is `timeline_create`, then `jobs_wait`, then `timeline_get` (`packages/mcp-server/src/mcp.ts`). `timeline_get` is `GET /v1/jobs/:id/result`. Writes need `idempotency_key` (and `mcp:write` under OAuth). #### Plan (unbilled) `POST /v1/timeline-1.0/plan` (`planTimelineV1`) runs schema + Sume-host URL checks + the pure compiler. It returns `object: timeline_plan` with `duration_seconds`, `segment_count`, `billable_minutes`, `estimated_cost_usd_micros`, and a `filtergraph_summary`. It does **not** create a job, reserve credits, or download media. `Idempotency-Key` is not required. A plan cannot predict short-source pad/loop warnings. #### Render The required fields are `audio.duration_seconds` (1–**1800**), plus either `audio.url` or `audio.parts[]` (unless `audio.mode` is `"silence"`), and `video[]` (1–**200** slots). Every URL must already be this workspace’s `media.sume.com` artifact or asset. Import the files first (`POST /v1/media-imports`). `Idempotency-Key` is required. The default `mode` is **`async`** (`readCommunicationOptions`). To wait up to **30 seconds** for a `200` finished job, pass `mode: "sync"`. Otherwise, you get `202`. Then poll the job. ```bash curl -X POST https://api.sume.com/v1/timeline-1.0/render \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: timeline-001" \ -d '{ "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24 }, "video": [ { "source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8 }, { "source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16, "transition": { "type": "fade", "duration": 0.25 } } ] }' ``` A successful submit returns a job (`type: timeline_render`, `model: sume/timeline-1.0`). When the job is `result_ready`, `GET /v1/jobs/:id/result` returns `kind: timeline_render` with `video_url`, `duration_seconds`, `segment_count`, `billable_minutes`, and optional `warnings[]`. Soft warnings (padded/looped short sources, snapped transitions, ignored still motion) are **not** failures. The public rate is **$0.10 per ceil(output minute)** (`TIMELINE_PUBLIC_PRICING`). Confirm the live rate in `GET /v1/catalog`. The reserve is `ceil(audio.duration_seconds / 60)` minutes. The job uses no provider inference, only worker ffmpeg. The default output is **1080×1920** MP4. If you omit `output.fps`, the job renders at the rate of the sources (the longest video sources decide, stills have no rate, and 30 applies only when no source has a rate). If a rate is different from the rate of a source, the job repeats or drops a frame every few frames. This causes judder on motion. The job reports this as `output_fps_resamples_sources` with the rate that the sources wanted. #### Program | Field | Effect | |---|---| | `audio.duration_seconds` | Output length. Required. 1–1800 s. | | `audio.url` | One Sume-hosted spine. Not permitted with `parts`. | | `audio.parts[]` | ≤20 gapless slices (`url` + optional `source_in` / `duration`). Sample-domain join, no re-TTS. Not permitted with `url`. | | `audio.mode` | `"silence"` — declared length with no spine file. In this mode, do not send `url` / `parts` / `gain_db` / `source_in`. | | `audio.source_in` | In-point into a **single** `url` spine. Not permitted with `parts`. The output length is still `duration_seconds`. | | `audio.gain_db` | −60…12. Not permitted with silence. | | `video[].source_url` | Sume-hosted clip or still. Stills are static holds (the job accepts `motion`, ignores it, and reports `motion_ignored`). | | `video[].start` | On-spine start. `video[0].start` must be **0**. Later starts must increase. Declared starts are authoritative. The compiler compensates for xfade and never pre-shifts. | | `video[].duration` | On-screen length, ≥ 0.2 s. Coverage can stop at most 0.5 s before the end of the spine. | | `video[].source_in` | In-point into the file. | | `video[].fit` | `cover` (default) \| `contain` \| `stretch` \| `blur`. | | `video[].transition` | Use only on slots **after** the first. `type` ∈ `fade` \| `wipeleft` \| `wiperight` \| `slideup` \| `slidedown` \| `dissolve`. Duration ≤ 1 s, ≤ 50% of the shorter neighbor, and at least one output frame. | | `output.width` / `height` | Even integers 256–2160. | | `output.fps` | `24` \| `25` \| `30` \| `60`. Omit it to match the sources. | | `output.fade_in_seconds` / `fade_out_seconds` | 0–5 s. The sum must be ≤ output length. | | `soundtrack` | Optional bed: `url`, `gain_db`, `loop`, `fade_out_seconds` ≤ 10, `duck_db` 0–20 (needs a real spine, not silence). | | `render.strategy` | `auto` (default, chunks past 12 segments) \| `chunked` \| `single` (Sume refuses `single` above 12 slots: `render_strategy_unsafe`). | If you need sliced VO only **inside this render**, use `audio.parts[]`. For a reusable merged file, use [timeline audio](/models/timeline-audio). For two sources on screen at the same time, use [timeline compose](/models/timeline-compose). Then put that MP4 into `video[]`. #### Refusals (stable codes) | Code | When | |---|---| | `audio_url_required` | No `url` / `parts` and not silence. | | `audio_url_and_parts_exclusive` | Both `url` and `parts`. | | `silent_audio_takes_no_url` / `_parts` / `_gain` / `_source_in` | Silence plus a spine field. | | `audio_source_in_requires_single_spine` | `source_in` with `parts`. | | `audio_parts_shorter_than_duration` | The sum of the declared part lengths is less than `duration_seconds`. | | `timeline_must_start_at_zero` | `video[0].start` ≠ 0. | | `transition_on_first_segment` | `video[0].transition`. | | `invalid_segment_timing` / `segment_overlap` | Starts do not increase, or slots overlap past the xfade. | | `transition_too_long` / `transition_not_frame_aligned` | Duration compared with neighbors / fps. | | `too_many_chained_transitions` | More than 8 adjacent fades. Insert a hard cut. | | `edge_fades_exceed_output` / `soundtrack_fade_exceeds_output` | Fade longer than the spine. | | `duck_requires_audio_spine` | `soundtrack.duck_db` with silence. | | `render_strategy_unsafe` | `strategy: "single"` with more than 12 slots. | | `unsupported_media_source` / `source_not_found` | Off-host or dead URL. | | Provider / ffmpeg keys | 400 — `model`, `filtergraph`, `ffmpeg_args`, `codec`, `crf`, and similar keys (`TIMELINE_REJECTED_PROVIDER_KEYS`). | Sume rejects off-host URLs (`https://example.com/…`) at admit. Import the files first. A dim / crop / pixel pass is **not** a render option. Use [video filter](/models/video-filter). #### Not this surface | Need | Use | |---|---| | `[start, end)` of one clip | [Video trim](/models/video-trim) | | Audio track as a durable wav / mp3 | [Audio detach](/models/audio-detach) | | Concat / split audio into reusable files | [Timeline audio](/models/timeline-audio) | | Still + video in one frame (반배너) | [Timeline compose](/models/timeline-compose) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | Probe / stills / optional STT | [Video inspect](/models/video-inspect) | #### Related - [Timeline compose](/models/timeline-compose) - [Timeline audio](/models/timeline-audio) - [Video trim](/models/video-trim) - [Audio detach](/models/audio-detach) - [Video filter](/models/video-filter) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Timeline compose Source: https://docs.sume.com/models/timeline-compose.md Put one Sume-hosted still and one Sume-hosted video on screen at the same time. Returns one MP4 shot for Timeline 1.0. > **Current SoT.** Timeline 1.0 compose (`sume/timeline-1.0/compose`, > #3976). It is available on dest and prod. This job builds **one shot**. > Assembly (the sequence, transitions, and the audio spine) stays on > [Timeline 1.0](/models/timeline). Sequential image-then-video is **not** > a compose mode. Adjacent `video[]` slots on the render already do that. Compose takes **one** still + **one** video and returns **one** MP4 with both on screen at the same time. The server compiles ffmpeg on the worker media runtime (`apps/api/src/routes.ts` `createTimelineV1Compose` / `submitSumeTimelineComposeJob`, and `packages/timeline-compiler/src/compose.ts`). Callers never send filtergraphs, codecs, or shell fragments. ```text POST /v1/timeline-1.0/compose ``` This surface has **no** `GET /v1/timeline-1.0/compose/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` The hosted MCP tool is `timeline_compose` (`packages/mcp-server/src/mcp.ts`). Writes need `idempotency_key` (and `mcp:write` under OAuth). The flow is `timeline_compose` → `jobs_wait` → `jobs_result` → put `video_url` into `timeline_create` `video[]`. The HTTP field that selects stack or overlay is **`operation`**, not `mode`. `mode` is the usual `async` / `sync` / `webhook` communication option. #### Create a compose The required fields are `operation` (`stack` \| `overlay`), `image.url`, and `video.url`. Both URLs must already be this workspace’s `media.sume.com` artifact or asset. `image.url` must probe as a still (`compose_image_not_still`). `video.url` must probe as video (`compose_video_not_video`). Import the files first (`POST /v1/media-imports`). `Idempotency-Key` is required. The optional fields are `layout`, `output`, `video.source_in`, and `video.duration`, plus the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. The default `mode` is **`async`**. To wait up to **30 seconds** for a `200` finished job, pass `mode: "sync"`. Otherwise, you get `202`. Then poll the job. The output length **always** comes from the video layer (`video.duration`, or else the rest of the file from `source_in`). The still stays on screen for the whole clip and can never make the clip longer. The ceiling is **300** s (`TIMELINE_COMPOSE_MAX_DURATION_SECONDS`). If the duration goes past the end of the source, the job clamps it (`compose_duration_clamped_to_source`). ```bash curl -X POST https://api.sume.com/v1/timeline-1.0/compose \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: timeline-compose-001" \ -d '{ "operation": "stack", "image": { "url": "https://media.sume.com/artifacts/artf_demo/banner.png" }, "video": { "url": "https://media.sume.com/artifacts/artf_demo/talk.mp4" }, "layout": { "split": "horizontal", "image_region": "top", "ratio": 0.5 }, "output": { "width": 720, "height": 1280, "fps": 25 } }' ``` A successful submit returns a job (`model: sume/timeline-1.0/compose`). When the job is `result_ready`, `GET /v1/jobs/:id/result` returns `kind: timeline_compose` with `video_url` (new `artf_`) and `duration_seconds`. Put that MP4 into [Timeline 1.0](/models/timeline) `video[]`. The public rate is **$0.02 flat per job** (`TIMELINE_COMPOSE_PUBLIC_PRICING`). Confirm the live rate in `GET /v1/catalog`. The rate is flat because `video.duration` can be absent until the worker probes the video. The job uses no provider inference, only worker ffmpeg. The default output is **1080×1920** MP4 at the frame rate of the video layer. This is the only rate that does not repeat or drop a frame. Set `output.width` / `output.height` to the timeline that you assemble into. This prevents a second rescale of the shot. If an explicit `output.fps` is different from the rate of the clip, the job warns `output_fps_resamples_sources`. The output uses the audio of the video. A mute video causes a **warning** (`compose_video_has_no_audio`), not a failure. The clip still renders, and the Timeline 1.0 spine supplies the audio at assemble time. #### Layout `stack` tiles two regions of one frame. The defaults `horizontal` / `top` / `0.5` are 반배너 (half-banner: still on top, video below). `ratio` is the share of the still (0.1–0.9). The video takes the exact remainder. | `layout` key | `stack` | `overlay` | |---|---|---| | `split` | `horizontal` \| `vertical` | not permitted | | `image_region` | `top` \| `bottom` on a horizontal split. `left` \| `right` on a vertical split | not permitted | | `ratio` | still’s share of the frame | not permitted | | `image_fit` / `video_fit` | `cover` \| `contain` \| `stretch` (`blur` is not a compose fit) | `video_fit` only | | `position` | not permitted | `top` \| `center` \| `bottom` | | `width_ratio` | not permitted | 0.05–1 of width (default 0.9). The plate keeps its aspect | | `margin_ratio` | not permitted | 0–0.45 of height (default 0.05) | If you mix stack keys with overlay keys, the request returns 400 (`compose_stack_takes_no_overlay_layout` / `compose_overlay_takes_no_stack_layout`). If the region does not match the split, the request returns `compose_image_region_wrong_axis`. #### Refusals (stable codes) | Code | When | |---|---| | `compose_image_not_still` | `image.url` is not a still. | | `compose_video_not_video` | `video.url` is not a video. | | `compose_image_region_wrong_axis` | `left`/`right` on a horizontal split, or `top`/`bottom` on a vertical split. | | `compose_stack_takes_no_overlay_layout` / `compose_overlay_takes_no_stack_layout` | Mixed layout vocabulary. | | `compose_duration_clamped_to_source` | Warning: `video.duration` went past the end of the file. The job still succeeds. | | `compose_video_has_no_audio` | Warning: the video source is mute. The job still succeeds. | | `unsupported_media_source` / `source_not_found` | Off-host or dead URL. | | Provider / ffmpeg keys | 400 — `filtergraph`, `ffmpeg_args`, `codec`, `crf`, and similar keys. | Sume rejects off-host URLs (`https://example.com/…`) at admit. Import the files first. #### Not this surface | Need | Use | |---|---| | Sequence several clips | [Timeline 1.0](/models/timeline) | | Concat / split audio into reusable files | [Timeline audio](/models/timeline-audio) | | `[start, end)` of one clip | [Video trim](/models/video-trim) | | Pixel pass (dim / crop) | [Video filter](/models/video-filter) | | Audio track as a durable wav / mp3 | [Audio detach](/models/audio-detach) | #### Related - [Timeline 1.0](/models/timeline) - [Timeline audio](/models/timeline-audio) - [Video trim](/models/video-trim) - [Video filter](/models/video-filter) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Timeline audio Source: https://docs.sume.com/models/timeline-audio.md Concat or split Sume-hosted audio into durable media.sume.com files. Sample-domain join, no re-synthesis. > **Current SoT.** Timeline 1.0 audio (`sume/timeline-1.0/audio`, #3516). > It is available on dest and prod. This job mints a **reusable** audio file. > If you need a join only inside one render, use > [Timeline 1.0](/models/timeline) `audio.parts[]`, and do not use this job. Timeline audio joins Sume-hosted audio into one gapless file (`operation: concat`) or slices one file into ranges (`operation: split`). It returns durable `media.sume.com` URLs plus timings. The join is sample-domain, with no re-TTS and no silence at the seams. The server compiles ffmpeg on the same worker media runtime as the render (`apps/api/src/routes.ts` `createTimelineV1Audio` / `submitSumeTimelineAudioJob`). ```text POST /v1/timeline-1.0/audio ``` This surface has **no** `GET /v1/timeline-1.0/audio/:id`. Poll the job envelope: ```text GET /v1/jobs/:id/status GET /v1/jobs/:id/result ``` The hosted MCP tool is `timeline_audio` (`packages/mcp-server/src/mcp.ts`). Writes need `idempotency_key` (and `mcp:write` under OAuth). The flow is `timeline_audio` → `jobs_wait` → `jobs_result`. #### Concat The required fields are `operation: "concat"` and `parts[]` (1–**20**, ordered). Each part is `{ url, source_in?, duration? }`. Do **not** send `url` or `ranges` at the top level (`audio_concat_takes_no_url` / `audio_concat_takes_no_ranges`). All URLs must already be this workspace’s `media.sume.com` audio. Import the files first (`POST /v1/media-imports`). `Idempotency-Key` is required. The default `mode` is **`async`**. To wait up to **30 seconds** for a `200` finished job, pass `mode: "sync"`. Otherwise, you get `202`. Then poll the job. ```bash curl -X POST https://api.sume.com/v1/timeline-1.0/audio \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: timeline-audio-concat-001" \ -d '{ "operation": "concat", "parts": [ { "url": "https://media.sume.com/artifacts/artf_demo/line1.wav" }, { "url": "https://media.sume.com/artifacts/artf_demo/line2.wav", "source_in": 0.1, "duration": 1.8 } ] }' ``` The result is `kind: timeline_audio` with one `audio_url`, `duration_seconds`, and `segments[]` (`index`, `start`, `duration_seconds`). The segments are the concat offsets that you use to re-base Timeline 1.0 `video[].start`. Use that file as `audio.url` on the render, or as Avatar 1.0 image-to-video audio. All parts must have the same channel layout (`audio_parts_channel_mismatch`). #### Split The required fields are `operation: "split"`, a top-level `url`, and `ranges[]` (1–**20**). Each range is `{ start, end? }` (if you omit `end`, the range goes to the end of the file). Do **not** send `parts` (`audio_split_takes_no_parts`). Ranges can overlap. ```bash curl -X POST https://api.sume.com/v1/timeline-1.0/audio \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: timeline-audio-split-001" \ -d '{ "operation": "split", "url": "https://media.sume.com/artifacts/artf_demo/spine.wav", "ranges": [{ "start": 0, "end": 12.4 }, { "start": 12.4 }] }' ``` The result is `kind: timeline_audio` with `segments[]`, and each segment has its own `audio_url`. For many ranges from a talking-head MP4, use [audio detach](/models/audio-detach) one time. Then split the audio here. #### Output format The optional `output.format` is **`wav`** (default, `pcm_s16le`, sample-exact) or **`mp3`** (smaller, adds priming padding again at every edge). Keep wav if you will join the file again or if the file drives lip-sync. The produced audio is ≤ **1800** s. The public rate is **$0.01 flat per job** (`TIMELINE_AUDIO_PUBLIC_PRICING`). Confirm the live rate in `GET /v1/catalog`. The job uses no provider inference, only worker ffmpeg. #### Refusals (stable codes) | Code | When | |---|---| | `audio_concat_requires_parts` | Concat without `parts`. | | `audio_concat_takes_no_url` / `audio_concat_takes_no_ranges` | Concat plus a split field. | | `audio_split_requires_url` / `audio_split_requires_ranges` | Split without `url` or `ranges`. | | `audio_split_takes_no_parts` | Split plus `parts`. | | `audio_range_end_before_start` | A range `end` ≤ `start`. | | `audio_parts_channel_mismatch` | Concat parts do not share a channel layout (worker). | | `unsupported_media_source` / `source_not_found` | Off-host or dead URL. | | Provider / ffmpeg keys | 400 — `filtergraph`, `ffmpeg_args`, `codec`, `crf`, and similar keys. | Sume rejects off-host URLs (`https://example.com/…`) at admit. Import the files first. #### Not this surface | Need | Use | |---|---| | Audio track of one video as wav / mp3 | [Audio detach](/models/audio-detach) | | Join only for one render | [Timeline 1.0](/models/timeline) `audio.parts[]` | | Sequence several clips | [Timeline 1.0](/models/timeline) | | Still + video in one frame | [Timeline compose](/models/timeline-compose) | | Speech-to-text | `POST /v1/stt-1.0/transcribe` | #### Related - [Timeline 1.0](/models/timeline) - [Timeline compose](/models/timeline-compose) - [Audio detach](/models/audio-detach) - [Media inputs](/workflows/asset-library) - [Jobs and results](/workflows/jobs-and-results) - [MCP tools and gates](/mcp/tools-and-gates) ### Video analyses Source: https://docs.sume.com/models/video-analyses.md Legacy scene-by-scene analysis resource. Dest create is retired (410); production still accepts until > **Current SoT.** Do not start new work on this surface. > > - **Default on dest and prod:** [Video inspect](/models/video-inspect) — > `POST /v1/video-inspect` (MCP `video_inspect`) for probe + stills + > optional transcript. Probe and stills are free. The default is `mode: sync`. > - **Dest (`api.dev.sume.com`):** `SUME_COM_VIDEO_ANALYSIS_ENABLED=false`. > `POST /v1/video-analyses` answers `410 video_analysis_retired`. The MCP > tools `video-analyses_*` are not in `tools_list`. Stored `vana_` rows stay > readable through GET. For semantic questions / intervals on dest, use > `video_analyze` / `video_segment` only when those names appear in > `tools_list` (trusted `api.dev.sume.com` origin + TwelveLabs key, never > production). > - **Prod (`api.sume.com`):** create stays on until #5953 PR-C2. Dest `410` > is not a production outage. This page documents the legacy `video_analysis` / `vana_` resource (typed `scenes[]`). This resource is not trend discovery or virality prediction, and it does **not** generate video. A reusable Format from clip structure is a remix job, not this analysis. When create is still enabled, a video-understanding model (TwelveLabs Pegasus) watches the MP4 and segments it into typed scenes. Then ffmpeg extracts at least one JPEG still per second of each scene (`keyframes`), plus a representative `keyframe_url`. The current `analysis_version` is `1.1`. ```text POST /v1/video-analyses # dest: 410; prod until PR-C2 GET /v1/video-analyses/:id # stored rows, both environments GET /v1/video-analyses # stored rows, both environments ``` The remote MCP wrappers `video-analyses_create` / `_get` / `_list` obey the same flag. Production lists them until PR-C2, and dest does not list them. Poll the remaining jobs with `jobs_wait` (usual runtime **3–5 minutes**, poll every **30–60s**). **Accuracy caveat:** longer videos give less accurate results. Short clips give the most reliable results. #### Create an analysis job (production until PR-C2) Dest callers: do not send a POST to this endpoint. Use [video inspect](/models/video-inspect), or dest-only `video_analyze` / `video_segment` when Sume lists them. The required field is `video_url`. The optional fields are `max_scenes` (2–40, default 24), `include_transcript` (when `true`, each scene's `audio` carries the spoken `speech` for that scene, and if not, `audio` is `null`), plus the usual `mode` / `webhook_url` / `wait_timeout_seconds` communication fields. ```bash curl -X POST https://api.sume.com/v1/video-analyses \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-analysis-001" \ -d '{ "video_url": "https://media.sume.com/artifacts/example/ad.mp4", "max_scenes": 24 }' ``` A successful submit returns `202` with a `vana_…` resource id, a `video_analysis` job (`request_id`), and `status_url` / `result_url`. If possible, use a durable `media.sume.com` URL from media imports (`POST /v1/media-imports`) or a completed Sume generation artifact. This surface has no Higgsfield-style `media_id` intake gate. #### Pricing note Each accepted analysis reserves and captures **$0.30 USD** of Sume usage under the current fixed estimate. Confirm the live price in `GET /v1/catalog` and OpenAPI. #### Duration limits | Cap | Behavior | |---|---| | Soft (~90s) | The job succeeds. The response can include the warning `low_confidence_long_video`. | | Hard (300s) | Sume rejects the job after the probe with a stable duration error. | #### Unsupported inputs | Input | Result | |---|---| | YouTube URLs | `422 unsupported_source` | | Non-HTTPS / private / localhost | Sume rejects them at admit | | Non-video assets | They fail with a stable, agent-legible code | #### Poll and read ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/video-analyses/vana_123 \ -H "Authorization: Bearer $SUME_API_KEY" curl "https://api.sume.com/v1/video-analyses?limit=20" \ -H "Authorization: Bearer $SUME_API_KEY" ``` When the resource is ready, it includes whole-video metadata plus a typed `scenes[]` array. Each scene has contiguous `start_seconds` / `end_seconds`, `duration_seconds`, `summary`, `visual`, optional `shot_type` / `camera_motion`, `on_screen_text`, and `confidence` (0–1). Each scene also has `audio` (`{ speech, has_speech, has_music }` when the request asked for `include_transcript`, otherwise `null`). Each scene also has `keyframe_url`, the representative JPEG still. This still is the 1s sample nearest 40% of the scene, or the historical 40% pick on sub-1s scenes. The value can also be `null` with a `keyframe_mirror_failed:scene_N` warning. The `keyframes` field is an array of `{ t, url }` stills for every whole second of that scene. A failed second is `url: null` plus `keyframe_mirror_failed:scene_N:t_T`, and it does not drop the scene. The 300s hard duration cap is also the stills ceiling (~300 JPEGs). This surface has no separate 1-second text/context track. #### Scopes | Scope | Operations | |---|---| | `video_analyses:write` | `POST /v1/video-analyses` | | `video_analyses:read` | `GET /v1/video-analyses`, `GET /v1/video-analyses/:id` | #### Related - [Video inspect](/models/video-inspect) (current clip inspection) - [Video captions](/models/video-captions) - [Trending videos](/models/trending-videos) (only discovery metadata, not analysis) - [Jobs and results](/workflows/jobs-and-results) - [Media inputs](/workflows/asset-library) ### Trending videos Source: https://docs.sume.com/models/trending-videos.md Browse or search TikTok trending video metadata for brand, product, creator, or keyword research. Trending video search returns ranked public TikTok video **metadata**. It is a paid research utility, not a generation model. The rebuilt browse feed (no-query default grid, higher limit, quality filters) is **on for DEV** (`api.dev` / Railway `development`, and `www.dev` for the Assets page). Production stays **OFF** until `SUME_COM_TRENDING_VIDEOS_REBUILD_ENABLED` or `NEXT_PUBLIC_SUME_COM_TRENDING_VIDEOS_UI_ENABLED` is exactly `1` or `true`. Production keeps the legacy contract: `query` is required, and the limit is 1–50 (default 10). ```text POST /v1/trending-videos/search ``` #### Search and browse Production (flag off) requires `query`. The optional fields are `platform`, `window`, `limit`, `region`, `summary_mode`, `download`, and `download_limit`. Omit `query` only when the rebuild is on (api.dev, or the explicit flag). ##### Fields | Field | Notes | |---|---| | `platform` | MVP supports only `tiktok` (default). | | `query` | Required in production. Brand, product, creator, or keyword (up to 200 chars). Omit it on api.dev or when `SUME_COM_TRENDING_VIDEOS_REBUILD_ENABLED` is on. | | `window` | `yesterday`, `this-week`, `this-month`, `last-3-months`, `last-6-months`, `all-time`. In production, the default for generic queries is `this-month`. The default for known SaaS profiles is `last-3-months`. | | `limit` | Production: 1–50, default `10`. Rebuild flag on: 1–100, search default `20`, browse default `48`. | | `region` | Optional two-letter country code. A browse without a region sends requests to the US, GB, and KR trending feeds. | | `summary_mode` | `none` (default), `metadata`, or `transcript`. At this time, `transcript` returns metadata plus an unsupported warning. | | `download` / `download_limit` | Reserved for a future mirror workflow. The MVP does not download or mirror videos. Values more than zero return an unsupported warning. | When the rebuild flag is on, search results must match the query (caption, hashtag, mention, or author) and meet a minimum engagement floor. The search does not use off-topic and stale hits to pad the results back to the requested limit. Production (flag off) keeps the legacy generic ranker, with no extra relevance or view floor. #### Response shape The response includes ranked videos with public watch URLs, optional cover thumbnails, author handles, metrics, relevance scores, and optional lightweight summaries. It does not include raw TikTok video CDN URLs. The usual fields of a video entry are: - `url` — canonical public TikTok watch URL - `cover_url` — optional display thumbnail (not a downloadable video) - `description`, `created_at`, `region` - `author.handle` / `author.nickname` - `metrics`, `relevance`, `scores` - `summary` when `summary_mode` is not `none` When the rebuild flag is on, `params.mode` is `browse` if the request omits `query`, and `search` in other cases. `params.cached` is `true` when the in-process feed cache served the response. Production (flag off) omits these fields. The Assets → Trending page (`/trending-videos`) is on for `www.dev` (when `NEXT_PUBLIC_APP_URL` is the DEV host). It stays off on production www until `NEXT_PUBLIC_SUME_COM_TRENDING_VIDEOS_UI_ENABLED` is exactly `1` or `true`. For the exact schema, refer to the live [OpenAPI](https://api.sume.com/reference/json). #### Pricing note Each accepted call reserves and captures **$0.10 USD** of Sume usage. The same per-call price includes `summary_mode: metadata`. The cache can serve repeat browse calls, so that they do not hit ScrapeCreators again. Confirm the live price in `GET /v1/catalog`. #### How this fits workflows Use trending search for research. Then generate with Avatar / other generators, and use your own public HTTPS media inputs. Today, this endpoint does not return downloadable source files for face-swap or captions. #### Related - [Generate avatar video](/models/avatar-videos) - [Face swap (Beta)](/models/face-swap) - [Video captions](/models/video-captions) - [API recipes](/api/cookbook) ### Image 1.0 Source: https://docs.sume.com/models/image.md Retiring soon. Compatibility alias for Image Router Auto. **Sume will retire this model soon.** Image 1.0 is a compatibility alias for [Image Router Auto](/models/images). For new integrations, use `POST /v1/images` with `model: "sume/auto"`. The URLs below continue to accept their legacy request shape, which includes avatar references and transparency. But these URLs use the same Auto model selection and return `job.model: "sume/auto"`. Primary invoke URL: ```text POST /v1/image-1.0/generate ``` Model-run alias (same body): ```text POST /v1/models/sume/image-1.0/runs ``` The public model id is `sume/image-1.0`. #### When to use | Goal | Approach | |---|---| | Text → image | `prompt` only | | Edit / reference | `prompt` + `image_urls` (1–10 public HTTPS URLs) | | Masked edit | add `mask_image_url` with `image_urls` | Use only public HTTPS image URLs. Sume rejects localhost, private-network, and non-HTTPS URLs before submission. #### Request fields | Field | Required | Notes | |---|---|---| | `prompt` | Yes | Non-empty string. | | `image_urls` | No | 1–10 reference/edit image URLs. Use this field, not the deprecated `input_urls`. | | `mask_image_url` | No | Mask image URL for edit flows. | | `aspect_ratio` | No | Per-model native catalog. `4:5` is Instagram portrait (1080×1350), not 4:3. gpt-image-2 also accepts `5:4` / `9:8` / `4:5`. Custom-pixel models ignore this field when `image_size` is set. | | `image_size` | No | Named presets or `{ width, height }` / `WIDTHxHEIGHT` on models that accept custom pixels (GPT, Seedream, Flux, Qwen, Recraft). GPT custom: both edges ×16, max edge 3840, aspect ≤3:1, 655,360–8,294,400. On Nano Banana, WxH maps to the native `aspect_ratio` (1080×1350 → `4:5`). Exact pixels come from a documented post-step through job `target_pixels`. | | `quality` | No | `low` (default), `medium`, `high`. Use a higher value for finals, dense text, or packaging. | | `num_images` | No | Integer 1–4. Use this field, not the deprecated `n`. | | `output_format` | No | `png`, `jpeg`, `jpg`, `webp`. Use this field, not the deprecated `format`. | | `metadata` | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider. | | `mode` | No | `async` (default behavior on most clients when you omit the field), `sync`, `subscribe`, `webhook`. | | `webhook_url` | No | Public HTTPS callback for terminal delivery when you use webhook mode. | | `wait_timeout_seconds` | No | 0–30. Maximum time that `sync` / `subscribe` waits before it returns. | Sume still accepts these deprecated aliases: `input_urls`, `n`, `format`. Use the non-deprecated names above. #### Create an image job Reference / edit example: ```bash curl -X POST https://api.sume.com/v1/image-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: image-edit-001" \ -d '{ "prompt": "Keep the product identical; swap the background to a soft daylight studio", "image_urls": ["https://example.com/product.png"], "quality": "medium", "aspect_ratio": "4:3" }' ``` When you retry after a client timeout, use the same `Idempotency-Key` again only for the same operation and payload. Refer to [Jobs and results](/workflows/jobs-and-results). #### Poll and fetch the result Submit responses include `status_url`, `result_url`, `events_url`, and optional `cancel_url`. Poll until the job is terminal. Then fetch the result. ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` #### Artifacts Completed Image 1.0 jobs return Sume-hosted media under `result.artifacts[]`: ```json { "id": "job_...", "status": "completed", "result": { "artifacts": [ { "id": "artifact_...", "type": "image", "url": "https://media.sume.com/artifacts/...", "content_type": "image/png" } ] } } ``` Use the returned Sume media URLs. Raw provider URLs are not part of the public result contract. #### Next - [Video 1.0](/models/video) for motion from prompts or first frames - [Jobs and results](/workflows/jobs-and-results) for modes, cancellation, and events - [Media inputs](/workflows/asset-library) for HTTPS URL rules - [Recipes](/api/cookbook) for short copy-paste flows ### Image API Source: https://docs.sume.com/models/images.md Generate images from text prompts and reference images across the Sume image model catalog. Sume has a dedicated Image API that generates images from text prompts and optional reference images. The service covers model discovery, per-endpoint capabilities, and generation features. To find the available models and their prices, use `GET /v1/images/models`. The catalog publishes the parameters that each model accepts as capability descriptors. Thus, you can find what a model supports before you call it. [Sume specifics](#sume-specifics) lists the behavior that is specific to Sume. ```text POST /v1/images GET /v1/images/models GET /v1/images/models/{model_id}/endpoints ``` To select a family, use a catalog model ID. To let Image Router choose, use `model: "sume/auto"`. Sume will retire [Image 1.0](/models/image) soon. Its public URLs stay compatibility aliases for this Auto pipe. #### Model discovery ChatGPT Image 2.5 is available as `openai/gpt-image-2.5` (Flare) and `openai/gpt-image-2.5-sunburst`. Both support text-to-image, up to 16 image references, an optional `mask_url`, and `background: auto|transparent|opaque`. Use `quality: auto|low|medium|high|xhigh|max`. If you omit quality, the default is `high`. `image_size` accepts named presets, `auto`, or custom pixels. For custom pixels, both edges must be multiples of 16, and the maximum edge is 3840. The aspect ratio must be at most 3:1, and the image must have 655,360–8,294,400 pixels. You can still select ChatGPT Image 2. Flare and Sunburst use the same [Fal token rates](https://fal.ai/models/openai/gpt-image-2.5/flare/text-to-image): $30 per million output image tokens, $8 per million input image tokens, and $5 per million input text tokens. Output estimates use OpenAI's ChatGPT Image 2.5 size/quality calculator. At 1024×1024, `xhigh` output is $0.09366 and `max` output is $0.21072, before input tokens and Sume pricing. Input token counts are estimates. Fal rounds the total up to $0.0001. `auto` quality reserves `max`. `auto` size and named presets without a verified GPT-specific pixel mapping reserve the upper bound of output tokens. Auto model routing continues to use Flare. Fal does not advertise a separate Effortless endpoint. Its documented automatic quality option is `auto`. Ideogram 4.5 is available as `ideogram/ideogram-v4.5`. Without `input_references`, it generates from text. With references, it edits the first image and uses up to 4 more as references (5 total). Use `quality: low|medium|high` (if you omit it, the default is `medium`) and `resolution: 1K|2K`. The [Fal list price](https://fal.ai/models/ideogram/v4.5) is $0.03, $0.06, or $0.22 per image by quality, for all sizes. An edit without `aspect_ratio` keeps the shape of the source image. ##### Via the Image Models API To list the available models and their capabilities, use the image models endpoint: ```bash curl "https://api.sume.com/v1/images/models" \ -H "Authorization: Bearer $SUME_API_KEY" ``` Key response fields include: - **id**: Model slug for generation requests - **architecture**: Supported input/output modalities - **supported_parameters**: Union of capabilities across endpoints - **supports_streaming**: Shows if native SSE streaming is available - **endpoints**: URL for per-endpoint records ```json { "data": [ { "id": "bytedance-seed/seedream-4.5", "name": "Seedream 4.5", "description": "Text-to-image and reference-guided image editing.", "created": 1748372400, "architecture": { "input_modalities": ["text", "image"], "output_modalities": ["image"] }, "supported_parameters": { "prompt": { "type": "boolean" }, "aspect_ratio": { "type": "enum", "values": ["1:1", "16:9", "9:16", "4:3", "3:4"] }, "n": { "type": "range", "min": 1, "max": 4 }, "input_references": { "type": "range", "min": 0, "max": 10 }, "output_format": { "type": "enum", "values": ["png", "jpeg", "webp"] } }, "supports_streaming": false, "endpoints": "/v1/images/models/bytedance-seed/seedream-4.5/endpoints" } ] } ``` ##### Per-endpoint records To get the definitive capabilities and prices for a model, use this call: ```bash curl "https://api.sume.com/v1/images/models/bytedance-seed/seedream-4.5/endpoints" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The important fields are: - **provider_slug**: Use this value for provider-specific parameters - **provider_tag**: Use this value to pin requests to specific providers - **supported_parameters**: Definitive parameter set for this endpoint - **allowed_passthrough_parameters**: Provider-specific keys - **pricing**: Billable lines with cost information ```json { "id": "bytedance-seed/seedream-4.5", "endpoints": [ { "provider_name": "Sume", "provider_slug": "sume", "provider_tag": "sume", "supported_parameters": { "prompt": { "type": "boolean" }, "aspect_ratio": { "type": "enum", "values": ["1:1", "16:9", "9:16", "4:3", "3:4"] }, "n": { "type": "range", "min": 1, "max": 4 }, "input_references": { "type": "range", "min": 0, "max": 10 }, "output_format": { "type": "enum", "values": ["png", "jpeg", "webp"] } }, "allowed_passthrough_parameters": [], "supports_streaming": false, "pricing": [{ "billable": "output_image", "unit": "image", "cost_usd": 0.033 }] } ] } ``` In v1, Sume serves every catalog model through a single `sume` endpoint. Thus, the model-level and endpoint-level `supported_parameters` are identical. ##### Capability descriptors Parameters use typed descriptors: - **enum**: Discrete allowlist of string values - **range**: Any integer in min/max bounds - **boolean**: Supported (present) or unsupported (absent) If a request sets a parameter that the selected model does not list, Sume rejects the request with `400 unsupported_parameter`. Sume does not silently drop the parameter. #### API usage Send a POST request to `/v1/images` with a model and a prompt: **Python (requests):** ```python import requests url = "https://api.sume.com/v1/images" headers = { "Authorization": f"Bearer {SUME_API_KEY}", "Content-Type": "application/json" } payload = { "model": "bytedance-seed/seedream-4.5", "prompt": "a red panda astronaut floating in space, studio lighting" } response = requests.post(url, headers=headers, json=payload) result = response.json() for image in result["data"]: print(f"Generated image: {image['url']}") ``` **TypeScript (fetch):** ```typescript const response = await fetch('https://api.sume.com/v1/images', { method: 'POST', headers: { Authorization: `Bearer ${SUME_API_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: 'bytedance-seed/seedream-4.5', prompt: 'a red panda astronaut floating in space, studio lighting', }), }); const result = await response.json(); for (const image of result.data) { console.log(`Generated image: ${image.url}`); } ``` **cURL:** ```bash curl -X POST "https://api.sume.com/v1/images" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "bytedance-seed/seedream-4.5", "prompt": "a red panda astronaut floating in space, studio lighting" }' ``` ##### Response format The response gives the images as Sume-hosted URLs with usage data: ```json { "created": 1748372400, "model": "bytedance-seed/seedream-4.5", "data": [ { "url": "https://media.sume.com/img/01J.../0.png", "media_type": "image/png" } ], "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0, "cost": 0.04 } } ``` For non-PNG formats, the response is: ```json { "created": 1748372400, "model": "bytedance-seed/seedream-4.5", "data": [ { "url": "https://media.sume.com/img/01J.../0.webp", "media_type": "image/webp" } ], "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0, "cost": 0.04 } } ``` `model` echoes the id that you requested. For example, `sume/auto` stays `sume/auto`. `cost` is the USD amount that Sume bills to your wallet. In v1, token counts are always `0`. Sume meters image models per image, and per-token usage data is not available yet. ##### Long-running requests `POST /v1/images` blocks for up to 30 seconds and returns the response above with `200`. Most catalog models complete in that budget. If the generation does not complete before the budget expires, Sume returns `202` with the standard job envelope instead. Sume also returns this envelope if you send `mode: "async"`, or `mode: "webhook"` with a `webhook_url`: ```json { "data": { "job": { "id": "job_01J...", "model": "bytedance-seed/seedream-4.5", "status": "queued" }, "status_url": "https://api.sume.com/v1/jobs/job_01J.../status", "result_url": "https://api.sume.com/v1/jobs/job_01J.../result" } } ``` Poll `GET /v1/jobs/{id}/status`. Then fetch `GET /v1/jobs/{id}/result` to get the generated images. These are the standard Sume job endpoints. They return the standard job result shape, not the image body above. Refer to [Jobs and results](/workflows/jobs-and-results). Examine the status code, not the body shape. `200` is the image response, and `202` is the job envelope. Slow configurations (4K, high `quality`, large `n`) are the most likely to degrade to `202`. #### Image configuration options ##### Resolution and aspect ratio ```json { "model": "bytedance-seed/seedream-4.5", "prompt": "a landscape photo", "resolution": "2K", "aspect_ratio": "16:9" } ``` - **resolution**: Normalized tier (512, 1K, 2K, 4K) - **aspect_ratio**: Normalized ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 1:2, 2:1, 1:4, 4:1, 1:8, 8:1, 9:21, 21:9). Use "auto" to let the provider choose - **size**: Shorthand for a resolution tier. Do not put custom pixels on `size`. Use `image_size` / `aspect_ratio`. `4:5` is Instagram portrait (1080×1350), not 4:3. Banana Pro sends `aspect_ratio: "4:5"` (native ~928×1152 at 1K). Exact 1080×1350 comes from a documented post-step through job `target_pixels`. A model accepts only the values that its catalog descriptors list. Thus, read `supported_parameters` before you pin a tier or ratio. On edit and image-to-image calls, use `aspect_ratio: "auto"` to match the reference. If you omit the field, the result is not the same as `auto`. ##### Quality and output format ```json { "model": "openai/gpt-image-2", "prompt": "a product photo", "quality": "high", "output_format": "png" } ``` - **quality**: auto, low, medium, high, xhigh, or max (catalog-gated) - **output_format**: png, jpeg, webp, or svg - **background**: auto, transparent, or opaque (ChatGPT Image 2.5 supports this field) - **output_compression**: 0–100 for webp/jpeg (Sume does not serve this field in v1) `output_compression` and `seed` are part of the schema, but no model advertises them yet. Thus, a request that sends one of them returns `400 unsupported_parameter`. For transparent stills today, use [Image 1.0](/models/image) with `transparency: true`. ##### Multiple images ```json { "model": "openai/gpt-image-2", "prompt": "a cute cat", "n": 4 } ``` Use `n` to request up to 10 images per call. Per-model ceilings are lower. Read the `n` range descriptor from the catalog. ##### Image-to-image (reference images) ```json { "model": "openai/gpt-image-2", "prompt": "make this scene look like a watercolor painting", "input_references": [ { "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } } ] } ``` Reference URLs must be public HTTPS. Sume rejects localhost, private-network, and non-HTTPS URLs before submission. If the `input_references` descriptor of a model is `{"min": 0, "max": 0}`, the model is text-to-image only and rejects references. ##### Provider routing ```json { "model": "bytedance-seed/seedream-4.5", "prompt": "a red panda astronaut floating in space", "provider": { "only": ["sume"], "allow_fallbacks": false } } ``` The routing fields are: - **only**: Allow only the listed provider slugs - **order**: Try providers in the listed order - **ignore**: Do not use the listed provider slugs - **sort**: Sort by price, throughput, or latency - **allow_fallbacks**: If false, stop after the primary provider In v1, Sume publishes a single `sume` endpoint for each model. Thus, `only` and `order` accept only `"sume"`. Sume accepts `ignore`, `sort`, and `allow_fallbacks`, but they have no effect. Any other slug returns `400 provider_not_available`. ##### Provider-specific options ```json { "model": "black-forest-labs/flux.2-pro", "prompt": "a dramatic portrait", "provider": { "options": { "sume": {} } } } ``` In v1, `allowed_passthrough_parameters` is empty for every endpoint. Thus, you must omit `provider.options` or send it empty. #### Streaming image generation In v1, Sume does not serve native SSE streaming. Every catalog row reports `supports_streaming: false`, and `stream: true` returns `400 streaming_not_supported`. The field is in the schema so that clients can use streams without a code change when Sume ships the feature. Until then, submit with `mode: "async"`. Then read `GET /v1/jobs/:id/events` for progress. As an alternative, use a [webhook](/workflows/webhooks) for the terminal event. `mode: "subscribe"` is **not** a progress stream. It is an alias of `sync` and gives you one bounded 30-second wait. Refer to [what "subscribe" means](/workflows/jobs-and-results#subscribe-means-three-different-things). #### Billing and cancellation Sume bills image generation on an all-or-nothing basis. Either a generation completes and Sume bills it in full, or the generation fails and Sume does not bill it. - **Completed generations** get the full charge from the endpoint pricing - **Failed or cancelled generations** get no charge. Failed requests return 502 Bad Gateway - **Client disconnects**: Sume bills a request that ends early as a failed generation (no charge) Endpoint `pricing` lines are the amount that Sume charges to your wallet. These lines already include the Sume margin. Thus, you pay `cost_usd × n`. #### Request parameters | Parameter | Type | Required | Description | |-----------|------|----------|-------------| | `model` | string | Yes | Model slug (for example, bytedance-seed/seedream-4.5), or `sume/auto` | | `prompt` | string | Yes | Text that describes the image | | `n` | integer | No | Number of images to generate (1–10) | | `resolution` | string | No | Resolution tier (512, 1K, 2K, 4K) | | `aspect_ratio` | string | No | Aspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 1:4, 4:1, etc.) | | `size` | string | No | Shorthand for a resolution tier. Sume does not serve explicit pixels in v1 | | `quality` | string | No | auto, low, medium, high, xhigh, or max (catalog-gated) | | `output_format` | string | No | png, jpeg, webp, or svg | | `mask_url` | string | No | Optional public HTTPS mask URL for ChatGPT Image 2.5 edits. | | `background` | string | No | auto, transparent, or opaque (ChatGPT Image 2.5) | | `output_compression` | integer | No | Compression level (0–100) for webp/jpeg (not served in v1) | | `seed` | integer | No | Seed for deterministic generation (not served in v1) | | `stream` | boolean | No | Stream partial images through SSE (not served in v1) | | `input_references` | array | No | Reference images for image-to-image | | `provider.only` | string[] | No | Allow only these provider slugs | | `provider.order` | string[] | No | Try provider slugs in this order | | `provider.ignore` | string[] | No | Exclude these provider slugs | | `provider.sort` | string or object | No | Sort by price, throughput, or latency | | `provider.allow_fallbacks` | boolean | No | Allow fallback provider on failure | | `provider.options` | object | No | Provider-specific parameters by slug | | `metadata` | object | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider | | `mode` | string | No | `sync` (default on this route), `async`, `subscribe`, `webhook` | | `webhook_url` | string | No | Public HTTPS callback for terminal delivery in webhook mode | | `wait_timeout_seconds` | integer | No | 0–30, default 30 on this route. Maximum time that `sync` / `subscribe` waits before it returns | #### Sume specifics | Area | Behavior | |---|---| | `sume/auto` | Sume-only `model` value. Sume selects the family for you and never discloses which one ran. `GET /v1/images/models` does not list it, and `job.model` stays `sume/auto`. | | Result payload | `data[].url` (Sume-hosted, signed), not inline base64. Sume already mirrors generated media, and URLs keep responses small. | | Async | Sume limits the wait of a call to 30s. Generations that exceed this limit, and `mode: "async"` / `"webhook"`, return the Sume job envelope with `202`. You then read the images from the standard job result endpoint. | | Provider | One `sume` endpoint for each model in v1. Sume does not disclose the upstream provider identity. Sume accepts the multi-provider routing fields, but they have no effect. | | `stream` | The schema accepts this field. Until native SSE ships, Sume rejects it at runtime with `400 streaming_not_supported`. | | Catalog-gated parameters | `output_compression`, `seed`, and explicit pixel `size` are in the schema, but no model advertises them in v1. Thus, they return `400 unsupported_parameter`. | | `usage` | `cost` is the billed USD amount. Token counts are `0` in v1. | | Legacy ids | Sume accepts the bare Image Router ids (`gpt-image-2`, `nano-banana-2`, …) as aliases for their `org/slug` equivalents. | The legacy `POST /v1/image-router/generate` and `GET /v1/image-router/models` routes still work without change. But they are deprecated, and this surface replaces them. These routes will not get new parameters. ### Video 1.0 Source: https://docs.sume.com/models/video.md Retiring soon. Compatibility alias for Video Router Auto. **Sume will retire this API soon.** Video 1.0 is a compatibility alias for [Video Router Auto](/models/videos). For new integrations, use `POST /v1/videos` with `model: "sume/auto"`. The URLs below continue to accept their legacy request shape. But they use the same Auto model selection, capability validation, and prices. Job receipts show `sume/auto`. The API accepts retired `routing_preset` values and ignores them. Primary invoke URL: ```text POST /v1/video-1.0/generate ``` Model-run alias (same body): ```text POST /v1/models/sume/video-1.0/runs ``` Public model ID: `sume/video-1.0`. #### When to use | Goal | Approach | |---|---| | Text → video | `prompt` only | | Image → video | `prompt` + `image_url` (first frame) | | Start + end frames | `image_url` + `end_image_url` | | Reference-guided | `reference_image_urls` and/or `reference_video_urls` (optional `reference_audio_urls` with a minimum of one image or video reference) | Use only public HTTPS media URLs. #### Request fields | Field | Required | Notes | |---|---|---| | `prompt` | Yes | Non-empty string. | | `image_url` | No | First-frame image URL. We recommend it over the deprecated `first_frame_url`. | | `end_image_url` | No | End-frame image URL. You must also send `image_url` (or the deprecated `first_frame_url`). We recommend it over the deprecated `last_frame_url`. | | `reference_image_urls` | No | 1–9 image URLs. | | `reference_video_urls` | No | 1–3 video URLs. | | `reference_audio_urls` | No | 1–3 audio URLs. You must also send a minimum of one reference image or video. | | `resolution` | No | Sume validates it against the Auto serving family. From this legacy vocabulary, that family supports `720p` (the default if you do not send it) and `1080p`. This URL rejects `4k`. For 4K, use `POST /v1/videos` with `sume/auto`. | | `duration` | No | Integer seconds. Sume validates it against the Auto serving family. We recommend it over `duration_seconds`. If you set the two fields, they must agree. | | `duration_seconds` | No | Alias of `duration`. | | `bitrate_mode` | No | The legacy shape keeps it, but Auto rejects it. | | `aspect_ratio` | No | `auto`, `adaptive`, `21:9`, `16:9`, `4:3`, `1:1`, `3:4`, `9:16`. | | `generate_audio` | No | Boolean. The current Auto family always generates audio. Do not send this field. | | `routing_preset` | No | Deprecated. Sume ignores these values: `cost`, `speed`, `quality`, `grok`, `kling`. All of them use Auto. | | `metadata` | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider. | | `mode` | No | `async`, `sync`, `subscribe`, `webhook`. | | `webhook_url` | No | Public HTTPS callback for webhook mode. | | `wait_timeout_seconds` | No | 0–30 for `sync` / `subscribe`. | The legacy URL does **not** accept a `model` body field. To select a family, use `POST /v1/videos` with a catalog model ID. #### Create a video job Image-to-video example: ```bash curl -X POST https://api.sume.com/v1/video-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-i2v-001" \ -d '{ "prompt": "Gentle camera drift; keep the product locked in frame", "image_url": "https://example.com/first-frame.png", "resolution": "1080p", "duration": 6, "aspect_ratio": "16:9" }' ``` #### Poll and fetch the result ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/events \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Obey `next_action` from the submit envelope (`poll_status`, then `fetch_result`). For the full lifecycle details, refer to [Jobs and results](/workflows/jobs-and-results). #### Artifacts Completed Video 1.0 jobs return Sume-hosted video artifacts (and in some cases, image artifacts): ```json { "id": "job_...", "status": "completed", "result": { "artifacts": [ { "id": "artifact_...", "type": "video", "url": "https://media.sume.com/artifacts/...", "content_type": "video/mp4" } ] } } ``` Artifact URLs are opaque. Do not parse their paths to find workspace, job, or provider identifiers. #### Next - [Image 1.0](/models/image) for stills / first frames - [Avatar video](/models/avatar-videos) for talking-head scripts on a reusable avatar - [Generation admission](/workflows/generation-admission) for queue and concurrency behavior ### Video generation Source: https://docs.sume.com/models/videos.md How to generate videos with Sume models via the asynchronous /v1/videos API. Sume supports video generation from text prompts (and optional reference images) through a dedicated asynchronous API. To find the supported models, their capabilities, and their prices, use `GET /v1/videos/models`. This surface agrees field-for-field with the [OpenRouter Video Generation API](https://openrouter.ai/docs/guides/overview/multimodal/video-generation). Thus, a client that you wrote from their docs works here after you change the base URL and the API key. The few differences of Sume are in [Sume differences](#sume-differences). #### Model Discovery You can find video generation models in these ways: ##### Via the Video Models API To get a list of all available video generation models and their supported parameters, use the dedicated video models endpoint: ```bash curl "https://api.sume.com/v1/videos/models" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The response returns a `data` array. Each model in the array includes these fields: ```json { "data": [ { "id": "seedance-2", "canonical_slug": "seedance-2", "name": "Seedance 2.0", "description": "Seedance 2.0 — text/image/reference-to-video with optional audio.", "created": 1767225600, "supported_resolutions": ["480p", "720p", "1080p"], "supported_aspect_ratios": ["21:9", "16:9", "4:3", "1:1", "3:4", "9:16"], "supported_sizes": null, "supported_durations": [4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15], "supported_frame_images": ["first_frame", "last_frame"], "supported_input_references": ["image_url", "video_url", "audio_url"], "generate_audio": true, "seed": false, "pricing_skus": { "per-1000-video-tokens": "0.0154" }, "allowed_passthrough_parameters": [], "hugging_face_id": null } ] } ``` | Field | Description | | -------------------------------- | ---------------------------------------------------------------------------- | | `id` | Model slug to use in generation requests | | `canonical_slug` | Permanent model identifier | | `supported_resolutions` | List of supported output resolutions (for example, `720p`, `1080p`) | | `supported_aspect_ratios` | List of supported aspect ratios (for example, `16:9`, `9:16`) | | `supported_sizes` | List of supported pixel dimensions (for example, `1280x720`), or `null` | | `supported_durations` | Supported video lengths in whole seconds | | `supported_frame_images` | The `frame_type` values that the model accepts | | `supported_input_references` | The `input_references` types that the model accepts | | `generate_audio` | Shows if the model can generate an audio track | | `seed` | Shows if the model accepts a `seed` | | `pricing_skus` | Price information for each SKU | | `allowed_passthrough_parameters` | Provider-specific parameters that you can send through the `provider` option | Before you submit a generation request, use this endpoint to find the resolutions, aspect ratios, and durations that each model supports. Limits are different for each model. `seedance-2.5` accepts 4–30 seconds at 480p/720p/1080p. `wan-3.0` accepts 2–30 seconds. `higgsfield-genjutsu` (Motion Transfer: one source video plus 1–8 reference images, 480p/720p, in the catalog only when its provider is configured) accepts 4–30 seconds. H3 Max Recast (`h3-max-recast`: one source video plus 1–4 person photos, 768p/1080p, optional prompt) accepts 5–30 seconds. `minimax-h3` accepts 5–15 seconds at native 480p/768p (768p is first-class, not 720p). For this model, Sume bills 2K/4K upscales if your request includes them. Each other catalog model has a maximum of 15 seconds. `seedance-2` also has 1080p. MiniMax H3 Max (`minimax-h3-max`) is the faster 768p variant. It does text-to-video, first/last-frame image-to-video, and reference-to-video at 480p/768p/1080p (1080p is a latent refinement from native 768p) for 5–15 seconds, with native stereo audio. It reports image, video, and audio `supported_input_references`. Gemini Omni Flash 1.1 (`gemini-omni-flash-1.1`) accepts 3–10 seconds at 360p/720p/1080p/4K in 16:9 or 9:16, with native synced audio. It accepts image and video `input_references` (no audio). It makes its video edit mode available through the Video Router `video_url` field. ##### Via the Models API You can also use the [public model catalog](/models) to find video generation models: ```bash curl "https://api.sume.com/v1/catalog" \ -H "Authorization: Bearer $SUME_API_KEY" ``` ##### On the Models Page Go to the [Models page](/models). Find the models that show `video` as an output modality. #### How It Works Different from text or image generation, video generation is **asynchronous**, because a video takes much more time to generate. The workflow has these steps: 1. **Submit** a generation request to `POST /v1/videos` 2. **Receive** a job ID and a polling URL immediately 3. **Poll** the polling URL (`GET /v1/videos/{jobId}`) until the status is `completed` 4. **Download** the video from the content URL (`GET /v1/videos/{jobId}/content`) #### API Usage ##### Submitting a Video Generation Request ```python title="Python" import requests import time url = "https://api.sume.com/v1/videos" headers = { "Authorization": f"Bearer {SUME_API_KEY}", "Content-Type": "application/json", } payload = { "model": "seedance-2", "prompt": "A golden retriever playing fetch on a sunny beach with waves crashing in the background", } # Step 1: Submit the generation request response = requests.post(url, headers=headers, json=payload) result = response.json() job_id = result["id"] polling_url = result["polling_url"] print(f"Job submitted: {job_id}") print(f"Status: {result['status']}") # Step 2: Poll until completion while True: time.sleep(30) # Wait 30 seconds between polls poll_response = requests.get(polling_url, headers=headers) status = poll_response.json() print(f"Status: {status['status']}") if status["status"] == "completed": # Step 3: Download the video content_url = status["unsigned_urls"][0] video_response = requests.get(content_url) with open("output.mp4", "wb") as f: f.write(video_response.content) print("Video saved to output.mp4") break elif status["status"] == "failed": print(f"Generation failed: {status.get('error', 'Unknown error')}") break ``` ```typescript title="TypeScript (fetch)" const headers = { Authorization: `Bearer ${SUME_API_KEY}`, 'Content-Type': 'application/json', }; // Step 1: Submit the generation request const response = await fetch('https://api.sume.com/v1/videos', { method: 'POST', headers, body: JSON.stringify({ model: 'seedance-2', prompt: 'A golden retriever playing fetch on a sunny beach with waves crashing in the background', }), }); const result = await response.json(); const jobId = result.id; const pollingUrl = result.polling_url; console.log(`Job submitted: ${jobId}`); console.log(`Status: ${result.status}`); // Step 2: Poll until completion while (true) { await new Promise((resolve) => setTimeout(resolve, 30000)); // Wait 30 seconds const pollResponse = await fetch(pollingUrl, { headers }); const status = await pollResponse.json(); console.log(`Status: ${status.status}`); if (status.status === 'completed') { // Step 3: Download the video const contentUrl = status.unsigned_urls[0]; console.log(`Video ready: ${contentUrl}`); break; } else if (status.status === 'failed') { console.error(`Generation failed: ${status.error ?? 'Unknown error'}`); break; } } ``` ```bash title="cURL" # Step 1: Submit the generation request curl -X POST "https://api.sume.com/v1/videos" \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "seedance-2", "prompt": "A golden retriever playing fetch on a sunny beach with waves crashing in the background" }' # Response: # { # "id": "", # "polling_url": "https://api.sume.com/v1/videos/", # "status": "pending", # "model": "seedance-2" # } # Step 2: Poll for status curl "https://api.sume.com/v1/videos/" \ -H "Authorization: Bearer $SUME_API_KEY" # Step 3: Once status is "completed", download from unsigned_urls[0] ``` ##### Request Parameters | Parameter | Type | Required | Description | | ------------------ | ------- | -------- | ----------------------------------------------------------------------------------------------------------------------------- | | `model` | string | Yes | The model for video generation (for example, `seedance-2.5`), or `sume/auto` to let Sume select the model | | `prompt` | string | Yes | Text description of the video to generate | | `duration` | integer | No | Duration of the generated video in seconds | | `resolution` | string | No | Resolution of the output video (for example, `720p`, `1080p`) | | `aspect_ratio` | string | No | Aspect ratio of the output video (for example, `16:9`, `9:16`, `3:2`) | | `size` | string | No | Accurate pixel dimensions in `WIDTHxHEIGHT` format (for example, `1280x720`). An alternative to `resolution` + `aspect_ratio` | | `frame_images` | array | No | Images for first/last frames (image-to-video) | | `input_references` | array | No | Reference images for the style (reference-to-video) | | `generate_audio` | boolean | No | Tells the model to generate audio with the video, or not. The default is the audio capability of the model | | `seed` | integer | No | Seed for deterministic generation (the result is not deterministic on all providers) | | `callback_url` | string | No | URL that receives a webhook notification when the job completes. It must be HTTPS | | `provider` | object | No | Provider-specific passthrough configuration | ##### Supported Resolutions - `480p` - `720p` - `768p` - `1080p` - `1K` - `2K` - `4K` Each model shows the subset that it accepts in `supported_resolutions`. ##### Supported Aspect Ratios - `16:9` — Widescreen landscape - `9:16` — Vertical/portrait - `1:1` — Square - `4:3` — Standard landscape - `3:4` — Standard portrait - `3:2` — Photography landscape - `2:3` — Photography portrait - `21:9` — Ultra-wide - `9:21` — Ultra-tall Each model shows the subset that it accepts in `supported_aspect_ratios`. ##### Using Images You can send images in two ways. Each way starts a different generation mode: - **`frame_images`** — Gives first or last frame images for **image-to-video** generation. Each entry must include a `frame_type` of `first_frame` or `last_frame`. - **`input_references`** — Gives style or content reference images for **reference-to-video** generation. The model uses these images as visual guidance, not as accurate frames. If you send the two fields, `frame_images` controls the mode, and Sume processes the request as image-to-video. ###### Image-to-Video (frame_images) ```json { "model": "seedance-2", "prompt": "A character walking through a forest", "frame_images": [ { "type": "image_url", "image_url": { "url": "https://example.com/first-frame.png" }, "frame_type": "first_frame" } ], "resolution": "1080p" } ``` ###### Reference-to-Video (input_references) ```json { "model": "seedance-2", "prompt": "A colossal solar flare beside a planet", "input_references": [ { "type": "image_url", "image_url": { "url": "https://example.com/style-ref.png" } } ], "resolution": "1080p" } ``` A model accepts a reference type only if its `supported_input_references` includes that type. The Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max accept audio and video references. Gemini Omni Flash 1.1, `higgsfield-genjutsu`, and `h3-max-recast` accept video references but not audio. ##### Provider-Specific Options You can send provider-specific options in the `provider` parameter. Each provider slug is the key for its options. Sume forwards only the options for the matched provider: ```json { "model": "seedance-2", "prompt": "A time-lapse of a flower blooming", "provider": { "options": {} } } ``` To find the passthrough parameters that each model supports, read the `allowed_passthrough_parameters` field in the [Video Models API](#via-the-video-models-api). In v1, that list is empty for each model. Thus, Sume rejects `provider.options` entries with an error, and does not ignore them. Refer to [Sume differences](#sume-differences). #### Response Format ##### Submit Response (202 Accepted) When you submit a video generation request, you immediately receive a response with the job details: ```json { "id": "job_01HXYZ", "polling_url": "https://api.sume.com/v1/videos/job_01HXYZ", "status": "pending", "model": "seedance-2" } ``` ##### Poll Response When you poll the job status, the response includes more fields as the job continues: ```json { "id": "job_01HXYZ", "generation_id": "job_01HXYZ", "polling_url": "https://api.sume.com/v1/videos/job_01HXYZ", "status": "completed", "model": "seedance-2", "unsigned_urls": ["https://api.sume.com/v1/videos/job_01HXYZ/content?index=0"], "usage": { "cost": 0.25, "is_byok": false } } ``` ##### Job Statuses | Status | Description | | ------------- | ------------------------------------------------- | | `pending` | You submitted the job, and it is in the queue | | `in_progress` | The video generation is in progress | | `completed` | You can download the video | | `failed` | The generation failed (examine the `error` field) | | `cancelled` | The job is canceled, and it did not complete | ##### Downloading the Video When the job status is `completed`, the `unsigned_urls` array contains URLs for the download of the generated video content. You can also use the content endpoint directly: ```bash curl "https://api.sume.com/v1/videos/{jobId}/content?index=0" \ -H "Authorization: Bearer $SUME_API_KEY" \ --output video.mp4 ``` The default of the `index` query parameter is `0`. If the model generates more than one video output, use this parameter. #### Webhooks If you do not want to poll for the job status, you can receive a webhook notification when a video generation job completes. Send `callback_url` in the request body. When the job gets to a terminal state, Sume POSTs to that URL. Sume signs the raw JSON body and sends `x-sume-webhook-timestamp` and `x-sume-webhook-signature` headers. The payload is the standard job webhook envelope of Sume, not the OpenRouter `video.generation.*` envelope. For the accurate shape and the verification steps, refer to [Sume differences](#sume-differences) and the [webhooks guide](/api/reference). #### Sume differences All the information above agrees with the OpenRouter Video Generation API. This table gives the only differences. | Area | OpenRouter | Sume | | ------------------- | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | Base path | `https://openrouter.ai/api/v1/videos` | `https://api.sume.com/v1/videos` (no `/api` segment) | | Auth | `Authorization: Bearer $OPENROUTER_API_KEY` | `Authorization: Bearer $SUME_API_KEY` | | Model ids | `org/slug` (for example, `google/veo-3.1`) | bare catalog IDs (for example, `seedance-2`). The published contract of Sume never has a provider-org prefix | | Auto routing | no generate-time auto | `model: "sume/auto"` lets Sume select the family. Responses show `sume/auto`. Sume never discloses the family that ran | | `size` | accepted when the model shows `supported_sizes` | each v1 model reports `supported_sizes: null`. Thus, `size` returns `400 unsupported_parameter`. Use `resolution` + `aspect_ratio` | | `provider.options` | OpenRouter forwards it to the matched upstream provider | v1 runs one backend for each model. Thus, a non-empty `provider.options` returns `400 unsupported_parameter` | | `seed` | many models accept it | no v1 model accepts `seed`. Each model reports `seed: false` and rejects the field | | Webhook envelope | `video.generation.*` events, `X-OpenRouter-Signature` | Sume's standard job webhook envelope with `x-sume-webhook-signature` | | Idempotency | none on this route | send `Idempotency-Key` to make retries safe. A replay returns the original job | | Job lifecycle | only the polling URL | you can also see the same job at `GET /v1/jobs/{id}/status` and `GET /v1/jobs/{id}/result` | | Zero Data Retention | video generation is ZDR-ineligible | Sume has no ZDR toggle. Refer to the [privacy docs](/the-basics) | | Billing | credits | workspace USD balance. At submit, Sume reserves provider list × 1.25 (each model, `minimax-h3-max` included). `usage.cost` is the Sume billable amount | ##### `sume/auto` `sume/auto` is an addition that only Sume has. If you do not want to pin a family, send it as `model`: ```json { "model": "sume/auto", "prompt": "A vertical UGC-style product clip on a desk, natural light", "aspect_ratio": "9:16", "duration": 5 } ``` Resolution is a pure function of the normalized request and the catalog version. Thus, an idempotent replay gets the same price and the same route. The poll response reports `"model": "sume/auto"`. Sume does not disclose which family served the request. Do not use observable traits of the output to find the family. ##### Legacy `/v1/video-router/*` `POST /v1/video-router/generate` and `GET /v1/video-router/models` still work as before. They create the same jobs, with the same model IDs. The difference is the wire. Video Router returns the `{ "data": ... }` job envelope of Sume and accepts the flat `image_url` / `reference_image_urls` fields of Sume. `/v1/videos` returns the response shape in the sections above. The two APIs use the same model vocabulary. Thus, a migration is a path-and-body change, and you do not have to map IDs again. We recommend that new integrations use `/v1/videos`. #### Best Practices - **Detailed Prompts**: For better video quality, write specific prompts with much detail. Include details of motion, camera angles, lighting, and scene composition. - **Appropriate Resolution**: A higher resolution takes more time to generate and has a higher price. Select the resolution that is applicable to your use case. - **Polling Interval**: Use a moderate polling interval (for example, 30 seconds) to prevent too many API calls. Video generation usually takes from 30 seconds to several minutes, as a function of the model and the parameters. - **Error Handling**: Always examine the job status for the `failed` state. Make sure that your code processes the `error` field correctly. - **Reference Images**: When you use reference images, make sure that they have high quality and are applicable to the video that you want. #### Troubleshooting **Job stays in `pending` for a long time?** - Video generation can take several minutes, as a function of the model, the resolution, and the server load - Continue to poll at regular intervals **Generation failed?** - Examine the `error` field in the poll response for details - Make sure that the model guidelines permit your prompt - Make sure that all reference images are available over public HTTPS and are in supported formats **Model not found?** - Use the [Video Models API](#via-the-video-models-api) to find available video generation models - Make sure that the model ID is correct (for example, `seedance-2`). Sume uses bare catalog IDs, not `org/slug` #### Workspace fal keys In the development preview, workspace creators and admins can connect a fal API key in **Dashboard → Integrations → Model keys**. fal bills video models directly. Sume charges only **5.5% of fal list price**, under the workspace fee terms. Image jobs do not change. For BYOK video jobs, the poll response returns `usage.is_byok: true`, and `usage.cost` is only the Sume fee. `/v1/usage` includes `byok: { provider: "fal", basis_usd_micros: ... }`. Usage shows the label **Billed by provider** on the informational estimate. There is no fallback to Sume credentials. If fal rejects the key, or the key balance is empty, the job fails with `provider_byok_rejected`, and Sume refunds the fee. If you disconnect the key, the videos that still run on that key fail. But fal can still complete them and bill them. If Sume cannot resolve the key storage, it returns `503 byok_unavailable` before the submission. On production, this feature stays disabled by default. ##### Vercel AI Gateway keys for Agents In **Keys → Bring your own key**, immediately below fal, you can connect a Vercel AI Gateway API key. This key is for the agent model calls of your workspace. This development preview uses the same owner/admin controls, encrypted storage, key tests, replacement, and disconnect actions as fal. A Gateway key test uses its credit endpoint and does not generate paid content. Auto and pinned agent models use your Gateway account. Vercel bills the model. Sume charges only your workspace Agent Fee, which is 5.5% by default. Usage shows **BYOK · Gateway**, the Sume fee, and the approximate customer-paid provider basis as separate items. Gateway generation receipts give the actual Gateway debit. An upstream provider-key list cost or a token fallback is an estimate. If a turn fails, Sume refunds the Sume fee. But Vercel can still bill the requests that completed. In a BYOK turn, Sume never uses its own Gateway key as a fallback. If you remove or replace the key, an in-flight turn fails when its next credential check runs. Requests that Sume already accepted can still complete. After a disconnect, new turns use the usual Sume billing. This key is for the main agent model. Video/media and auxiliary services keep their current providers and billing. The key does not add upstream provider keys to your Vercel team. ### Video Router Source: https://docs.sume.com/models/video-router.md Explicit pass-through video generation via the Video Router catalog, including Seedance. We recommend that new integrations use [Video generation](/models/videos) (`POST /v1/videos`) in place of this API. It has the same catalog and the same jobs behind an OpenRouter-compatible wire. Video Router stays available and does not change. Video Router is the **explicit model catalog** of Sume for video generation. You select a catalog `model` ID from `GET /v1/video-router/models` (for example Seedance). Sume bills list × 1.25 on each model and returns the usual async job envelope. For new integrations, we recommend [Video Router Auto](/models/videos) through `POST /v1/videos` with `model: "sume/auto"`. Sume will retire [Video 1.0](/models/video) soon. Until then, it stays a compatibility alias for the same Auto pipe. The Auto create controls have a default of 720p and 8s, with 3–10s clips at 16:9 or 9:16. Primary invoke URL: ```text POST /v1/video-router/generate ``` Catalog: ```text GET /v1/video-router/models GET /v1/video-router/models/{model_id} ``` #### When to use | Goal | Approach | |---|---| | Let Sume select the model | Use [Video Router Auto](/models/videos) + `model: "sume/auto"` | | Pin a catalog model (for example, Seedance) | Video Router + `model` from the catalog | #### Create a Video Router job ```bash curl -X POST https://api.sume.com/v1/video-router/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-router-001" \ -d '{ "model": "seedance-2.5", "prompt": "A vertical UGC-style product clip on a desk, natural light", "resolution": "720p", "duration": 12, "aspect_ratio": "9:16", "mode": "async" }' ``` Limits are different for each model. `seedance-2.5` accepts 4–30s at 480p/720p/1080p. `wan-3.0` accepts 2–30s. `higgsfield-genjutsu` (Motion Transfer: `video_url` plus 1–8 `reference_image_urls`, 480p/720p) accepts 4–30s. This model is in the catalog only when its provider is configured. `h3-max-recast` (H3 Max Recast) accepts 5–30s. It replaces the people in a source `video_url` with the people from 1–4 `reference_image_urls`, with one photo for each person, at 768p/1080p. The prompt is optional, and the duration is the source length. `minimax-h3` accepts 5–15s at native 480p/768p (768p is first-class, not 720p). Each other catalog model has a limit of 15s. `seedance-2` also has 1080p. `minimax-h3-max` (MiniMax H3 Max) is the faster 768p variant. It does text-to-video, start/end-frame image-to-video, and reference-to-video (`reference_*_urls`) at 480p/768p/1080p (1080p is a latent refinement from native 768p) for 5–15s. Sume bills it at the × 1.25 house margin, as for each model. `gemini-omni-flash-1.1` (Gemini Omni Flash 1.1) is 3–10s at 360p/720p/1080p/4K, 16:9 or 9:16, and native synced audio is always on. Read `capabilities` from `GET /v1/video-router/models`, because each model can have a different envelope. #### Gemini Omni Flash 1.1 `gemini-omni-flash-1.1` is one catalog ID. Sume routes it by the shape of the request, and you never select an endpoint: | Capability | Send | Notes | |---|---|---| | `text_to_video` | `prompt` | 3–10s, `resolution` 360p–4K, `aspect_ratio` 16:9 / 9:16 | | `image_to_video` | `image_url` (+ optional `end_image_url`) | The same envelope | | `reference_to_video` | `reference_image_urls` (≤10) and/or `reference_video_urls` (≤3, each ≤3s) | Refer to the media as ``, `` (0-based, in list order) | | `video_to_video` (edit) | `video_url` | The prompt gives the instructions for the edit. `resolution` is optional (default 720p). No `aspect_ratio` / `duration` | Native audio is always on (the API rejects `generate_audio: false`). There is no `bitrate_mode` and no `reference_audio_urls`. Sume bills the provider list × 1.25 for each output second, as a function of the resolution. Edit example: ```bash curl -X POST https://api.sume.com/v1/video-router/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: video-router-edit-001" \ -d '{ "model": "gemini-omni-flash-1.1", "prompt": "Replace the bottle with an apple. Keep everything else the same.", "video_url": "https://example.com/clip.mp4", "resolution": "720p", "mode": "async" }' ``` `video_url` is the edit source, not a reference. You cannot send it with `image_url`, `end_image_url`, or `reference_*_urls`. For the full request schema and catalog fields, refer to the OpenAPI paths under **Video Router** in the [API reference](/api/reference). ### Music Router Source: https://docs.sume.com/models/music-router.md Music generation through the Music Router — sume/music-auto picks the engine (Lyria 3.5 today); explicit Lyria ids pass through. Music Router is the Sume model surface for music generation. It is in the same family as [Image Router](/models/images) and [Video Router](/models/videos). Send a prompt. If you omit `model` or set it to `sume/music-auto`, Sume selects the engine (Lyria 3.5 today). To pass through to one engine, pin an explicit catalog id from `GET /v1/music-router/models`. The retirement of [Music 1.0](/models/music) (`sume/music-1.0`) is gradual. Its routes continue to work and keep `job.model = sume/music-1.0`. But every request now resolves through the Music Router. For new integrations, we recommend the router. Primary invoke URL: ```text POST /v1/music-router/generate ``` Catalog: ```text GET /v1/music-router/models GET /v1/music-router/models/{model_id} ``` The public model id is `sume/music-router`. The routable ids are `sume/music-auto` (default), `lyria-3.5`, and `lyria-3-pro`. #### Request fields The body is the same as for [Music 1.0](/models/music), plus an optional `model`: | Field | Required | Notes | |---|---|---| | `model` | No | Routable id from the catalog. If you omit it, the value is `sume/music-auto`. Unknown ids fail with `400 model_not_found` and `catalog_url`. | | `prompt` | Yes | 1–5000 characters. Put exclusions in the positive prompt. Control the length in the prompt (“a 2-minute track”, `[0:00-0:30] Intro: …`). | | `image_url` | No | Optional public HTTPS image, or `null` to clear. | | `negative_prompt` | No | Sume does not support a non-empty value. Omit the field or send `""`. | | `metadata`, `mode`, `webhook_url`, `wait_timeout_seconds` | No | Same as on Music 1.0. | Sume rejects `duration` / `duration_seconds`. #### Create a Music Router job ```bash curl -X POST https://api.sume.com/v1/music-router/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: music-router-001" \ -d '{ "model": "sume/music-auto", "prompt": "Warm lo-fi hip hop, 84 BPM, C minor. Dusty Rhodes chords, brushed boom-bap drums, a muted trumpet answer at 0:10. A 30-second track. Instrumental, no vocals." }' ``` `job.model` echoes the routable id that you requested (`sume/music-auto` stays `sume/music-auto`). `job.request.routed_model` names the catalog engine that ran (for example `lyria-3.5`), on both the router and the Music 1.0 routes. #### Pricing Every Music Router model charges the fixed Music price per audio generation. The catalog lists the provider list price per model for reference. #### Poll and fetch the result ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` Read the audio artifact from `result.artifacts[]` where `type` is `audio`. When they are present, `result.lyrics` carries the model-reported lyrics or section map. ### Music 1.0 Source: https://docs.sume.com/models/music.md Generate music from text (and optional image conditioning) with Sume Music 1.0. The retirement of Music 1.0 is gradual. Its routes continue to work and keep `job.model = sume/music-1.0`. But every request now resolves through the [Music Router](/models/music-router) (`sume/music-auto`). For new integrations, we recommend `POST /v1/music-router/generate`. Use Music 1.0 for prompt-driven music generation. The request is text-first, and the model also supports optional image conditioning. Provider model ids stay internal. Music 1.0 runs on Google Lyria 3.5. It makes full-length structured songs of up to a few minutes, and the prompt controls them. For example, write "a 2-minute track" or use section markers such as `[0:00-0:30] Intro: ...`. Primary invoke URL: ```text POST /v1/music-1.0/generate ``` Model-run alias (same body): ```text POST /v1/models/sume/music-1.0/runs ``` The public model id is `sume/music-1.0`. #### When to use | Goal | Approach | |---|---| | Text → music | `prompt` only | | Visual conditioning | `prompt` + optional `image_url` | | Avoid specific styles | put exclusions in the positive `prompt` (for example, “no vocals, no spoken word”) | #### Writing the prompt (music brief) Music 1.0 accepts a text prompt and optional image conditioning. It has no seed, temperature, guidance, or duration parameter. To prevent generic musical choices, write a scene-specific brief with seven axes. These axes are creative directions, not guaranteed output values. Examine the generated audio. | Axis | Example | |---|---| | Emotion (precise, darker permitted) | “hushed, slightly melancholic”, “proud and nostalgic”, “cocky, restless” | | Genre / lineage | neo-soul, drill, bossa nova, gugak fusion, synthwave, chamber folk | | Tempo as a number | “72 BPM”, “142 BPM half-time” | | Key and mode | “D minor”, “E phrygian”, “G major with a lydian lift” | | Instruments with texture (2–4) | “Rhodes through tape wow”, “gayageum plucks”, “808 with long glide” | | Arc with one named moment | “breakdown to bass and claps at 0:20, full return at 0:28” | | Era / production | “1998 production, dry and close”, “2024 hyper-clean” | End the prompt with one clause: “Instrumental, no vocals.” Add “no spoken word” only under narration. For scenes in one project that contrast, change the broad genre family, the tempo (at least 12 BPM apart), and the lead instrument. If the user requests one consistent score, keep continuity. When it is applicable, pass the accepted scene still as `image_url`. After a policy rejection, change the flagged content, but keep the musical brief. Try again only in the authorized budget. Do not reduce the request to a generic bed. Provider `lyrics` can describe tempo and structure. These lyrics are model-reported metadata, not an audio measurement. ```text Hushed and slightly melancholic neo-soul nocturne, 72 BPM, D minor. Rhodes through tape wow, soft sub bass, brushed snare with rimshots, a single muted trumpet line. Sparse first half; the trumpet answers the Rhodes from 0:12 and the bass thickens for the last pass. Late-night, dry and close, 1998 production. Instrumental, no vocals. ``` #### Hard constraints - Do **not** send `duration` or `duration_seconds`. Music 1.0 does not recognize these fields and rejects them. - Do **not** send a non-empty `negative_prompt`. Music 1.0 / Lyria does not support negative prompts. A non-empty value returns HTTP 400 with `public_reason=negative_prompt_unsupported`. Omit the field or send `""`. - The maximum prompt length is 5000 characters. - Image URLs must be public HTTPS. Send `null` for `image_url` only when you intentionally clear an image input on a client that reuses request objects. #### Request fields | Field | Required | Notes | |---|---|---| | `prompt` | Yes | 1–5000 characters. Include exclusions in the positive prompt. | | `negative_prompt` | No | Sume does not support a non-empty value. Omit the field or send `""`. | | `image_url` | No | Optional public HTTPS image URL, or `null` to clear. | | `metadata` | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider. | | `mode` | No | `async`, `sync`, `subscribe`, `webhook`. | | `webhook_url` | No | Public HTTPS callback for webhook mode. | | `wait_timeout_seconds` | No | 0–30 for `sync` / `subscribe`. | #### Create a music job Image-conditioned example: ```bash curl -X POST https://api.sume.com/v1/music-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: music-image-001" \ -d '{ "prompt": "Cinematic ambient underscore matching the mood of the reference still, instrumental only", "image_url": "https://example.com/moodboard.png" }' ``` #### Poll and fetch the result ```bash curl https://api.sume.com/v1/jobs/job_123/status \ -H "Authorization: Bearer $SUME_API_KEY" curl https://api.sume.com/v1/jobs/job_123/result \ -H "Authorization: Bearer $SUME_API_KEY" ``` If the job is successful, read the audio artifact from `result.artifacts[]` where `type` is `audio` (usually `audio/mpeg` on `media.sume.com`). #### Artifacts Completed Music 1.0 jobs return Sume-hosted audio artifacts: ```json { "id": "job_...", "status": "completed", "result": { "artifacts": [ { "id": "artifact_...", "type": "audio", "url": "https://media.sume.com/artifacts/...", "content_type": "audio/mpeg" } ] } } ``` Use Sume media URLs from the result. Raw provider URLs are not public outputs. #### Pricing The price is a fixed **$0.125 USD** for each accepted Music 1.0 generation. The price does not change with prompt length or optional image conditioning. #### Next - [Jobs and results](/workflows/jobs-and-results) for polls and webhooks - [Recipes](/api/cookbook) for short copy-paste flows ## Dashboard ### API keys Source: https://docs.sume.com/dashboard/api-keys.md Manage Sume Developer API keys from the dashboard. Use the [API Keys dashboard](https://www.sume.com/dashboard/api-keys) to create and manage API keys for the Developer API. API keys are workspace-scoped. The dashboard shows the full secret only when you create a key. Store the secret immediately in a secure secret manager. For request header examples and safety rules, refer to [Authentication](/authentication). ### Playground Source: https://docs.sume.com/dashboard/playground.md Use the Sume Avatar playground for scoped API experiments. Open the [Avatar playground](https://www.sume.com/playground) to try Avatar generation manually in the context of your signed-in workspace. (`/playground` routes to the current Avatar playground surface.) Use the playground to validate payloads before you move them into server code, CLI commands, or agent workflows. For the accurate request and response fields, use the live [OpenAPI schema](https://api.sume.com/reference/json) (local docs snapshot: [/api/openapi.json](/api/openapi.json)). ### Jobs Source: https://docs.sume.com/dashboard/jobs.md Inspect generation jobs and operational status in the dashboard. Open the [Jobs dashboard](https://www.sume.com/dashboard/jobs) to examine recent Developer API jobs. The dashboard shows the status, the mode, the public result or error data, and the timeline information when available. For programmatic access, use: ```text GET /v1/jobs GET /v1/jobs/:id GET /v1/jobs/:id/status GET /v1/jobs/:id/result GET /v1/jobs/:id/events POST /v1/jobs/:id/cancel ``` The live [OpenAPI](https://api.sume.com/reference/json) gives the accurate response shapes. The [API reference](/api/reference) gives a readable map. ### Usage Source: https://docs.sume.com/dashboard/usage.md Inspect Sume usage ledger entries and balance. Open the [Usage dashboard](https://www.sume.com/dashboard/usage) for a summary of metered Developer API activity that a person can read. You can get programmatic usage reads from: ```text GET /v1/balance GET /v1/usage ``` The live [OpenAPI](https://api.sume.com/reference/json) gives the accurate schemas. #### Balance ```bash curl https://api.sume.com/v1/balance \ -H "Authorization: Bearer $SUME_API_KEY" ``` The balance is in USD. Compatibility fields can show rounded cent values as credits. #### Usage ledger ```bash curl https://api.sume.com/v1/usage \ -H "Authorization: Bearer $SUME_API_KEY" ``` Usage entries can include reservations, captures, refunds, top-ups, and grants. The entries change with the workspace and the environment. | Status | Meaning | |---|---| | `reserved` | Sume reserved the estimated usage before provider execution. | | `captured` | Sume captured the billable usage after a successful completion. | | `refunded` | Sume released the reserved usage after a failure or a cancellation before capture. | Use job ids and request ids to correlate usage rows with generation workflows. ##### One thread, one run, one job Add `thread_id`, `run_id` or `job_id` to the request. Then the response gives the accurate cost of one Studio Agent thread, one Format / Action / Agent run, or one generation job: ```bash curl "https://api.sume.com/v1/usage?run_id=arun_…&limit=50" \ -H "Authorization: Bearer $SUME_API_KEY" ``` The response adds `summary`. The summary folds over every ledger row that the scope caused (`limit` only caps the listed rows): | Field | Meaning | |---|---| | `debited_usd_micros`, `debited_usd` | The amount that the wallet deducted: captured rows of every operation type (the agent's own turns are included). Quote this figure. | | `held_usd_micros` | Holds that are still open, with parked `pending_*` rows included. This is not spend yet. | | `refunded_usd_micros` | Holds that Sume gave back after a failure, cancellation or `queue_full`. This is not spend. | | `final` | `true` when no hold is open. | | `includes` | Row counts by kind: LLM turns, sidecars, Browser sessions, generation jobs, `script_run_children` (jobs that a `script_run` call dispatched), refunded rows, and BYOK rows. | | `script_runs` | The `script_run` calls in the scope. Each call has its rows and money. | | `by_operation_type`, `runs` | The same money by operation type. For a thread, also by run. | | `cap` | For a run: its generation cap, the amount that counts against the cap, and the amount that is left. The amount against the cap is reserved + captured generation rows, without LLM (the receipt's `billable_amount_usd_micros`). This value is never a cost. | Rows keep their ledger ids. Rows now also carry `thread_id`, `run_id`, `turn_job_id` (the agent turn that commissioned the job), `script_run_id` / `script_run_call_index` (when a `script_run` dispatched the job) and `settle_state`. `job_id` also accepts the job id of a turn. Then the sum includes the turn's own row and every job that the turn commissioned. Never sum the rows yourself. A refunded row keeps its hold amount in `billable_amount_usd_micros`. The [API pricing page](https://www.sume.com/pricing/api) gives the metered rates for each capability. This page does not include them. Sume generates the pricing page from the same pricing constants that the API uses for bills. For plan changes and top-ups, use [Billing & subscription](https://www.sume.com/dashboard/subscription). ### Billing & subscription Source: https://docs.sume.com/dashboard/credits.md Billing, subscription, and credit top-up behavior in the Sume dashboard. Open [Billing & subscription](https://www.sume.com/dashboard/subscription) to examine the plan status and the available balance. You can also buy credits there for Developer API usage. (The product nav label is **Billing & subscription**. Older docs called this surface “Credits”.) The [API pricing page](https://www.sume.com/pricing/api) shows the cost of a generation. [Pricing](https://www.sume.com/pricing) shows the subscription plans. #### Current public reads ```text GET /v1/balance GET /v1/usage ``` These endpoints are read-only. Their scope is the workspace of the authenticated API key. The live [OpenAPI](https://api.sume.com/reference/json) gives the accurate schemas. #### Manual top-up When billing is configured for the workspace, the dashboard can start a Stripe-backed manual credit top-up. After checkout completes, the usage ledger can show the top-up activity and the updated balance. For API clients, a top-up is a dashboard operation. At this time, the public Developer API gives balance and usage reads. It does not give a public endpoint to create a top-up. ## CLI ### Overview Source: https://docs.sume.com/cli.md Current Sume CLI surface for the sume.com developer platform. The Sume CLI is an agent-first wrapper over the current `api.sume.com/v1` Developer API. The CLI is intentionally separate from older consumer-product CLI surfaces. #### Current status The shipped CLI (`sumelabs/cli`) supports account verification, catalog discovery, jobs, and input assets. The CLI also supports Avatar 1.0, Avatar Video 1.0, bundled agent skills, schema discovery, and local diagnostics. **Most of the generation coverage is for Avatar**. At this time, the CLI has no first-class subcommands for Image 1.0, Video 1.0, or Music 1.0. For those families, call the [Developer API](/public-api) directly (refer to [Image](/models/image), [Video](/models/video), [Music](/models/music)). Local `sume mcp` is **not available** in current CLI releases (`coming_soon`). For remote MCP in Cursor/Claude, use the [hosted MCP](/mcp) connector at `https://mcp.sume.com/mcp`. Use the hosted installer to install the current native binary: ```bash curl https://cli.sume.com/install -fsS | bash ``` For Windows PowerShell: ```powershell irm https://cli.sume.com/install.ps1 | iex ``` The installer finds the latest `sumelabs/cli` GitHub Release. Then it downloads the binary for your platform and compares the binary with `checksums.txt`. Last, it installs `sume` in `~/.sume-com/bin`. Then use the browser flow to sign in: ```bash sume login sume auth status ``` #### API base ```text https://api.sume.com/v1 ``` #### Common commands ```bash sume login sume auth status sume account get sume balance sume usage get --limit 20 sume catalog list sume jobs list sume jobs status sume jobs result sume avatars create --confirm-paid --avatar-handle studio_presenter --type prompt --prompt "Friendly presenter" sume avatar-videos create --confirm-paid --avatar-handle sume_clawra --script "Say hello" --quality plus ``` For stable machine-readable output, use `--json`. When automation reads output that can contain sensitive URLs or account metadata, use `--agent`. The estimated length of an avatar-video script must be 4-60 seconds inclusive. #### Next pages 1. [Install and update](/cli/install-login) 2. [Command reference](/cli/commands) 3. [Generation workflows](/cli/generation-workflows) 4. [Hosted MCP](/mcp) — remote connector (at this time, we recommend it over local `sume mcp`) ### Install and update Source: https://docs.sume.com/cli/install-login.md Install or run the Sume CLI for current developer-platform workflows. The recommended public install path is the hosted installer for the native [`sumelabs/cli`](https://github.com/sumelabs/cli) binary. The binary is the `sume.com` developer-platform CLI for `api.sume.com` workflows. #### Install on macOS or Linux ```bash curl https://cli.sume.com/install -fsS | bash ``` The installer downloads the latest release binary for your OS and architecture. Then it compares the binary with `checksums.txt` and installs `sume` in `~/.sume-com/bin`. If a different `sume` is already on your `PATH`, the installer does not silently overwrite it. #### Install on Windows PowerShell ```powershell irm https://cli.sume.com/install.ps1 | iex ``` #### Direct GitHub binary fallback The names of the current release assets are `sume-darwin-arm64`, `sume-darwin-x64`, `sume-linux-arm64`, `sume-linux-x64`, and `sume-windows-x64.exe`. Each GitHub Release also has `checksums.txt` attached. For a pinned install, replace `latest` in the URL with a release tag such as `v0.1.6`. ```bash OS="$(uname -s | tr '[:upper:]' '[:lower:]')" ARCH="$(uname -m | sed 's/x86_64/x64/;s/aarch64/arm64/')" curl -fsSL "https://github.com/sumelabs/cli/releases/latest/download/sume-${OS}-${ARCH}" -o /tmp/sume chmod +x /tmp/sume sudo mv /tmp/sume /usr/local/bin/sume ``` #### Sign in For the default auth flow, use browser login: ```bash sume login sume auth status ``` For remote or headless environments: ```bash sume login --no-browser ``` Manual API-key setup stays available for CI and controlled server environments. But it is not the recommended first-run path for local users. #### Development build Run these commands from the CLI source repository: ```bash pnpm install pnpm run build pnpm dev -- --help ``` When necessary, build a local binary: ```bash pnpm run build:binary ./dist/sume --version ``` #### Verify ```bash sume version sume doctor --agent --json sume update --check ``` `sume update --check` shows if a newer GitHub Release exists. It does not change local files. When you decide to upgrade, run the hosted installer again. ### Authentication Source: https://docs.sume.com/cli/authentication.md Configure Sume API keys for the CLI. For the default local CLI flow, use browser login. The CLI opens the device approval page at [www.sume.com/cli/login](https://www.sume.com/cli/login) (with a `user_code`). Then the CLI waits for approval and stores a CLI-scoped API key in the local config. As an alternative, you can create a key manually in the [API Keys dashboard](https://www.sume.com/dashboard/api-keys). ```bash sume login sume auth status ``` For a remote or headless terminal, use this command. It prints the approval URL and does not open a browser: ```bash sume login --no-browser ``` The CLI still supports manual API-key setup and environment variables for CI, server automation, and advanced local workflows: ```bash sume auth setup --api-key "$SUME_API_KEY" ``` ```bash export SUME_API_KEY="sume_live_..." export SUME_API_BASE_URL="https://api.sume.com/v1" export SUME_API_AUTH_MODE="x-api-key" ``` `x-api-key` is the current CLI default. The CLI also supports `SUME_API_AUTH_MODE=bearer` for clients that select `Authorization: Bearer`. By default, the CLI stores the local config in `~/.sume-com/config.json`. ### Configuration Source: https://docs.sume.com/cli/configuration.md Environment variables and local configuration used by the Sume CLI. #### Environment variables | Variable | Purpose | |---|---| | `SUME_API_KEY` | Sume Developer API key. | | `SUME_API_BASE_URL` | API base URL. The default is `https://api.sume.com/v1`. | | `SUME_API_AUTH_MODE` | `x-api-key` or `bearer`. | | `SUME_APP_BASE_URL` | App base URL that `sume login` uses. For the production API, the default is `https://app.sume.com`. For other APIs, the default is the origin of `SUME_API_BASE_URL`. | | `SUME_CONFIG_DIR` | Overrides the local config directory for tests or isolated environments. | #### Local config The CLI stores the local configuration in this file: ```text ~/.sume-com/config.json ``` Use the recommended browser login flow to create the local config: ```bash sume login ``` To examine the local readiness without a call to the API, use `sume doctor --agent --json`. ### Command reference Source: https://docs.sume.com/cli/commands.md Current Sume CLI command groups and safety gates. This page lists the current developer-platform CLI surface from `sumelabs/cli`. This page intentionally does not include older consumer-product commands that are not supported. Image/Video/Music generators do not exist as CLI subcommands at this time. Use these commands to find the exact schemas at runtime: ```bash sume tools list --json sume tools schema --json ``` #### Auth and account ```bash sume login sume login --no-browser sume logout sume auth status sume auth setup --api-key "$SUME_API_KEY" sume me sume account get sume balance sume usage get --limit 20 ``` `sume me` and `sume account get` both read account context for the configured key. #### Read commands ```bash sume catalog list sume health sume health service sume health v1 sume doctor --agent --json sume jobs list sume jobs get sume jobs status sume jobs events sume jobs result sume assets list sume assets get sume avatars list sume avatars get sume avatar-videos list sume avatar-videos get sume tools list --json sume tools schema jobs.result --json sume skills list sume version sume update --check ``` `sume models list` stays as a deprecated alias for `sume catalog list`. #### Jobs helpers ```bash sume jobs watch sume jobs download --output-dir ./out sume jobs cancel --confirm-submit ``` #### Assets ```bash sume assets create --source-url https://example.com/reference.png --confirm-submit sume assets upload-url --content-type image/png --size-bytes 12345 --confirm-submit sume assets complete --confirm-submit sume assets download-url sume assets download --output-dir ./assets ``` Asset registration and upload paths are write-gated. When first-party asset ids are not necessary, it is better to use public HTTPS URLs in generation payloads. Refer to [Jobs and media](/cli/jobs-assets). #### Write and paid generation (Avatar) Each write command must have explicit confirmation. At this time, the paid Avatar generators are the only first-class path to submit a generation in the CLI. ```bash sume avatars create --confirm-paid --avatar-handle studio_presenter --type prompt --prompt "Friendly presenter" sume avatars create --confirm-paid --avatar-handle reference_presenter --type photo --image-url https://example.com/reference.png sume avatar-videos create --confirm-paid --avatar-handle sume_clawra --script "Say hello" --quality plus sume avatar-videos create --confirm-paid --avatar-handle sume_clawra --product-image https://example.com/product.png --script "Introduce this product naturally" --quality plus ``` When provider execution can spend credits, use `--confirm-paid`. For non-paid writes such as job cancellation or asset registration, use `--confirm-submit`. Before you submit, make sure that the estimated length of the avatar-video script is 4-60 seconds inclusive. If you do not send Avatar Video `quality`, the default is **`plus`**. The CLI also has batch helpers in `sume avatars batch` and `sume avatar-videos batch` (plan / create / watch / result against local state files). To examine them, use `sume avatars batch --help`. #### Not available as CLI generators The CLI does **not** ship `sume image`, `sume video`, `sume music`, or similar subcommands. For these families, use the HTTP API: | Family | Guide | |---|---| | Image 1.0 | [Image](/models/image) | | Video 1.0 | [Video](/models/video) | | Music 1.0 | [Music](/models/music) | After you use the API to submit those jobs, you can still recover them. Use `sume jobs status` / `sume jobs result` / `sume jobs download`. #### MCP (CLI vs hosted) ```bash sume mcp doctor --json ``` In current releases, local `sume mcp` reports `coming_soon` / `launched: false`. `sume mcp install` is not available at this time. For Cursor, Claude, and other remote MCP clients, use the [hosted MCP](/mcp) endpoint at `https://mcp.sume.com/mcp`. That endpoint is a surface that is separate from the local CLI command. #### Machine-readable output ```bash sume catalog list --json sume jobs result --agent --json ``` `--agent` redacts URL-like and sensitive account/workspace fields where the CLI supports agent-safe output. #### Skills ```bash sume skills list sume skills install sume skills update sume skills remove sume skills export ``` Refer to [Agent skills](/cli/agent-skills). ### Jobs and media Source: https://docs.sume.com/cli/jobs-assets.md Recover jobs, download artifacts, and register public HTTPS media inputs with the Sume CLI. #### Jobs ```bash sume jobs list --agent --json sume jobs get --agent --json sume jobs status --agent --json sume jobs events --agent --json sume jobs result --agent --json sume jobs watch sume jobs download --output-dir ./out sume jobs cancel --confirm-submit ``` If a process restarts or a local timeout occurs, use the job recovery commands. Do not submit paid generation again. `jobs watch` polls until the job gets to a terminal state or until a timeout occurs. `jobs download` writes completed media artifacts into a local directory. These job helpers work for Avatar CLI submits **and** for Image/Video/Music jobs that you created through the [Developer API](/public-api). #### Assets If a workflow must use asset ids in place of bare public URLs (or together with them), register or upload first-party input assets: ```bash sume assets list --agent --json sume assets get --agent --json sume assets create --source-url https://example.com/reference.png --confirm-submit sume assets upload-url --content-type image/png --size-bytes 12345 --confirm-submit sume assets complete --confirm-submit sume assets download-url sume assets download --output-dir ./assets ``` To examine the exact flags, use `sume tools schema assets.create --json` and the related schemas. When the first-party asset lifecycle is not necessary, it is better to use public HTTPS media URLs in generation payloads. #### Use media in Avatar generation Avatar launch fields such as `input.image_url`, `product_image`, and `scene.image_url` usually accept public HTTPS URLs. To find the exact request-body fields, examine these schemas: ```bash sume tools schema avatars.create --json sume tools schema avatar-videos.create --json ``` For Image/Video/Music media fields, obey the model guides ([Image](/models/image), [Video](/models/video), [Music](/models/music)). For the CLI, those families are API-first. ### Generation workflows Source: https://docs.sume.com/cli/generation-workflows.md Submit Avatar 1.0 and Avatar Video 1.0 from the CLI; use the API for Image, Video, and Music. #### CLI generation coverage | Family | CLI submit today | Recommended path | |---|---|---| | Avatar 1.0 | `sume avatars create` | CLI or [Avatar guide](/models/avatar) | | Avatar Video 1.0 | `sume avatar-videos create` | CLI or [Avatar Video guide](/models/avatar-videos) | | Image 1.0 | **No CLI subcommand** | [Developer API](/models/image) | | Video 1.0 | **No CLI subcommand** | [Developer API](/models/video) | | Music 1.0 | **No CLI subcommand** | [Developer API](/models/music) | Do not try to use Image/Video/Music CLI commands. These commands do not exist. Call `POST /v1/{image,video,music}-1.0/generate` (or the model-run alias `POST /v1/models/sume/-1.0/runs`) and send an `Idempotency-Key`. If you want the CLI to poll or download the result, use `sume jobs …` to recover the job. Hosted MCP is a different surface. Hosted MCP has `generate_image` / `generate_video` / `music_create`. The CLI still has no Image/Video/Music subcommands. Refer to [MCP overview](/mcp). #### Avatar 1.0 ```bash sume avatars create \ --confirm-paid \ --avatar-handle studio_presenter \ --type prompt \ --prompt "A friendly presenter in neutral studio lighting" \ --json ``` The command submits the primary Avatar route `POST /v1/avatar-1.0/generate` (the API keeps legacy model-run aliases for compatibility). Use `--type prompt`, `--type photo`, or `--type props`. To send an exact request body, use `--payload-json` or `--payload-file`. The optional `--model` flag accepts an exact model-run id when that path is necessary. #### Avatar Video 1.0 ```bash sume avatar-videos create \ --confirm-paid \ --avatar-handle sume_clawra \ --product-image https://example.com/product.png \ --script "Say hello to the Sume developer platform." \ --quality plus \ --json ``` The command submits `POST /v1/avatar-1.0/talking-video`. To read avatar handles, use `sume avatars list --agent --json` or `sume avatars get --agent --json`. The estimated length of an avatar-video script must be 4-60 seconds inclusive. The CLI validates this length before it submits the job. `quality` accepts `standard`, `plus`, or `max`. If you do not set it, the CLI default is **`plus`** (the same as the public API). For the fastest path, use `--quality standard`. When quality is more important than turnaround, use `--quality max`. #### Image, Video, and Music (API-first) This example uses curl to submit an Image job. Then it uses the CLI to recover the job: ```bash curl -sS https://api.sume.com/v1/image-1.0/generate \ -H "Authorization: Bearer $SUME_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{"prompt":"Studio product photo on white background"}' sume jobs status --agent --json sume jobs result --agent --json ``` Full request fields and constraints: - [Image 1.0](/models/image) - [Video 1.0](/models/video) - [Music 1.0](/models/music) #### Recover after submit All submit commands return or print a job id. To recover, use the job commands. Do not submit paid work again: ```bash sume jobs status --agent --json sume jobs events --agent --json sume jobs result --agent --json sume jobs watch sume jobs download --output-dir ./out ``` Where the CLI exposes those flags, submit commands support exact payloads, idempotency keys, communication modes, webhook URLs, and bounded wait timeouts. ### Troubleshooting Source: https://docs.sume.com/cli/troubleshooting.md Debug common Sume CLI auth, config, job, media input, and release issues. Start with read-only checks: ```bash sume version sume auth status sume doctor --agent --json sume account get --json ``` #### Missing API key Start with the browser login: ```bash sume login sume auth status ``` For remote or headless terminals: ```bash sume login --no-browser ``` For CI and server automation, you can use a manual API-key setup and environment variables: ```bash sume auth setup --api-key "$SUME_API_KEY" ``` ```bash export SUME_API_KEY="sume_live_..." export SUME_API_BASE_URL="https://api.sume.com/v1" ``` #### Wrong API base Do a check of the local configuration: ```bash sume doctor --agent --json ``` The current production API base is: ```text https://api.sume.com/v1 ``` #### Job created but no result yet Jobs are async. Poll the job status: ```bash sume jobs status --agent --json ``` Get the result only after the job is complete: ```bash sume jobs result --agent --json ``` If a local wait timed out, examine the job before you submit another paid job. #### Image / Video / Music CLI command missing The shipped CLI has no `sume image`, `sume video`, or `sume music` subcommand. Submit those families through the Developer API ([Image](/models/image), [Video](/models/video), [Music](/models/music)). Then use `sume jobs status` / `sume jobs result` to recover the job. #### Local MCP not launched ```bash sume mcp doctor --json ``` Current releases report `coming_soon`. For Cursor/Claude, use [hosted MCP](/mcp) at `https://mcp.sume.com/mcp`. Or, use the CLI commands directly. #### Media input rejected Avatar media fields must contain public HTTPS image URLs. For Avatar 1.0 photo input, use `--type photo --image-url https://...`. For Avatar Video, use `--product-image https://...` or `--scene-image-url https://...` when necessary. Frequent causes: 1. The URL is not HTTPS. 2. The URL points to localhost or to a private network. 3. The response is not an image. 4. The URL must have cookies, auth headers, or a short-lived signature. The signature expires before Sume can fetch the URL. Do not paste private media URLs, signed URLs, or auth headers into public logs. #### Report an issue Include this data: - The command name and flags, with the secrets redacted. - `sume version` - The sanitized error code/message. - The request id, if there is one. - The job id, if the issue is job-specific. - The source of the auth: env or local config. Do not include API keys, signed URLs, private media URLs, raw provider payloads, emails, or workspace/user ids. ### Security Source: https://docs.sume.com/cli/security.md Handle Sume CLI secrets, media outputs, paid calls, and agent automation safely. We designed the Sume CLI for agent use. But CLI output can still contain sensitive data or data that users own. Be careful with auth, job, and media output. #### Secrets Do not print or commit these items: - `SUME_API_KEY`, - local `~/.sume-com/config.json`, - raw provider payloads. For automation, use environment variables or secret managers: ```bash export SUME_API_KEY="sume_live_..." sume account get --json ``` For local development, the best choice is `sume login`. Then the CLI uses the [CLI login approval](https://www.sume.com/cli/login) flow to create and store a scoped key. Manual keys are in the [API Keys dashboard](https://www.sume.com/dashboard/api-keys). #### Write and paid gates You must add a confirmation flag to current write commands. | Gate | Use for | |---|---| | `--confirm-submit` | Non-paid writes such as job cancellation. | | `--confirm-paid` | Provider-backed Avatar 1.0 and Avatar Video 1.0 runs that can reserve or spend credits. Image/Video/Music are API-first at this time. Apply the same confirmation discipline in your HTTP client. | Agents must confirm the intent of the user. When agents do a test, they must submit one bounded job first. Agents must recover the jobs that already exist and must not blindly retry paid commands. #### Media URLs Sume job results can include first-party media URLs. These URLs are public, but they are still user data. When you report results: - summarize the media counts and the file types, - when possible, use local filenames, not full remote URLs, - redact query strings and private identifiers, - do not dump large raw result payloads. #### Bug reports When they help, include sanitized command names, error codes, request ids, and job ids. Do not include API keys, signed URLs, full private media URLs, raw provider payloads, emails, or workspace/user ids. The only exception is when engineering explicitly asks for them. ### Agent skills Source: https://docs.sume.com/cli/agent-skills.md Agent-safe Sume CLI usage patterns and bundled skills install. We designed the Sume CLI commands to make agent automation explicit and auditable. #### Use read-only first While you plan, use read commands first: ```bash sume doctor --agent --json sume catalog list --json sume tools list --json sume jobs status --agent --json sume jobs result --agent --json ``` #### Require confirmation for writes Agents must not create resources without explicit operator confirmation. Agents must not submit paid provider work without explicit operator confirmation. | Gate | Use for | |---|---| | `--confirm-submit` | Non-paid writes such as job cancellation or asset registration. | | `--confirm-paid` | Generation that can reserve or spend credits (Avatar / Avatar Video). | At this time, the CLI has no submit commands for Image, Video, and Music generation. For those families, agents must call the [Developer API](/public-api). Agents can still use CLI job recovery when it helps. #### Redact sensitive output When automation reads command output, use `--agent --json`. Where the CLI supports it, agent mode redacts or summarizes URL-like fields and account/workspace details. #### Bundled skills Use these commands to install or refresh the packaged Sume skill in the local agent skill directories: ```bash sume skills list sume skills install sume skills update sume skills export sume sume skills remove sume ``` `sume skills install` writes into `.agents/skills` or `.claude/skills`. Use `export` to examine the source files before a custom install. #### MCP for agents For Cursor/Claude remote connectors, the best choice is [hosted MCP](/mcp) (`https://mcp.sume.com/mcp`). Local `sume mcp` is not available in current CLI releases. Refer to `sume mcp doctor --json` and the [CLI command reference](/cli/commands). #### Recommended loop 1. Examine the catalog and the local readiness. 2. Use a read-only/schema command to validate the payload. 3. Before writes or paid work, ask for confirmation. 4. Submit one bounded job (the CLI for Avatar, the API for Image/Video/Music). 5. Use the job status/events/result commands to recover. 6. Summarize the outputs. Do not paste secrets or signed URLs. ## MCP ### Overview Source: https://docs.sume.com/mcp.md Hosted Sume MCP for Cursor, Claude, and other MCP clients. Sume has a hosted [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) endpoint. Agents can use this endpoint to call Sume account, catalog, job, asset, generation, crawl, and Avatar tools. It is not necessary to wrap the HTTP API yourself. #### Production endpoint ```text https://mcp.sume.com/mcp ``` Use this URL in Cursor, Claude Code, Codex, and other remote MCP clients. Internal/dev note: `https://mcp.dev.sume.com/mcp` is for Sume development environments. Public docs and customer configs must use `mcp.sume.com`. #### Choose the right surface | Surface | What it is | When to use | |---|---|---| | Hosted MCP | Remote HTTP MCP at `https://mcp.sume.com/mcp` | Connect Cursor/Claude/Codex to Sume through OAuth or an API key. **Preferred MCP path today.** | | Local `sume mcp` | CLI MCP command (stdio client setup) | **Not launched** in the current `sumelabs/cli` releases (`sume mcp doctor` → `coming_soon`). Use hosted MCP or direct CLI commands. | | Developer API / CLI | `https://api.sume.com/v1` and `sume` binary | Backend integrations, scripts, and HTTP submits. Hosted MCP covers most generation families. Image 1.0 / Video 1.0 (`images_create` / `videos_create`) stay REST-only. | | Studio Agent | Product agent experience in the Sume app | In-product agent workflows. It is not a public MCP connector, and this page does not document it. | Hosted MCP is the supported remote connector today. Local `sume mcp` is a future CLI surface. Do not document it or configure it as an operational stdio server at this time. The Studio Agent product surface is also not a supported remote connector. Also refer to the [CLI overview](/cli). #### Auth at a glance | Auth | Hosted MCP capability today | |---|---| | OAuth | `mcp:read` (required) sees read-only tools. To expose the tools that change data and the paid tools, opt in to `mcp:write` on the MCP-host consent page. There is **no** `mcp:paid` scope. | | API key | Full hosted tool set. Paid/write calls still need `idempotency_key`. Wallet/admission is the spend gate. | OAuth is the preferred path for interactive clients, for example Cursor and Claude. API-key remote MCP is still available for current automation. For details, refer to [OAuth and API keys](/mcp/oauth). #### What hosted MCP can do today Hosted MCP does **not** have full parity with the HTTP API. The live tool ids are the underscore names in `tools_list` (`packages/mcp-server/src/mcp.ts` `remoteMcpTools`). The server canonicalizes dotted aliases (`tools.list`) to the underscore names. The shipped paid generation tools include more than Avatar: - `generate_image` / `generate_video` (omit `payload.model` to route to `sume/auto`, and get catalog ids from `image-models_list` / `video-router_models`) - `music_create`, `tts_create`, `stt_create` - `avatars_create`, `avatar-videos_create`, Stage P preview tools - `kling-motion-control_create`, `image_upscale_create`, `rmbg_create`, `video_upscale_create`, `timeline_create` / `timeline_audio` / `timeline_compose` / `timeline_get` Web and social reads use the `crawl_*` family (`crawl_scrape` / `crawl_map` / `crawl_search` / `crawl_site` / `crawl_get`, plus `crawl_profile` / `crawl_feed` / `crawl_media` / `crawl_find`). **REST-only (not in `tools_list`):** Sume Image 1.0 and Video 1.0 (`images_create` / `videos_create`). Use the [Developer API](/public-api) for those two products. Discovery tools, for example `catalog_list`, `tools_list`, and `tools_schema`, tell the client accurately which MCP tools are available in the current session. #### Next pages 1. [Quickstart](/mcp/quickstart) — connect Cursor or Claude in a few minutes. 2. [Tools and gates](/mcp/tools-and-gates) — tool inventory, safety flags, and playbooks. 3. [OAuth and API keys](/mcp/oauth) — auth matrix and `mcp:read` / `mcp:write`. 4. [Agents](/agents) · [Safe automation](/agents/safe-automation) — how agents must treat API / CLI / MCP boundaries. ### Quickstart Source: https://docs.sume.com/mcp/quickstart.md Connect Cursor or Claude to hosted Sume MCP. Connect a remote MCP client to the Sume production endpoint. Complete OAuth (or use an API key). Then call a read-only discovery tool. #### 1. Use the production URL ```text https://mcp.sume.com/mcp ``` Do not paste API keys into chat. For interactive clients, the OAuth connector flow is the preferred method. #### 2. Connect Claude Code ```bash claude mcp add --transport http sume https://mcp.sume.com/mcp claude mcp login sume ``` In the Claude Code MCP status UI, make sure that the server is connected. Then ask Claude to call `tools_list` or `mcp_health`. #### 3. Connect Cursor Add a remote MCP server entry (Cursor Settings → MCP, or your MCP config file): ```json { "mcpServers": { "sume": { "url": "https://mcp.sume.com/mcp" } } } ``` When Cursor prompts you, complete the OAuth sign-in. The consent page is on the MCP host, not `app.sume.com`. After you connect, call `tools_list` one time to make sure that the session works. #### 4. Connect Codex or other HTTP MCP clients Set the URL of a streamable HTTP MCP server to: ```text https://mcp.sume.com/mcp ``` Run the OAuth login of the client for the configured server name. The client will discover the Sume protected-resource metadata from the MCP endpoint. Then the client will send you to `https://mcp.sume.com/oauth/authorize`. This URL continues to `https://mcp.sume.com/oauth/consent`. #### 5. Verify with a read-only tool Ask the agent: ```text Call the Sume MCP tool tools_list and summarize the available tools. Live ids use underscores (tools_list, generate_image). Dotted aliases also work. ``` Useful first calls: | Tool | Why | |---|---| | `mcp_health` | Confirms the endpoint, the auth source, and the safety posture. | | `tools_list` | Lists all the tools that this session can see. | | `tools_schema` | Returns one tool contract by `name`. | | `account_me` | Confirms the workspace account context. | | `catalog_list` | Lists the public API capabilities (more than the MCP tools). | #### 6. Know the OAuth scopes By default, hosted OAuth grants **read-only** access (`mcp:read`). To also get `mcp:write`, set the **Write** toggle to on at consent. There is no `mcp:paid` scope. If the session has only `mcp:read`: - Read tools, for example `jobs_list`, `assets_get`, `catalog_list`, and `crawl_scrape`, work. - Write tools, for example `jobs_cancel` or `assets_create`, return `insufficient_scope`. - Paid tools, for example `generate_image` or `avatars_create`, return `insufficient_scope`. To run write or paid MCP tools, grant `mcp:write` on consent. As an alternative, use an API-key remote MCP session or the Developer API / CLI. Refer to [OAuth and API keys](/mcp/oauth) and [Tools and gates](/mcp/tools-and-gates). #### Local CLI alternative If the agent can run shell commands on your machine, direct CLI commands are the preferred method today (`sume login`, `sume tools list --json`, Avatar submit + `sume jobs …`). Refer to the [CLI overview](/cli). ```bash sume login sume mcp doctor --json ``` Local `sume mcp` is **not launched** in the current CLI releases (`coming_soon`). It is a future surface, separate from the hosted connector at `mcp.sume.com`. When the client supports remote HTTP MCP, hosted MCP is the preferred method. #### Studio Agent is different Studio Agent is the in-app Sume agent product. It is **not** the public hosted MCP connector that these pages document. Do not configure Studio Agent internals as a customer MCP server URL. ### Tools and gates Source: https://docs.sume.com/mcp/tools-and-gates.md Hosted MCP tool inventory, safety gates, and agent playbooks. Hosted MCP tools wrap selected public API capabilities. Always use `tools_list` and `tools_schema` to discover the live contract. Do not assume parity with the HTTP API. The live tool ids are **underscore** names from `packages/mcp-server/src/mcp.ts` (`remoteMcpTools`). On each call, the server canonicalizes `.` → `_`. Thus, dotted aliases, for example `tools.list`, still work. Retired aliases: `image-generations_create` → `generate_image`, `video-router_create` → `generate_video`. #### Discover tools | Tool | Purpose | |---|---| | `tools_list` | List all the tools that this session can see, and their safety metadata. | | `tools_schema` | Get one tool contract by `name`. | | `mcp_health` | Endpoint readiness, auth source, and safety posture. | Example agent instruction: ```text Call tools_schema with name "generate_image" and explain idempotency_key and dry_run before submitting any paid generation. ``` #### Safety gates Under OAuth `mcp:read`, hosted MCP gives read-only **visibility** by default. Hosted MCP hides the tools that change data and the paid tools until the session has `mcp:write` (or an API key). Spend is wallet/admission. There is no `mcp:paid` scope. | Gate | Required? | Meaning | |---|---|---| | `idempotency_key` | **Required** on write and paid tools | Stable key for transport/dedup, not human approval. | | `dry_run=true` | Optional | Admission/cost preview only. It does not submit the job. | | `max_spend_usd` | Optional | Sume enforces it only when you provide it. | | `allow_write` / `allow_paid` | Optional (legacy) | Sume accepts them for back-compat. They are **not** required. They cannot bypass a missing `mcp:write` scope. | Before expensive bursts, we recommend `generation_admission_preview` and/or `dry_run`. A normal single create does not need these admission steps. ##### Auth interaction | Session auth | What you see / can call | |---|---| | OAuth `mcp:read` only | Read-only tools. Calls to write/paid tools return `insufficient_scope`. | | OAuth `mcp:read` + `mcp:write` | Full hosted tool set. Paid submits still need `idempotency_key` and wallet/admission. | | API key | Full hosted tool set. Same `idempotency_key` / admission rules. | #### Programmatic tool calling (`script_run`) `script_run` runs a short JavaScript program on the Sume side. The program calls the tools below in a loop, in parallel, or with conditions, and returns one value. Use it when a turn needs three or more independent calls of the same shape (one `tts_create` per sentence, one `generate_image` per scene). Inside the script, `await sume.call(name, arguments)` runs any listed tool. The call has the same gates, redaction and errors as a direct call. Paid creates still need their own `idempotency_key`. `timeout_seconds` (5–55), `max_calls` and `max_paid_calls` set the limits of the run. The response contains the returned value, a `calls[]` journal and the child `jobs[]` to use with `jobs_wait`. A script cannot call discovery tools or `script_run` itself. #### Tool inventory (hosted) This list groups the tools from the current hosted registry. The names are live tool ids. To get the subset that the session can see, call `tools_list`. ##### Meta and health - `mcp_health` - `tools_list` - `tools_schema` - `script_run` (programmatic tool calling, refer to the section above) - `health_service` - `health_v1` ##### Account and catalog - `account_me` - `balance_get` - `usage_get` - `catalog_list` - `image-models_list` / `image-models_get` - `video-router_models` - `generation_admission_preview` ##### Jobs Read: `jobs_list`, `jobs_get`, `jobs_status`, `jobs_result`, `jobs_events`, `jobs_wait`. Write (`idempotency_key`): `jobs_cancel`. ##### Assets Read: `assets_list`, `assets_get`, `assets_download_url`. Write (`idempotency_key`): `assets_create`, `assets_upload_url`, `assets_complete`. Hosted MCP cannot read files from your laptop. The upload flow is: create upload URL → client PUT bytes → `assets_complete`. ##### Image, video, audio generation Paid (`idempotency_key`, and omit `payload.model` to route to `sume/auto` unless the user named a family): - `generate_image` - `generate_video` - `music_create` - `tts_create` - `stt_create` - `image_upscale_create` - `rmbg_create` - `video_upscale_create` - `kling-motion-control_create` Read (free): `tts_source_get` (accepted-script manifest for `tts_create` with `transcript_source`), `tts_source_verify_spine` (compares the selected TTS jobs with the accepted script). ##### Avatars and talking-head Read: `avatars_list`, `avatars_get`, `avatars_search`, `avatar-videos_list`, `avatar-videos_get`. Paid: `avatars_create`, `avatar-videos_create`, `avatar-image-to-video_create`, `avatar-video-previews_create` / `_get` / `_regenerate` / `_generate_video`. ##### Crawl (web + social) Read: `crawl_scrape`, `crawl_map`, `crawl_search`, `crawl_get`, `crawl_profile`, `crawl_feed`, `crawl_media`, `crawl_find`. Write (`idempotency_key`, unbilled utility): `crawl_site` (then `jobs_wait` → `crawl_get` on the same id). Social discovery skill: `crawl-social`. Web research skill: `crawl-web`. ##### Media inspect / import / captions / timeline - `media-imports_create` / `media-imports_get` - `video_inspect` (default for “what is this clip”, dest and prod) — [Video inspect](/models/video-inspect) - `video_frames_create` / `video_frames_get` — [Video frames](/models/video-frames) - `video_trim` — [Video trim](/models/video-trim) - `audio_detach` — [Audio detach](/models/audio-detach) - `video_filter` — [Video filter](/models/video-filter) (`check_only: true` is the unbilled `/check`) - `video-captions_get` (the legacy `video-captions_create` / `video-caption-overlay_create` are unlisted. In-flight clients can still call them by name, but they are not in `tools_list`) - `timeline_create` / `timeline_get` — [Timeline 1.0](/models/timeline) - `timeline_compose` — [Timeline compose](/models/timeline-compose) - `timeline_audio` — [Timeline audio](/models/timeline-audio) - `trending-videos_search`, `trending-research_search` `video-analyses_*` (#5953): dest (`SUME_COM_VIDEO_ANALYSIS_ENABLED=false`) delists them from `tools_list` and `video-analyses_create` answers `410 video_analysis_retired`. Production still lists them until PR-C2. Do not call create on dest. Dest-only (`mcp.dev.sume.com` / `api.dev.sume.com`, never production): `video_analyze` and `video_segment` appear in `tools_list` only when `videoUnderstand.enabled` is on (Railway `development` + trusted origin `https://api.dev.sume.com` + `SUME_COM_TWELVELABS_API_KEY`). Both must have `idempotency_key` and `max_spend_usd`. They do not replace `video_inspect` for lightweight probe/stills. #### Not on hosted MCP These names are **not** in `tools_list`: - `images_create` / `videos_create` — Sume Image 1.0 and Video 1.0 stay REST-only. Use the [Developer API](/public-api). - Higgsfield-only names (`get_workflow_instructions`, `models_explore`, `media_import_url`, `remove_background` as an HF tool). Cutouts are `rmbg_create`. The social URL mirror is `media-imports_create`. `catalog_list` can still show HTTP capabilities that do not have an MCP tool. #### Playbooks ##### Playbook A — OAuth read-only discovery (Cursor / Claude) 1. Use OAuth to connect to `https://mcp.sume.com/mcp`. If you do not need to change data, keep Write off. 2. Call `mcp_health`. Make sure that `authenticated.auth_source` is `mcp_oauth`. 3. Call `tools_list`. When Write is off, only the `read_only` tools are available. 4. If necessary, call `catalog_list`, `balance_get`, and `jobs_list`. 5. If Write was off, stop before you call tools that change data. These tools return `insufficient_scope`. ##### Playbook B — Inspect one tool before paying 1. Call `tools_schema` with `name: "generate_image"` (or `avatars_create`). 2. Call `generation_admission_preview` or the paid tool with `dry_run=true`. 3. Examine the estimate, the balance, and the queue behavior. 4. Submit with a new `idempotency_key` on a session that has `mcp:write` or an API key. If you want a cap, add the optional `max_spend_usd`. ##### Playbook C — Paid avatar create (write session) Use this playbook only when the user explicitly confirms spend. ```json { "idempotency_key": "avatar-create-2026-07-21-001", "dry_run": true, "max_spend_usd": 2, "payload": { "avatar_handle": "studio_presenter", "input": { "type": "prompt", "prompt": "A friendly studio presenter in neutral lighting" } } } ``` You can still send `allow_write` / `allow_paid`. They are not required. 1. First, call with `dry_run=true`. Then examine the preview. 2. To submit, call again with `dry_run` omitted or `false`. 3. Poll with `jobs_status` / `jobs_wait`. Then read `jobs_result`. 4. In agent reports, Sume public ids and `media.sume.com` URLs are preferred. Do not paste signed URLs, OAuth tokens, or API keys into chat logs. ##### Playbook D — Local CLI instead of hosted MCP ```bash sume login sume mcp doctor --json sume tools list --json ``` Use this playbook when the agent already runs local shell commands. CLI tool ids stay dotted (`avatars.create`). The CLI registry is not the hosted MCP catalog. In the current CLI releases, local `sume mcp` is still `coming_soon`. Hosted OAuth and local CLI login are different flows. `sume login` does not mint hosted MCP OAuth tokens. #### Related - [MCP quickstart](/mcp/quickstart) - [OAuth and API keys](/mcp/oauth) - [Generation admission](/workflows/generation-admission) - [Jobs and results](/workflows/jobs-and-results) ### OAuth and API keys Source: https://docs.sume.com/mcp/oauth.md Hosted MCP authentication matrix and mcp:read / mcp:write scopes. Hosted Sume MCP accepts OAuth access tokens or Sume API keys. You cannot use one of these credentials in place of the other. #### Auth matrix | Mode | How you connect | Hosted capability today | |---|---|---| | OAuth | The client uses MCP OAuth / protected-resource metadata and first-party consent on the **MCP host** | `mcp:read` is required and read-only. To also grant `mcp:write`, toggle Write on the consent page. There is **no** `mcp:paid` scope. | | API key | The client sends `Authorization: Bearer ` or `x-api-key` | Full hosted tool set. Spend is wallet/admission. Writes and paid calls must include `idempotency_key`. | | Local `sume mcp` | Future CLI MCP after `sume login` or local key | **Not launched** yet (`sume mcp doctor` → `coming_soon`). It is separate from hosted OAuth. | #### OAuth flow (hosted) 1. The MCP client connects to `https://mcp.sume.com/mcp`. 2. Sume returns an OAuth challenge and protected-resource metadata (`authorization_servers` is the MCP origin, not `www` / `app.sume.com`). 3. The client sends the user to `https://mcp.sume.com/oauth/authorize`. This URL redirects to the first-party consent page `GET /oauth/consent` on the MCP host (Clerk browser JS on that origin). 4. After sign-in, the consent page shows **Permissions**: Read is locked on, and the Write toggle is off by default. Continue posts to `POST /oauth/consent/decision`. 5. The client exchanges the authorization code (PKCE) for an access token. 6. The client calls `https://mcp.sume.com/mcp` with that bearer token. `www.sume.com` is still a secondary/deprecated authorization-server surface. The protected-resource metadata does not advertise it now. Do not send interactive clients to `app.sume.com` for MCP OAuth. Useful public metadata endpoints: ```text https://mcp.sume.com/.well-known/oauth-protected-resource/mcp https://mcp.sume.com/.well-known/oauth-authorization-server ``` OAuth resource audience: ```text https://mcp.sume.com/mcp ``` Development uses the same shape on `https://mcp.dev.sume.com`. #### Scopes Already shipped (`packages/mcp-oauth` + `packages/mcp-server/src/mcp-oauth-as.ts`): - Supported scopes: `mcp:read` (required) and `mcp:write` (opt-in). A grant of write always includes read. - There is **no** `mcp:paid` OAuth scope. Paid submits are wallet/admission. - `mcp:read` sessions see only **read-only** tools. If a session without write calls a tool that changes data, the call returns `insufficient_scope`. - `mcp:write` sessions see the tools that change data and the paid tools. - Paid/write submits must include `idempotency_key` (transport/dedup). The optional `dry_run` preflights cost. Sume enforces the optional `max_spend_usd` only when you provide it. - Sume accepts the legacy `allow_write` / `allow_paid` for back-compat. They are **not** required. They cannot bypass a missing `mcp:write` scope. API-key remote MCP is the other path for automation that does not use OAuth. #### API-key remote MCP API-key compatibility is still available for current users and automation. Send one of these: ```bash # Bearer Authorization: Bearer $SUME_API_KEY # or header x-api-key: $SUME_API_KEY ``` API-key sessions can see write and paid tools. Calls that change data and paid calls must still include `idempotency_key`. Before the first paid submit, we recommend `dry_run=true` or `generation_admission_preview`. If you pass the optional `max_spend_usd`, it sets the maximum spend. Create keys in the dashboard: [API keys](/dashboard/api-keys). #### Credential safety - An MCP OAuth token is **not** a Sume API key. - Do not store OAuth tokens in CLI config. Do not paste them into prompts. Do not forward them to third-party providers. - Do not mint API keys for hosted OAuth clients as a workaround. - `sume login` does **not** broker hosted MCP OAuth tokens. - If API keys appear in logs or chat history, rotate the keys. #### Hosted MCP vs local MCP vs Studio Agent | Question | Answer | |---|---| | Best interactive connector for Cursor/Claude? | Hosted MCP + OAuth at `https://mcp.sume.com/mcp`. | | Best for local shell agents already on the CLI? | Direct [CLI](/cli) commands after `sume login` (local `sume mcp` is not launched yet). | | Need Image 1.0 / Video 1.0 (`images_create` / `videos_create`)? | Not on hosted MCP. Use the Developer API. Router stills/clips are `generate_image` / `generate_video`. | | Is Studio Agent the same as hosted MCP? | No. Studio Agent is a separate product surface. | #### Related - [MCP overview](/mcp) - [MCP quickstart](/mcp/quickstart) - [Tools and gates](/mcp/tools-and-gates) - [CLI overview](/cli) - [Authentication](/authentication)