Agent Completions
An Agent Completion runs the Sume Agent on an ad-hoc prompt. It uses the same runtime as the Agents chat UI: full sandbox, tools, MCP bridge, and media generation. Your own backend can use an API key to start it, and no person has to monitor it.
A schedule stores what to do. A Format stores how to do it. An Agent Completion stores nothing. You send the task on each call.
Which one do I want?
| Surface | Use it when | Start with |
|---|---|---|
| Format | You have a saved packaged workflow and only the inputs change. | POST /v1/formats/{handle}/{slug}/runs |
| Scheduled | You want a schedule or a trigger on a saved automation. | POST /v1/actions/{handle}/{slug}/runs |
| Agent Completions | The task changes on each call. You only want the Agent to do this one task. | POST /v1/agent/completions |
All three run the same agent and return the same receipt shape. The only differences are the source of the instruction and what Sume saved for you.
Async, not a chat drop-in
An Agent Completion is not a synchronous chat completion. A real agent turn opens a sandbox,
calls tools, and can generate media. This work takes too long to keep one HTTP request open.
Thus, the create call returns 202 and a receipt, and you poll the receipt.
The request uses the OpenAI messages[] shape, so your current integration code fits. But the response is
a run receipt, not choices[]. Streaming and a sync OpenAI-compatible wire are not available at
this time.
Scopes
| Scope | Needed for |
|---|---|
agent_completions:read | Read and list runs. |
agent_completions:write | Create a completion, cancel a run. |
Keys that you created before Agent Completions shipped do not have these scopes. Each
request with an older key fails with 403 insufficient_scope. You cannot add scopes to a key that
already exists. Create a new key at API Keys. Then
rotate to the new key. Refer to Authentication.
Service-account keys cannot create Agent Completions. A request with a service-account key fails
with 403 insufficient_scope and details.reason of
service_account_agent_completions_unsupported.
Create a completion
To rewrite the example, put your values in the required fields. generation_spend_cap_usd has
no default. If you do not send it, the request fails.
Create an Agent Completion
POST /v1/agent/completions
Required
If Sume accepts the completion, it returns 202 and a receipt:
Request fields
| Field | Required | Notes |
|---|---|---|
instruction | one of | The task, as a plain string. |
messages | one of | system and user turns. Send exactly one of instruction or messages. Do not send both. |
generation_spend_cap_usd | yes | The maximum generation spend on this run. Refer to the section below. |
model | no | Only sume-agent. If you do not send it, you get the same agent. |
input | no | Caller data. Sume writes all of it to /workspace/inputs/sume-action-input.json. The prompt has a bounded pointer to the file. Sume uses this value only as data, not as instructions. |
attachments | no | A maximum of 30 images that the agent can see and use. Refer to Attachments. |
output_schema | no | Bind the output of the run to your own schema. The contract is the same as for Action runs. |
primary_output_key | no | The output key that holds the headline result. |
communication.webhook_url | no | Public HTTPS URL that Sume notifies when the run gets to a terminal status. Refer to Run webhooks. On api.dev.sume.com and api.sume.com, Sume accepts, stores, and delivers it. |
Idempotency-Key works the same as on Action runs. If you send a key again, the API returns the
original receipt with idempotency_hit: true. If you use the key again with a different payload,
the API returns 409 idempotency_conflict.
messages[]
content accepts a string or an OpenAI-style [{ "type": "text", "text": "..." }] array. Sume
joins the turns, in order, into one prompt.
content also accepts { "type": "input_text", "text": "..." } as an alias for text, and
{ "type": "input_image", ... } parts. Refer to Attachments.
The API rejects assistant turns. It does not ignore them. Acceptance of these turns implies
that Sume replays a prior conversation. This endpoint does not do that at this time. Each
completion runs in a new thread. The thread_id in the receipt identifies that thread.
Attachments
You can send images that the agent can actually look at. Send them at the top level or as
input_image content parts. Both forms use the same item shape. Sume merges the two sources into
one list.
A turn with only images is permitted. If you do not send the text part, Sume tells the agent to use the attached files.
You can use attachments and output_schema together. The images go to the agent. After the run
completes, Sume still parses the output of the run against your schema.
For the item shape, limits, upload path, and error codes, refer to Format runs. These items are the same on both surfaces.
The spend cap is required
generation_spend_cap_usd has no default. If you do not send it, the request fails with 400 invalid_request.
This rule is intentional. An Agent Completion is an unattended agent. It has tools and access to your generation wallet.
In the chat UI, an interactive spend-approval prompt protects you. A backend caller does not get this prompt. The cap replaces the prompt. Set the cap for each run to the maximum spend that you accept for that run.
To find the correct cap, read the metered rates that the run will use on the API pricing page.
Poll for the result
The statuses are the same as for Action runs: queued, processing, completed, failed,
canceled. Poll status_url until the value of next_action is not poll_status.
A completed run fills output. By default, this output has the sume/action-run-output/v1 shape.
The last text of the Agent is in output.text. Any generated media is in output.images,
output.videos, output.audio, and output.files. The run also fills artifacts, and it records
the spend in usage. Media URLs are durable media.sume.com HTTPS URLs.
To stop a run that is in progress:
GET /v1/agent-runs lists your completions, newest first.
Errors
| Status | Code | Cause |
|---|---|---|
400 | invalid_request | Missing generation_spend_cap_usd, neither or both of instruction/messages, an assistant turn, a malformed input, or a model other than sume-agent. |
400 | invalid_attachment | Bad attachment item: wrong type, missing or non-HTTPS URL, both image_url and asset_id, or a source that is not a permitted image. |
400 | attachment_not_found | asset_id is unknown in this workspace. |
413 | attachment_too_large | An image is more than 30 MB, or the total of the set is more than 500 MB. |
502 | attachment_fetch_failed | Sume could not fetch the image. Causes: unreachable host, hotlink protection, or a non-2xx response. |
403 | insufficient_scope | The key does not have agent_completions:*, or it is a service-account key. |
404 | agent_run_not_found | Unknown run id, or a run that another account owns. An Action or Format run id will not resolve here. |
409 | idempotency_conflict | You used the Idempotency-Key again with a different payload. |
Not available yet
- Non-image attachments. At this time,
input_imageis the onlytype. PDFs and other files will come later. - Streaming, and a synchronous OpenAI-compatible
choices[]response. - Continuation of a prior thread with
thread_id, andassistantturns inmessages[]. - Team-owned threads. Completions are user-owned.