Structured output
By default, a completed Format run gives you media and a paragraph of text. This result is good for a person, but it is not easy to use in a database. If you bind a schema, you get a typed object. It is the same run, but with a shape that you can write directly into your own records.
The exact request and response schemas come from the live OpenAPI
(https://api.sume.com/reference/json). The tables on this page are a summary that is easy to
read. They are not a second schema.
input
A run request has two JSON-shaped fields, and they operate in very different ways. Confusion between the two fields is the most frequent cause of errors in a first integration.
input | output_schema | |
|---|---|---|
| What it is | Caller data | A contract for the receipt |
| Direction | You → the run | The run → you |
| Shape | Any JSON object that your backend needs | JSON Schema, inside the supported subset |
| Checked for | Object type, key count, byte size | Every rule in the subset |
| A shape Sume does not expect | The run starts. Unknown keys are only more data. | 400 output_schema_invalid. Nothing runs, and Sume charges nothing. |
| Where it lands | The agent's prompt, as a fenced data block | The post-run projection |
Thus, input is a flexible concatenation, and output_schema is a strict typed receipt. You
can make your input as loose as you want. You cannot make your output_schema loose at all.
strict: false does not make it loose, and there is no other way to bypass the rules.
The two fields do not interact. The projection never sees your input (refer to the
sections below). Thus, a value that you sent cannot come back in output, unless the run
repeats it in its final text.
Where your object comes from
This part is different from a chat completion. Make sure that you understand it before you design a schema. There are two paths, and the receipt tells you which path you got.
filled_by: "agent" — the run answered. Sume gives your schema to the run as a tool that
the run must call before it stops. Your schema is the argument shape of that tool. The model
that does the work is also the model that fills your object. It fills the object while it still
knows what it made and why. It can see your input and your instruction, because they are
part of the run.
filled_by: "projection" — the fallback. If the run stops and did not submit a valid
object, a separate constrained pass builds one after the run. The pass uses the data that the
run left. This pass is an OpenAI-strict json_schema completion at temperature 0. It gets only
two facts:
| Fact | Detail |
|---|---|
| The run's generated media | All the artifacts that the run produced, with their durable URLs and metadata. |
| The run's closing text | The final assistant message, truncated to the first 8000 characters. |
The pass does not get your input, not your instruction, not the Format body, and
not the intermediate steps of the run. On this path, all data that you want in output
must be in one of these two facts.
Thus, it is important to read filled_by. It shows if the run wrote the object, or if a pass
built the object again from the data that the run left. It explains most unexpected results:
- On the projection path, your own identifiers do not round-trip. The projection cannot see an
order_idthat you sent ininput. Keep your identifiers on your side, withrun.idor theIdempotency-Keythat you sent as the key. Useoutputonly for the data that the run made. - Nothing in
outputis invented, on either path. The same checks gate both paths before the object gets to you (refer to the sections below). An object that the run wrote does not get more trust. - On the projection path, an empty text field is a signal. The projection never sees your
input. Thus, if the titles, descriptions or ids in a schema come from the brief, these fields come backnull. At the same time, the media fields are full.filled_by: "projection"with null prose shows a run that stopped early. It does not show a Format that did not write copy.
If you score runs automatically (a smoke matrix, a partner integration, a dashboard), read
filled_by before you count a run as delivered. For video, read the file, not the number next
to it. ffprobe on primary_output_url takes two seconds. It is the only method that shows the
difference between an assembled cut and a clip that has the shape of one.
The practical rule that comes from this is the same: require only what the Format actually
makes. If a schema requires a field that the recipe never produces, the schema will fail on
the projection path each time. The failure is silent until you read output_error.
Bind a schema
There are two spellings, with one behavior. Use the spelling that is applicable to your client.
Native — output_schema:
OpenAI-shaped alias — response_format:
Sume normalizes response_format into output_schema on the receipt. Thus, the receipt
always shows the native spelling. If you send both, the result is 400 invalid_request.
The alias uses the Chat Completions spelling: type at the top, and the binding nested
under json_schema. The OpenAI Responses API flattens the same fields into
text.format: { type, name, strict, schema }. Sume does not accept this flattened shape.
If your client builds Responses-style bodies, move the fields into output_schema, or nest
them again under json_schema.
| Field | Rules |
|---|---|
name | Required. 1–64 characters, ^[A-Za-z0-9._/-]+$. Give it a namespace, because it shows on every receipt. |
strict | The default is true. Refer to the note under Supported schemas. |
schema | Required. A JSON Schema object inside the supported subset. |
A Format can also have its own default schema, which you bind in the dashboard. A per-request
output_schema overrides it for that run. The receipt shows which schema applied, in
output_schema.source:
source | Meaning |
|---|---|
default | No schema is bound. output is the built-in schema. |
action_default | The Format's own bound schema. (action_ is the wire spelling. Scheduled uses the same spelling.) |
request_override | The output_schema that you sent on this run request. |
Supported schemas
Schemas must agree with the OpenAI strict-mode subset. This subset is a requirement, not a
recommendation. If a schema is outside the subset, Sume rejects it at submit with
400 output_schema_invalid. The response has a details.violations[] array that names each
problem. Nothing runs. Thus, Sume does not charge you.
The supported keywords, exactly
The subset uses an allowlist. A keyword that is not on this list is a violation. Sume does not ignore it silently. An ignored constraint is a part of the schema that Sume cannot promise to satisfy.
| Group | Accepted |
|---|---|
| Structure | type, properties, required, additionalProperties, items, $defs, $ref, anyOf |
| Values | enum, const |
| Strings | format, pattern, minLength, maxLength |
| Numbers | minimum, maximum, exclusiveMinimum, exclusiveMaximum, multipleOf |
| Arrays | minItems, maxItems |
| Annotation | title, description, default, examples, $schema, $id |
Types are string, number, integer, boolean, object, array, null.
Sume rejects all other keywords. These are the keywords that cause problems for real integrations:
| Rejected | Instead |
|---|---|
oneOf | anyOf. Only anyOf is on the list. Schemas ported from OpenAPI usually use oneOf. |
allOf | Flatten the branches into one object. |
not, if / then / else, dependentRequired, dependentSchemas | You cannot express these. Model the alternatives as anyOf, or validate on your side after you read output. |
nullable: true | A nullable union: "type": ["string", "null"]. |
patternProperties, propertyNames, unevaluatedProperties, additionalItems | Declare the properties that you want. additionalProperties: false covers the rest. |
type
A node can also have a $ref or an anyOf in its place. A bare { "description": "…" } is
legal JSON Schema, and it means "anything". But here it is a missing_type violation. An
array must also declare items.
$ref and anyOf each short-circuit the node that contains them. Sume checks the sibling
keywords next to them against the allowlist, but these keywords have no other effect. Put
constraints inside the anyOf branches or inside the $defs entry, not next to the $ref.
The root must be an object
The type of the root must be "object", as a single type. Thus, Sume rejects even
{ "type": ["object", "null"] }. Sume also rejects a top-level array, string, or union. Wrap
it:
additionalProperties: false
This rule is applicable to all objects in the schema, not only to the root. This includes the
objects nested inside array items and inside $defs.
required
There is no optional property. Each property that you declare must be present.
Use a nullable union to express optionality:
This rule causes problems for more ported schemas than any other rule. Read null as "the
run had nothing to put here". This is the same case for which you wanted optional.
Size and nesting limits
| Limit | Value | Violation |
|---|---|---|
| Nesting depth | 10 levels | max_depth |
| Total properties | 5000, counted across the whole document | max_properties |
| Enum values | 1000 per enum | max_enum_values |
| Total string length | 120,000 characters, summed over every property name, key, and string value in the document | max_string_length |
The last limit is a document-wide budget, not a per-field cap. Thus, long description
annotations on a large schema can use all of it, even when no single string is very long.
is limited
Only two targets resolve:
| Target | Use |
|---|---|
#/$defs/* | Your own definitions, declared at the root of the schema document. |
SumeMediaFile# | Sume's media shape. Refer to the section below. |
Sume rejects an external $ref (a URL, a sibling document, #/components/...). Sume also
rejects $ref: "#". OpenAI strict mode permits root recursion in this way, but Sume does not.
Sume also rejects a #/$defs/* target that has no root $defs entry of that name. This rule
finds a $defs block that is nested inside a sub-schema and not declared at the root.
Recursion through a named definition is permitted. A $defs entry can $ref itself. The depth
limit counts literal nesting in the document. Thus, a self-referential definition does not
increase the depth count.
does not relax any of this
Sume accepts and stores it, but it does not change the subset above. If a schema is outside
the subset, Sume rejects it, whether strict is true or false. Do not use it to bypass the
subset. There is no way to bypass it.
There is also no equivalent of the OpenAI JSON mode ({"type": "json_object"}), the loose
"valid JSON, any shape" option. Bind a schema or use the built-in
one. These are the only two options.
details.violations[]
Each entry is { path, rule, message }. path is a JSON-Pointer-style location, for example
#/properties/scenes/items/properties/clip. message is text for a person, and it can change.
rule is a stable lowercase token that you can safely switch on:
rule | Meaning |
|---|---|
root_must_be_object | The root is missing, is not an object, or its type is not exactly "object". |
not_an_object | A schema node is not a JSON object. |
missing_type | A node has no type, $ref, or anyOf. |
unsupported_type | A type that is not one of the seven types above. |
unsupported_keyword | A keyword that is not on the allowlist. |
additional_properties_false | An object node without additionalProperties: false. |
required_completeness | A declared property that is not in required, or a required entry that has no property of that name. |
missing_items | An array node with no items. |
unsupported_ref | A $ref that is neither #/$defs/<name> nor SumeMediaFile#, or one that names a definition that does not exist. |
invalid_defs | $defs is present but is not an object of named schemas. |
max_depth, max_properties, max_enum_values, max_string_length | The limits above. |
Sume reports all the problems, not only the first. Thus, one 400 gives you sufficient data to fix the schema.
Coming from OpenAI structured outputs
If you used response_format: { type: "json_schema", … } or the Responses API's
text.format, most of what you know is applicable here. The subset rules are the same, and
they come from the same guide. The difference is where Sume applies the schema.
| OpenAI | Sume | Note |
|---|---|---|
response_format.json_schema | output_schema, or response_format verbatim | Chat Completions spelling only. Sume does not accept text.format. |
json_schema.name | output_schema.name | Required in both places. Give it a namespace, because it shows on every receipt. |
json_schema.strict | output_schema.strict | Accepted, the default is true, and it changes nothing. Sume always enforces the subset. |
{"type": "json_object"} (JSON mode) | (no equivalent) | Bind a schema, or use the built-in one. |
| The model emits the JSON | A post-run projection emits it | The schema constrains the projection, never the run. |
refusal on the message | output_error on the receipt | Different mechanism: not a safety refusal but a failed projection. |
incomplete_details.reason: "max_output_tokens" | (not applicable) | The projection is small and bounded. There is no truncated-JSON case to handle. |
| Streamed partial JSON | (not applicable) | output appears once, on the terminal receipt. |
$ref: "#" root recursion | Rejected | Recurse through a named #/$defs/* entry instead. |
| Nothing comparable | The URL gate | Sume checks each URL in output against the media that the run really produced. |
| Nothing comparable | SumeMediaFile# | A built-in $ref target for the run's media. |
The change in how you think is this: with OpenAI, you constrain what the model says. Here,
you constrain how Sume reads back a completed run. All other differences come from this. It
is why the schema cannot make the Format produce a video, and why a hallucinated URL cannot get
into output. It is also why a run can give you output: null.
SumeMediaFile
Use { "$ref": "SumeMediaFile#" } at each location where you want a piece of the media of the
run in your output. All fields are required. All fields except type and url are nullable.
| Field | Type | Notes |
|---|---|---|
type | "image" | "video" | "audio" | "file" | |
url | string (uri) | Must be a URL that this run actually produced. Refer to the URL gate. |
content_type | string | null | For example, video/mp4. |
file_name | string | null | |
size_bytes | integer | null | |
width, height | integer | null | Images and video. |
duration_ms | integer | null | Video and audio. |
expires_at | string (date-time) | null | null for durable media.sume.com URLs, which is the normal case. It has a value only when Sume returns a signed URL. |
Sume serves Sume-hosted media from media.sume.com as public, max-age=31536000, immutable,
and the media does not expire. Store the URL with your own record and render it later. It is
not necessary to refresh it. A durable URL is also a public URL. If this is important to your
product, refer to
Map artifacts into your UI.
The built-in schema
If you do not bind a schema, Sume projects output onto sume/action-run-output/v1:
text is nullable. The four arrays are always present, and they can be empty.
Sume fills the built-in schema deterministically from the generated media and the final text of the run. Sume does not use a model. Thus, it cannot fail in the way that a custom schema can fail. If you want the media but do not want to design a schema, this schema is sufficient to ship with.
The URL gate
Before a custom output gets to you, Sume checks each URL in it against the set of media that
this run actually produced. The comparison is exact string equality.
Two passes find these URLs. The first pass collects each http(s):// string in the object, at
any depth. It does this whether or not you declared the field as media. The second pass
collects the url of each SumeMediaFile-shaped value, whatever it
contains. Thus, a placeholder such as "none" or "" in a url cannot get past the gate
only because it does not look like a URL.
If this run did not produce a URL, the projection fails, and Sume does not return the URL. This
is also true for a well-formed media.sume.com URL that looks correct. Then Sume validates the
object against your schema.
Thus, a completed run returns output that agrees with your schema and has real media URLs. Or,
it returns output: null and tells you why. It never returns a schema-shaped guess.
The deliverable is not one of its parts
A schema that describes a whole made of parts (scenes, shots, segments) usually also has a field for the completed, assembled file. These are different files. Sume rejects a receipt that fills the whole with one of its own parts.
The check is narrow intentionally. It needs two or more parts that report
"status": "succeeded" with their own video. It also needs a video outside all parts that
uses one of those files again. The check does not affect a Format whose single clip really is
the deliverable. It also does not affect a poster or thumbnail that the run intentionally
reused from a part.
A run that could not assemble the file has an honest answer available. That answer is not "here is a scene". Report the parts that you made and leave the assembled field null. Or, report the real status of the assembly.
Durations are checked against the file
A duration_ms inside a SumeMediaFile comes from the same ledger that fills
artifacts[]. If that ledger recorded a length, the value in output must agree with it within
10%. A value that does not agree describes a different file. It fails the projection and does
not get to you.
If the ledger did not record a length, Sume does not check the value. In this case, null
means "not measured", not "zero". Sume does not fail a run because of a fact that nobody
measured. Thus, a duration_ms that you read back is the value of the artifact, or it is not
verified. It is never a number calculated from other data.
What the gate does not check
The gate checks only URLs, durations, and the shape. The projection pass reads all other values
(ids, labels, captions, counts) from the media metadata and the closing text of the run. These
values come from what the run reported, but Sume does not verify them against the run. Thus, treat
these fields as the report of the run about its work, not as measurements. The projection
writes a duration_seconds number that you declared yourself. The duration_ms on the media
file is the value that Sume checks.
The gate also admits only media that the run generated. A file that the run only uploaded
is not in that set. Thus, if a schema field holds the URL of an upload, the projection fails,
and all of the output fails with it. Do not put uploads in your schema.
and
Most integrations have one item to show. If you name it, the receipt resolves it for you:
Then primary_output_url on the receipt is the URL at that key. The resolution order is:
- The
primary_output_keyon the run request. - The Format's own
primary_output_key. - The first top-level key that holds a media object, or an array whose first element is a media object.
For the built-in schema, the fallback order is videos → images → audio → files.
If you named a key in step 1 or 2, the receipt echoes it back when output has a value under
it. This includes a value that is not media. If you point at a headline string, you get
primary_output_key: "headline" with primary_output_url: null, because there is no URL to
resolve. The resolution goes to step 3 only when the key is missing from output, or empty
there.
primary_output_key has a maximum of 64 characters. Both fields are null on any non-terminal
status. They are also null when output_error is set.
When output cannot be produced
Before you read output, examine output_error.
On a run over the API, a projection failure is a run failure. These runs are unattended,
as Call a Format states. A completed
receipt whose output is null looks like a success, but it is not one. Thus, status is
failed, and error has the same reason as output_error. In the Agents UI, a person reads
the thread. There, the same shape stays completed, because it is a draft, not a receipt.
A failed run still publishes what it produced. output is null when nothing satisfied
your schema, not only because the run failed. For example, a live-commerce show rendered 20 of
40 scenes and then stopped. The show reports those 20 on output. The other scenes have the
"failed" value that your own schema defines. The output_error next to them tells why the run
stopped.
A failure never gets a pointer: primary_output_key and primary_output_url are null on
every non-completed run. Thus, if (run.primary_output_url) stays a safe test for "the
deliverable exists", and a partial result cannot make it give a wrong answer.
For this to work, your schema must permit the partial result. Refer to Make a partial result legal.
artifacts[] is populated either way, with all the items that the run made. You always
have the media, even when the shape failed.
The exception is a projection that did not get access to the run. output_extraction_failed
records a transport failure, not a verdict. Thus, the run stays completed, and Sume projects
it again automatically on your next read.
output_error.code | Meaning | details |
|---|---|---|
output_schema_unsatisfied | The projection did not agree with your schema, or it referenced media that this run did not produce. | rejected_urls[] (first 10 only) or violations[], and a harvested count by media type. |
output_extraction_failed | The projection could not run. status stays completed. A reason of harvest_unavailable means that Sume could not read the media of the run while the run finalized. The receipt fills in on the next read. | reason |
unattended_blocked | The run stopped and did not claim a deliverable that it did not produce. message is the reason in the words of the run. | A harvested count by media type, and assembled_deliverable: false. The harvested media are intermediates, not the completed cut. |
deliverable_missing | The Format produces media (io.output_kind) that the run never made. Thus, Sume did not let any structured output claim it. | declared_output_kind, and a harvested count by media type. |
primary_output_missing | The run satisfied your schema, but the primary_output_key that you declared has no value. output still has the partial result. The run is failed because the item that you named is not in it. | primary_output_key |
agent_reported_failure | The accepted receipt of the run says that the run did not deliver: an explicit-fail payload, failed / stand-in media slots, or a primary of the wrong media type for the io.output_kind of the Format. output still has the ledger. The run is failed because the receipt says so. | reason (explicit_failed, slots_failed, primary_not_deliverable), non_delivered_slots[] or primary_media_type, and a harvested count by media type. |
Treat this set as open, because new codes can appear. Branch on the codes that you handle, and use a default path for the rest. Do not use an exhaustive switch.
Handle each case as this list shows:
output_schema_unsatisfied, repeatedly, on the same Format. The cause is almost always a schema that requires media that the recipe does not generate. Comparedetails.harvestedwith your required fields. The example above requires an image and got one, but it wanted a second image. Make the field a nullable union, or change theinstructionso that the run makes the media.output_schema_unsatisfiedwithviolations[]. The problem is the shape, not the media. The violations name the paths that cause the problem.output_extraction_failed. This failure is transient. Before you do anything else, read the run one more time. This step alone clears aharvest_unavailable. If the failure continues, retry the run with a new idempotency key. The old key is bound to the receipt that you already have.- Any of them, in your UI. You still have
artifacts[]. Show the media and log the shape failure. This is better than an error for a customer whose video exists.
Make a partial result legal
A run can produce some of its output and then fail. The run can report this only if your
schema says a partial is a legal shape. The platform does not add a minimum of its own. Ajv
enforces the keywords that you wrote, and only those keywords. Thus, if a schema requires every
scene, a 20-of-40 show returns output: null.
Two rules do all the work. Both rules are not intuitive under the strict subset:
- Optional means nullable, not absent. You must still list each property in
required(refer to Every property must be listed inrequired). To show that a value can be missing, use"type": ["string", "null"]. - Do not use
minItemson the arrays that you want to receive partially. Sume really enforces it. Thus,minItems: 1on a scene list will reject the ledger that you want to read.
Use it together with "primary_output_key": "full_video". This keeps the loose schema honest.
A run that fills scenes but leaves full_video null satisfied the schema, but it did not
produce the deliverable. Thus, the run terminalizes as failed with primary_output_missing,
and it does not report a success. The loose schema lets you see the partial result. It does not
let the run pass.
When you read one of these receipts, branch on full_video for "did I get a show". Branch on
each scenes[].status for "what do I need to retry". To retry, send the id of the failed run
as previous_run_id on a new run (refer to Continue a run).
The next turn continues the same conversation, and those clips are already in it.
Failures at submit
Sume finds these schema problems before anything runs. Sume does not charge you.
| Code | Status | What to do |
|---|---|---|
output_schema_invalid | 400 | Your schema is outside the supported subset. details.violations[] names each problem. |
invalid_request | 400 | This includes a request that sends both output_schema and response_format. |
Checklist
- The root is
{"type": "object"}only.additionalProperties: falseis on every object, also in$defs. - Every node has a
type, a$ref, or ananyOf. Everyarraydeclaresitems. - Every declared property is in
required. Optionality is a nullable union. - No
oneOf,allOf,not,if/then/else, ornullable: true. -
$reftargets are only#/$defs/*(declared at the root) andSumeMediaFile#, and never#. - The schema requires only media that the Format actually produces.
- No field expects a value that you sent in
input, because the projection cannot see it. -
nameis namespaced and stable. Thus, you can grep the receipts. -
primary_output_keynames the one item that your UI shows. - Your reader handles
output: nullwithoutput_errorset, on afailedrun and on acompletedone. - When
outputis null, your reader usesartifacts[]as a fallback.
Next
input— caller data — the other half of the run body, and the one with no schema- Runs and results — the receipt that contains all of this
- Calling a Format — the invoke contract
- Embed a Format in your product — how to map output into your own records