Structured output

By default, a completed Format run gives you media and a paragraph of text. This result is good for a person, but it is not easy to use in a database. If you bind a schema, you get a typed object. It is the same run, but with a shape that you can write directly into your own records.

The exact request and response schemas come from the live OpenAPI (https://api.sume.com/reference/json). The tables on this page are a summary that is easy to read. They are not a second schema.

input

A run request has two JSON-shaped fields, and they operate in very different ways. Confusion between the two fields is the most frequent cause of errors in a first integration.

inputoutput_schema
What it isCaller dataA contract for the receipt
DirectionYou → the runThe run → you
ShapeAny JSON object that your backend needsJSON Schema, inside the supported subset
Checked forObject type, key count, byte sizeEvery rule in the subset
A shape Sume does not expectThe run starts. Unknown keys are only more data.400 output_schema_invalid. Nothing runs, and Sume charges nothing.
Where it landsThe agent's prompt, as a fenced data blockThe post-run projection

Thus, input is a flexible concatenation, and output_schema is a strict typed receipt. You can make your input as loose as you want. You cannot make your output_schema loose at all. strict: false does not make it loose, and there is no other way to bypass the rules.

The two fields do not interact. The projection never sees your input (refer to the sections below). Thus, a value that you sent cannot come back in output, unless the run repeats it in its final text.

Where your object comes from

This part is different from a chat completion. Make sure that you understand it before you design a schema. There are two paths, and the receipt tells you which path you got.

filled_by: "agent" — the run answered. Sume gives your schema to the run as a tool that the run must call before it stops. Your schema is the argument shape of that tool. The model that does the work is also the model that fills your object. It fills the object while it still knows what it made and why. It can see your input and your instruction, because they are part of the run.

filled_by: "projection" — the fallback. If the run stops and did not submit a valid object, a separate constrained pass builds one after the run. The pass uses the data that the run left. This pass is an OpenAI-strict json_schema completion at temperature 0. It gets only two facts:

FactDetail
The run's generated mediaAll the artifacts that the run produced, with their durable URLs and metadata.
The run's closing textThe final assistant message, truncated to the first 8000 characters.

The pass does not get your input, not your instruction, not the Format body, and not the intermediate steps of the run. On this path, all data that you want in output must be in one of these two facts.

Thus, it is important to read filled_by. It shows if the run wrote the object, or if a pass built the object again from the data that the run left. It explains most unexpected results:

  • On the projection path, your own identifiers do not round-trip. The projection cannot see an order_id that you sent in input. Keep your identifiers on your side, with run.id or the Idempotency-Key that you sent as the key. Use output only for the data that the run made.
  • Nothing in output is invented, on either path. The same checks gate both paths before the object gets to you (refer to the sections below). An object that the run wrote does not get more trust.
  • On the projection path, an empty text field is a signal. The projection never sees your input. Thus, if the titles, descriptions or ids in a schema come from the brief, these fields come back null. At the same time, the media fields are full. filled_by: "projection" with null prose shows a run that stopped early. It does not show a Format that did not write copy.

If you score runs automatically (a smoke matrix, a partner integration, a dashboard), read filled_by before you count a run as delivered. For video, read the file, not the number next to it. ffprobe on primary_output_url takes two seconds. It is the only method that shows the difference between an assembled cut and a clip that has the shape of one.

The practical rule that comes from this is the same: require only what the Format actually makes. If a schema requires a field that the recipe never produces, the schema will fail on the projection path each time. The failure is silent until you read output_error.

Bind a schema

There are two spellings, with one behavior. Use the spelling that is applicable to your client.

Native — output_schema:

OpenAI-shaped alias — response_format:

Sume normalizes response_format into output_schema on the receipt. Thus, the receipt always shows the native spelling. If you send both, the result is 400 invalid_request.

The alias uses the Chat Completions spelling: type at the top, and the binding nested under json_schema. The OpenAI Responses API flattens the same fields into text.format: { type, name, strict, schema }. Sume does not accept this flattened shape. If your client builds Responses-style bodies, move the fields into output_schema, or nest them again under json_schema.

FieldRules
nameRequired. 1–64 characters, ^[A-Za-z0-9._/-]+$. Give it a namespace, because it shows on every receipt.
strictThe default is true. Refer to the note under Supported schemas.
schemaRequired. A JSON Schema object inside the supported subset.

A Format can also have its own default schema, which you bind in the dashboard. A per-request output_schema overrides it for that run. The receipt shows which schema applied, in output_schema.source:

sourceMeaning
defaultNo schema is bound. output is the built-in schema.
action_defaultThe Format's own bound schema. (action_ is the wire spelling. Scheduled uses the same spelling.)
request_overrideThe output_schema that you sent on this run request.

Supported schemas

Schemas must agree with the OpenAI strict-mode subset. This subset is a requirement, not a recommendation. If a schema is outside the subset, Sume rejects it at submit with 400 output_schema_invalid. The response has a details.violations[] array that names each problem. Nothing runs. Thus, Sume does not charge you.

The supported keywords, exactly

The subset uses an allowlist. A keyword that is not on this list is a violation. Sume does not ignore it silently. An ignored constraint is a part of the schema that Sume cannot promise to satisfy.

GroupAccepted
Structuretype, properties, required, additionalProperties, items, $defs, $ref, anyOf
Valuesenum, const
Stringsformat, pattern, minLength, maxLength
Numbersminimum, maximum, exclusiveMinimum, exclusiveMaximum, multipleOf
ArraysminItems, maxItems
Annotationtitle, description, default, examples, $schema, $id

Types are string, number, integer, boolean, object, array, null.

Sume rejects all other keywords. These are the keywords that cause problems for real integrations:

RejectedInstead
oneOfanyOf. Only anyOf is on the list. Schemas ported from OpenAPI usually use oneOf.
allOfFlatten the branches into one object.
not, if / then / else, dependentRequired, dependentSchemasYou cannot express these. Model the alternatives as anyOf, or validate on your side after you read output.
nullable: trueA nullable union: "type": ["string", "null"].
patternProperties, propertyNames, unevaluatedProperties, additionalItemsDeclare the properties that you want. additionalProperties: false covers the rest.

type

A node can also have a $ref or an anyOf in its place. A bare { "description": "…" } is legal JSON Schema, and it means "anything". But here it is a missing_type violation. An array must also declare items.

$ref and anyOf each short-circuit the node that contains them. Sume checks the sibling keywords next to them against the allowlist, but these keywords have no other effect. Put constraints inside the anyOf branches or inside the $defs entry, not next to the $ref.

The root must be an object

The type of the root must be "object", as a single type. Thus, Sume rejects even { "type": ["object", "null"] }. Sume also rejects a top-level array, string, or union. Wrap it:

additionalProperties: false

This rule is applicable to all objects in the schema, not only to the root. This includes the objects nested inside array items and inside $defs.

required

There is no optional property. Each property that you declare must be present.

Use a nullable union to express optionality:

This rule causes problems for more ported schemas than any other rule. Read null as "the run had nothing to put here". This is the same case for which you wanted optional.

Size and nesting limits

LimitValueViolation
Nesting depth10 levelsmax_depth
Total properties5000, counted across the whole documentmax_properties
Enum values1000 per enummax_enum_values
Total string length120,000 characters, summed over every property name, key, and string value in the documentmax_string_length

The last limit is a document-wide budget, not a per-field cap. Thus, long description annotations on a large schema can use all of it, even when no single string is very long.

is limited

Only two targets resolve:

TargetUse
#/$defs/*Your own definitions, declared at the root of the schema document.
SumeMediaFile#Sume's media shape. Refer to the section below.

Sume rejects an external $ref (a URL, a sibling document, #/components/...). Sume also rejects $ref: "#". OpenAI strict mode permits root recursion in this way, but Sume does not. Sume also rejects a #/$defs/* target that has no root $defs entry of that name. This rule finds a $defs block that is nested inside a sub-schema and not declared at the root.

Recursion through a named definition is permitted. A $defs entry can $ref itself. The depth limit counts literal nesting in the document. Thus, a self-referential definition does not increase the depth count.

does not relax any of this

Sume accepts and stores it, but it does not change the subset above. If a schema is outside the subset, Sume rejects it, whether strict is true or false. Do not use it to bypass the subset. There is no way to bypass it.

There is also no equivalent of the OpenAI JSON mode ({"type": "json_object"}), the loose "valid JSON, any shape" option. Bind a schema or use the built-in one. These are the only two options.

details.violations[]

Each entry is { path, rule, message }. path is a JSON-Pointer-style location, for example #/properties/scenes/items/properties/clip. message is text for a person, and it can change. rule is a stable lowercase token that you can safely switch on:

ruleMeaning
root_must_be_objectThe root is missing, is not an object, or its type is not exactly "object".
not_an_objectA schema node is not a JSON object.
missing_typeA node has no type, $ref, or anyOf.
unsupported_typeA type that is not one of the seven types above.
unsupported_keywordA keyword that is not on the allowlist.
additional_properties_falseAn object node without additionalProperties: false.
required_completenessA declared property that is not in required, or a required entry that has no property of that name.
missing_itemsAn array node with no items.
unsupported_refA $ref that is neither #/$defs/<name> nor SumeMediaFile#, or one that names a definition that does not exist.
invalid_defs$defs is present but is not an object of named schemas.
max_depth, max_properties, max_enum_values, max_string_lengthThe limits above.

Sume reports all the problems, not only the first. Thus, one 400 gives you sufficient data to fix the schema.

Coming from OpenAI structured outputs

If you used response_format: { type: "json_schema", … } or the Responses API's text.format, most of what you know is applicable here. The subset rules are the same, and they come from the same guide. The difference is where Sume applies the schema.

OpenAISumeNote
response_format.json_schemaoutput_schema, or response_format verbatimChat Completions spelling only. Sume does not accept text.format.
json_schema.nameoutput_schema.nameRequired in both places. Give it a namespace, because it shows on every receipt.
json_schema.strictoutput_schema.strictAccepted, the default is true, and it changes nothing. Sume always enforces the subset.
{"type": "json_object"} (JSON mode)(no equivalent)Bind a schema, or use the built-in one.
The model emits the JSONA post-run projection emits itThe schema constrains the projection, never the run.
refusal on the messageoutput_error on the receiptDifferent mechanism: not a safety refusal but a failed projection.
incomplete_details.reason: "max_output_tokens"(not applicable)The projection is small and bounded. There is no truncated-JSON case to handle.
Streamed partial JSON(not applicable)output appears once, on the terminal receipt.
$ref: "#" root recursionRejectedRecurse through a named #/$defs/* entry instead.
Nothing comparableThe URL gateSume checks each URL in output against the media that the run really produced.
Nothing comparableSumeMediaFile#A built-in $ref target for the run's media.

The change in how you think is this: with OpenAI, you constrain what the model says. Here, you constrain how Sume reads back a completed run. All other differences come from this. It is why the schema cannot make the Format produce a video, and why a hallucinated URL cannot get into output. It is also why a run can give you output: null.

SumeMediaFile

Use { "$ref": "SumeMediaFile#" } at each location where you want a piece of the media of the run in your output. All fields are required. All fields except type and url are nullable.

FieldTypeNotes
type"image" | "video" | "audio" | "file"
urlstring (uri)Must be a URL that this run actually produced. Refer to the URL gate.
content_typestring | nullFor example, video/mp4.
file_namestring | null
size_bytesinteger | null
width, heightinteger | nullImages and video.
duration_msinteger | nullVideo and audio.
expires_atstring (date-time) | nullnull for durable media.sume.com URLs, which is the normal case. It has a value only when Sume returns a signed URL.

Sume serves Sume-hosted media from media.sume.com as public, max-age=31536000, immutable, and the media does not expire. Store the URL with your own record and render it later. It is not necessary to refresh it. A durable URL is also a public URL. If this is important to your product, refer to Map artifacts into your UI.

The built-in schema

If you do not bind a schema, Sume projects output onto sume/action-run-output/v1:

text is nullable. The four arrays are always present, and they can be empty.

Sume fills the built-in schema deterministically from the generated media and the final text of the run. Sume does not use a model. Thus, it cannot fail in the way that a custom schema can fail. If you want the media but do not want to design a schema, this schema is sufficient to ship with.

The URL gate

Before a custom output gets to you, Sume checks each URL in it against the set of media that this run actually produced. The comparison is exact string equality.

Two passes find these URLs. The first pass collects each http(s):// string in the object, at any depth. It does this whether or not you declared the field as media. The second pass collects the url of each SumeMediaFile-shaped value, whatever it contains. Thus, a placeholder such as "none" or "" in a url cannot get past the gate only because it does not look like a URL.

If this run did not produce a URL, the projection fails, and Sume does not return the URL. This is also true for a well-formed media.sume.com URL that looks correct. Then Sume validates the object against your schema.

Thus, a completed run returns output that agrees with your schema and has real media URLs. Or, it returns output: null and tells you why. It never returns a schema-shaped guess.

The deliverable is not one of its parts

A schema that describes a whole made of parts (scenes, shots, segments) usually also has a field for the completed, assembled file. These are different files. Sume rejects a receipt that fills the whole with one of its own parts.

The check is narrow intentionally. It needs two or more parts that report "status": "succeeded" with their own video. It also needs a video outside all parts that uses one of those files again. The check does not affect a Format whose single clip really is the deliverable. It also does not affect a poster or thumbnail that the run intentionally reused from a part.

A run that could not assemble the file has an honest answer available. That answer is not "here is a scene". Report the parts that you made and leave the assembled field null. Or, report the real status of the assembly.

Durations are checked against the file

A duration_ms inside a SumeMediaFile comes from the same ledger that fills artifacts[]. If that ledger recorded a length, the value in output must agree with it within 10%. A value that does not agree describes a different file. It fails the projection and does not get to you.

If the ledger did not record a length, Sume does not check the value. In this case, null means "not measured", not "zero". Sume does not fail a run because of a fact that nobody measured. Thus, a duration_ms that you read back is the value of the artifact, or it is not verified. It is never a number calculated from other data.

What the gate does not check

The gate checks only URLs, durations, and the shape. The projection pass reads all other values (ids, labels, captions, counts) from the media metadata and the closing text of the run. These values come from what the run reported, but Sume does not verify them against the run. Thus, treat these fields as the report of the run about its work, not as measurements. The projection writes a duration_seconds number that you declared yourself. The duration_ms on the media file is the value that Sume checks.

The gate also admits only media that the run generated. A file that the run only uploaded is not in that set. Thus, if a schema field holds the URL of an upload, the projection fails, and all of the output fails with it. Do not put uploads in your schema.

and

Most integrations have one item to show. If you name it, the receipt resolves it for you:

Then primary_output_url on the receipt is the URL at that key. The resolution order is:

  1. The primary_output_key on the run request.
  2. The Format's own primary_output_key.
  3. The first top-level key that holds a media object, or an array whose first element is a media object.

For the built-in schema, the fallback order is videos → images → audio → files.

If you named a key in step 1 or 2, the receipt echoes it back when output has a value under it. This includes a value that is not media. If you point at a headline string, you get primary_output_key: "headline" with primary_output_url: null, because there is no URL to resolve. The resolution goes to step 3 only when the key is missing from output, or empty there.

primary_output_key has a maximum of 64 characters. Both fields are null on any non-terminal status. They are also null when output_error is set.

When output cannot be produced

Before you read output, examine output_error.

On a run over the API, a projection failure is a run failure. These runs are unattended, as Call a Format states. A completed receipt whose output is null looks like a success, but it is not one. Thus, status is failed, and error has the same reason as output_error. In the Agents UI, a person reads the thread. There, the same shape stays completed, because it is a draft, not a receipt.

A failed run still publishes what it produced. output is null when nothing satisfied your schema, not only because the run failed. For example, a live-commerce show rendered 20 of 40 scenes and then stopped. The show reports those 20 on output. The other scenes have the "failed" value that your own schema defines. The output_error next to them tells why the run stopped.

A failure never gets a pointer: primary_output_key and primary_output_url are null on every non-completed run. Thus, if (run.primary_output_url) stays a safe test for "the deliverable exists", and a partial result cannot make it give a wrong answer.

For this to work, your schema must permit the partial result. Refer to Make a partial result legal.

artifacts[] is populated either way, with all the items that the run made. You always have the media, even when the shape failed.

The exception is a projection that did not get access to the run. output_extraction_failed records a transport failure, not a verdict. Thus, the run stays completed, and Sume projects it again automatically on your next read.

output_error.codeMeaningdetails
output_schema_unsatisfiedThe projection did not agree with your schema, or it referenced media that this run did not produce.rejected_urls[] (first 10 only) or violations[], and a harvested count by media type.
output_extraction_failedThe projection could not run. status stays completed. A reason of harvest_unavailable means that Sume could not read the media of the run while the run finalized. The receipt fills in on the next read.reason
unattended_blockedThe run stopped and did not claim a deliverable that it did not produce. message is the reason in the words of the run.A harvested count by media type, and assembled_deliverable: false. The harvested media are intermediates, not the completed cut.
deliverable_missingThe Format produces media (io.output_kind) that the run never made. Thus, Sume did not let any structured output claim it.declared_output_kind, and a harvested count by media type.
primary_output_missingThe run satisfied your schema, but the primary_output_key that you declared has no value. output still has the partial result. The run is failed because the item that you named is not in it.primary_output_key
agent_reported_failureThe accepted receipt of the run says that the run did not deliver: an explicit-fail payload, failed / stand-in media slots, or a primary of the wrong media type for the io.output_kind of the Format. output still has the ledger. The run is failed because the receipt says so.reason (explicit_failed, slots_failed, primary_not_deliverable), non_delivered_slots[] or primary_media_type, and a harvested count by media type.

Treat this set as open, because new codes can appear. Branch on the codes that you handle, and use a default path for the rest. Do not use an exhaustive switch.

Handle each case as this list shows:

  • output_schema_unsatisfied, repeatedly, on the same Format. The cause is almost always a schema that requires media that the recipe does not generate. Compare details.harvested with your required fields. The example above requires an image and got one, but it wanted a second image. Make the field a nullable union, or change the instruction so that the run makes the media.
  • output_schema_unsatisfied with violations[]. The problem is the shape, not the media. The violations name the paths that cause the problem.
  • output_extraction_failed. This failure is transient. Before you do anything else, read the run one more time. This step alone clears a harvest_unavailable. If the failure continues, retry the run with a new idempotency key. The old key is bound to the receipt that you already have.
  • Any of them, in your UI. You still have artifacts[]. Show the media and log the shape failure. This is better than an error for a customer whose video exists.

A run can produce some of its output and then fail. The run can report this only if your schema says a partial is a legal shape. The platform does not add a minimum of its own. Ajv enforces the keywords that you wrote, and only those keywords. Thus, if a schema requires every scene, a 20-of-40 show returns output: null.

Two rules do all the work. Both rules are not intuitive under the strict subset:

  1. Optional means nullable, not absent. You must still list each property in required (refer to Every property must be listed in required). To show that a value can be missing, use "type": ["string", "null"].
  2. Do not use minItems on the arrays that you want to receive partially. Sume really enforces it. Thus, minItems: 1 on a scene list will reject the ledger that you want to read.

Use it together with "primary_output_key": "full_video". This keeps the loose schema honest. A run that fills scenes but leaves full_video null satisfied the schema, but it did not produce the deliverable. Thus, the run terminalizes as failed with primary_output_missing, and it does not report a success. The loose schema lets you see the partial result. It does not let the run pass.

When you read one of these receipts, branch on full_video for "did I get a show". Branch on each scenes[].status for "what do I need to retry". To retry, send the id of the failed run as previous_run_id on a new run (refer to Continue a run). The next turn continues the same conversation, and those clips are already in it.

Failures at submit

Sume finds these schema problems before anything runs. Sume does not charge you.

CodeStatusWhat to do
output_schema_invalid400Your schema is outside the supported subset. details.violations[] names each problem.
invalid_request400This includes a request that sends both output_schema and response_format.

Checklist

  • The root is {"type": "object"} only. additionalProperties: false is on every object, also in $defs.
  • Every node has a type, a $ref, or an anyOf. Every array declares items.
  • Every declared property is in required. Optionality is a nullable union.
  • No oneOf, allOf, not, if/then/else, or nullable: true.
  • $ref targets are only #/$defs/* (declared at the root) and SumeMediaFile#, and never #.
  • The schema requires only media that the Format actually produces.
  • No field expects a value that you sent in input, because the projection cannot see it.
  • name is namespaced and stable. Thus, you can grep the receipts.
  • primary_output_key names the one item that your UI shows.
  • Your reader handles output: null with output_error set, on a failed run and on a completed one.
  • When output is null, your reader uses artifacts[] as a fallback.

Next