Remix media

Dest only (#8076). Sume lists POST /v1/remix-media/:verb, GET /v1/remix-media/:id, and the sume_* tools only where SUME_COM_REMIX_MEDIA_ENABLED permits them (development auto-on, production opt-in). Every name in this lane has the remix- prefix. This lane does not change Reference ingest or video_inspect.

Remix media 1.0 returns the artifacts that a director's read of a reference leaves. The artifacts are a probe, a transcript.json with per-word confidence, and a mechanical boundaries series. They also include time- and word-labeled tile grids and frames, a time-labeled cut excerpt, and the agent's ANALYSIS.md / TIMELINE.md / PROGRESS.md.

One job type (remix_media) carries one verb in each request. The first verb, probe, mints the understanding id (its own job id). Every later verb binds to that id. The compute runs on its own media workers with a source cache. Thus, the tenth grid of one understanding never downloads the source again.

The hosted MCP tools are sume_probe, sume_transcribe, sume_boundaries, sume_tile, sume_cut, sume_notes_write, sume_align, and sume_understand_get. The Sume Agent host adds sume-agent__sume_*. These tools run the same verbs and attach sume_tile's pages inline as images (up to six per call). The host also adds the host-only sume-agent__sume_compose (below).

Verbs

Order: one probe, then batches

Every verb after probe needs only the understanding_id. The app runs these verbs in parallel containers. Thus, we recommend that an agent does not run the verbs one after another.

In stage 0, run probe alone. In stage 1, run one batch of transcribe + boundaries + the overview tile calls. In stage 2, run dense tile calls and cuts that you select from the boundaries and the overview (no transcript is necessary). In stage 3, when the transcript is available, run tile with around { phrase } and any word-labeled close read (transcript_job_id). In stage 4, write the three notes, then get the bundle.

Only around, word labels, and the notes wait on the transcript. You never need a second transcribe.

VerbProducesBilling
probeffprobe facts, 16 kHz wav (audio.wav_url), loudness gate (audio.silent)Modal compute
transcriberemix.transcript/1: passages[] → words[] with start_seconds, end_seconds, score (0–1). The default is engine: whisperx, WhisperX (faster-whisper ASR + wav2vec2 alignment on the media workers). This engine gives a score on every aligned word. engine: scribe_v2 is Sume STT 1.0, with no per-word score. You must send an explicit language. checkpoint is optional: small is the default (int8 on a CPU container), and medium, large-v3-turbo, large-v3 are opt-in on the GPU. This route has no model field, because model is the public model idwhisperx = its Modal compute, with no STT line. scribe_v2 = the STT per-minute rate on the probe's duration. scribe_v2 needs allow_billed_stt: true, and the charge settles to zero when the audio is silent
boundariesremix.boundaries/1: candidates[] {at, score} from a 12 fps 32×32 mean-absolute-difference pass (defaults rate 12, threshold 0.1)Modal compute
tileremix.tile/1: pages of exact frames. Each cell has the label HH:MM:SS.mmm. With transcript_job_id, each cell also shows the words spoken at that instant, with spans and context. layout: frames = one native-size labeled picture per instant. every_frame = every native frame of ≤ 4 sModal compute
cutremix.cut/1: an MP4 of [start, end) or joined keep[] spans (≤ 120 s), with mapping[]. When label_time is set, the cut burns the source time into each frameModal compute
notesremix.notes/1: ANALYSIS.md, TIMELINE.md, PROGRESS.md (the read of the reference) or FORMAT.md, TREATMENT.md (the project: why the format works and which spoken words its captions attach to, plus the target treatment) as a text artifact. A light lint applies: evidence links must be media.sume.com artifacts, and FORMAT.md must have a ## Captions sectionunbilled
alignremix.take/1 (#8076 steps 2–5): one script.svml segment aligned to a generated take's transcript (the take's own probe + transcribe). The output has every script token with the measured start_seconds / end_seconds / score of the transcript words that it paired with (an authored eojeol can pair with two heard words). It also has cues[] from || breaks and role changes, selections{} and moments{} resolved to word boundaries, and unmatched[] for words that the take did not say. No argument carries a time. An unmatched word stays untimed. take.json + script.svml are artifacts, and the bundle lists takes[]unbilled

To select samples on tile, use one of these: start/end with every (seconds) or frames (evenly spaced midpoints, approximately 1.5 a second by default, at least 4, at most 9), at[] explicit instants, around {phrase, occurrence, padding} resolved on the transcript, or ranges[] for several stretches. The verb sizes the pages so that the agent host's 1080 px inline downscale does not make a cell smaller. For example, a 480 px portrait cell stays 480 px, with two cells on each page. To trade size for count, pass cell and columns.

Every image, clip, and json is a durable media.sume.com artifact that the job's artifacts[] lists. Thus, a later step binds it by id. For idempotency, the host tools derive the key from the verb and its arguments. On REST, send Idempotency-Key.

Compose (steps 2–5, Sume Agent host only)

sume-agent__sume_compose takes composition.svml. This file has one <take id job src/> for each aligned segment. It binds <video> / <image> / <audio> / <text> with during / at + for / until + for / start + end. The bind targets are story.selection.<name>, story.moment.<name>, story.segment.<id>, take.<id>, or program. The clock literals are 2s / 12f / 250ms with ± offsets, and the tool refuses a literal alone in during. The file also has <captions> (a kit face, size, colors, position, karaoke: progress | active | off).

The tool compiles the HyperFrames document. In this document, every cue is a data-sume-caption root with attribute-timed word spans. The tool runs the document on the current compose engine. output: "check" = free lint + a Modal-metered Chromium audit with a checked_draft_ref for hyperframes_snapshot. output: "mp4" = the billed video_caption bake, with the thread composition revision on the Studio timeline. The result's svml receipt lists every take, cue, and anchor with the spoken word that it resolved through.

On dest, sume/remix carries the order: project notes → produce with the current generation tools → measure each take → SVML → align → check → snapshot → bake.

Errors that name the fix

CodeMeaning
source_too_long_for_remix_mediaThe clip is longer than 300 s.
source_no_video_streamThe file has no video stream.
remix_understanding_not_foundunderstanding_id is not a completed probe of this workspace.
remix_tile_too_many_cellsThe selectors resolve to more than 96 cells. Increase every, decrease frames, or split ranges[].
remix_transcript_not_foundtranscript_job_id is not a completed transcribe of this understanding in this workspace.
remix_tile_transcript_requiredThe request has around without transcript_job_id.
remix_transcribe_checkpoint_conflictThe request sent checkpoint with engine: scribe_v2. The checkpoint is a WhisperX knob.
remix_align_script_invalidscript.svml did not parse. The message carries the remix_svml_* code and line (a marker inside a word, a selection across segments, a reused name).
remix_align_segment_unknownsegment is not a segment of the script.
remix_svml_* (compose)composition.svml has a problem: an unknown selection or moment, a take without an align job, a clock literal in during, a non-kit font, or a non-media.sume.com source.
remix_media_unavailableThe media workers are not deployed in this environment.
ffmpeg_fields_rejectedThe request has vf / filtergraph / codec / argv fields. Sume compiles every pass.

Not this surface

NeedUse
Deterministic shots + OCR text tracks + audio facts in one callReference ingest
A quick probe or unlabeled stillsVideo inspect
A production trimVideo trim