Remix media
Dest only (#8076). Sume lists
POST /v1/remix-media/:verb,GET /v1/remix-media/:id, and thesume_*tools only whereSUME_COM_REMIX_MEDIA_ENABLEDpermits them (development auto-on, production opt-in). Every name in this lane has theremix-prefix. This lane does not change Reference ingest orvideo_inspect.
Remix media 1.0 returns the artifacts that a director's read of a
reference leaves. The artifacts are a
probe, a transcript.json with per-word confidence, and a mechanical
boundaries series. They also include time- and word-labeled tile grids
and frames, a time-labeled cut excerpt, and the agent's ANALYSIS.md /
TIMELINE.md / PROGRESS.md.
One job type (remix_media) carries one verb in each request. The first
verb, probe, mints the understanding id (its own job id). Every later
verb binds to that id. The compute runs on its own media workers with a
source cache. Thus, the tenth grid of one understanding never downloads the
source again.
The hosted MCP tools are sume_probe, sume_transcribe, sume_boundaries,
sume_tile, sume_cut, sume_notes_write, sume_align, and
sume_understand_get. The Sume Agent host adds sume-agent__sume_*. These
tools run the same verbs and attach sume_tile's pages inline as images
(up to six per call). The host also adds the host-only
sume-agent__sume_compose (below).
Verbs
Order: one probe, then batches
Every verb after probe needs only the understanding_id. The app runs these
verbs in parallel containers. Thus, we recommend that an agent does not run the verbs one after
another.
In stage 0, run probe alone. In stage 1, run one batch of transcribe +
boundaries + the overview tile calls. In stage 2, run dense tile calls
and cuts that you select from the boundaries and the overview (no transcript
is necessary). In stage 3, when the transcript is available, run tile with
around { phrase } and any word-labeled close read (transcript_job_id). In
stage 4, write the three notes, then get the bundle.
Only around, word labels, and the notes wait on the transcript. You never
need a second transcribe.
| Verb | Produces | Billing |
|---|---|---|
probe | ffprobe facts, 16 kHz wav (audio.wav_url), loudness gate (audio.silent) | Modal compute |
transcribe | remix.transcript/1: passages[] → words[] with start_seconds, end_seconds, score (0–1). The default is engine: whisperx, WhisperX (faster-whisper ASR + wav2vec2 alignment on the media workers). This engine gives a score on every aligned word. engine: scribe_v2 is Sume STT 1.0, with no per-word score. You must send an explicit language. checkpoint is optional: small is the default (int8 on a CPU container), and medium, large-v3-turbo, large-v3 are opt-in on the GPU. This route has no model field, because model is the public model id | whisperx = its Modal compute, with no STT line. scribe_v2 = the STT per-minute rate on the probe's duration. scribe_v2 needs allow_billed_stt: true, and the charge settles to zero when the audio is silent |
boundaries | remix.boundaries/1: candidates[] {at, score} from a 12 fps 32×32 mean-absolute-difference pass (defaults rate 12, threshold 0.1) | Modal compute |
tile | remix.tile/1: pages of exact frames. Each cell has the label HH:MM:SS.mmm. With transcript_job_id, each cell also shows the words spoken at that instant, with spans and context. layout: frames = one native-size labeled picture per instant. every_frame = every native frame of ≤ 4 s | Modal compute |
cut | remix.cut/1: an MP4 of [start, end) or joined keep[] spans (≤ 120 s), with mapping[]. When label_time is set, the cut burns the source time into each frame | Modal compute |
notes | remix.notes/1: ANALYSIS.md, TIMELINE.md, PROGRESS.md (the read of the reference) or FORMAT.md, TREATMENT.md (the project: why the format works and which spoken words its captions attach to, plus the target treatment) as a text artifact. A light lint applies: evidence links must be media.sume.com artifacts, and FORMAT.md must have a ## Captions section | unbilled |
align | remix.take/1 (#8076 steps 2–5): one script.svml segment aligned to a generated take's transcript (the take's own probe + transcribe). The output has every script token with the measured start_seconds / end_seconds / score of the transcript words that it paired with (an authored eojeol can pair with two heard words). It also has cues[] from || breaks and role changes, selections{} and moments{} resolved to word boundaries, and unmatched[] for words that the take did not say. No argument carries a time. An unmatched word stays untimed. take.json + script.svml are artifacts, and the bundle lists takes[] | unbilled |
To select samples on tile, use one of these: start/end with every
(seconds) or frames (evenly spaced midpoints, approximately 1.5 a second by
default, at least 4, at most 9), at[] explicit instants,
around {phrase, occurrence, padding} resolved on the transcript, or
ranges[] for several stretches. The verb sizes the pages so that the agent
host's 1080 px inline downscale does not make a cell smaller. For example, a
480 px portrait cell stays 480 px, with two cells on each page. To trade size
for count, pass cell and columns.
Every image, clip, and json is a durable media.sume.com artifact that the
job's artifacts[] lists. Thus, a later step binds it by id. For
idempotency, the host tools derive the key from the verb and its arguments.
On REST, send Idempotency-Key.
Compose (steps 2–5, Sume Agent host only)
sume-agent__sume_compose takes composition.svml. This file has one
<take id job src/> for each aligned segment. It binds
<video> / <image> / <audio> / <text> with
during / at + for / until + for / start + end. The bind targets
are story.selection.<name>, story.moment.<name>, story.segment.<id>,
take.<id>, or program. The clock literals are 2s / 12f / 250ms with
± offsets, and the tool refuses a literal alone in during. The file also
has <captions> (a kit face, size, colors, position,
karaoke: progress | active | off).
The tool compiles the HyperFrames document. In this document, every cue is a
data-sume-caption root with attribute-timed word spans. The tool runs the
document on the current compose engine. output: "check" = free lint + a
Modal-metered Chromium audit with a checked_draft_ref for
hyperframes_snapshot. output: "mp4" = the billed video_caption bake,
with the thread composition revision on the Studio timeline. The result's
svml receipt lists every take, cue, and anchor with the spoken word that it
resolved through.
On dest, sume/remix carries the order: project notes →
produce with the current generation tools → measure each take → SVML → align →
check → snapshot → bake.
Errors that name the fix
| Code | Meaning |
|---|---|
source_too_long_for_remix_media | The clip is longer than 300 s. |
source_no_video_stream | The file has no video stream. |
remix_understanding_not_found | understanding_id is not a completed probe of this workspace. |
remix_tile_too_many_cells | The selectors resolve to more than 96 cells. Increase every, decrease frames, or split ranges[]. |
remix_transcript_not_found | transcript_job_id is not a completed transcribe of this understanding in this workspace. |
remix_tile_transcript_required | The request has around without transcript_job_id. |
remix_transcribe_checkpoint_conflict | The request sent checkpoint with engine: scribe_v2. The checkpoint is a WhisperX knob. |
remix_align_script_invalid | script.svml did not parse. The message carries the remix_svml_* code and line (a marker inside a word, a selection across segments, a reused name). |
remix_align_segment_unknown | segment is not a segment of the script. |
remix_svml_* (compose) | composition.svml has a problem: an unknown selection or moment, a take without an align job, a clock literal in during, a non-kit font, or a non-media.sume.com source. |
remix_media_unavailable | The media workers are not deployed in this environment. |
ffmpeg_fields_rejected | The request has vf / filtergraph / codec / argv fields. Sume compiles every pass. |
Not this surface
| Need | Use |
|---|---|
| Deterministic shots + OCR text tracks + audio facts in one call | Reference ingest |
| A quick probe or unlabeled stills | Video inspect |
| A production trim | Video trim |