---
title: Remix media
description: The remix media verbs for one reference clip — probe, word-timed transcript with scores, cut candidates, time- and word-labeled grids, labeled excerpts, notes — each a durable artifact with an id. Dest only.
---

> **Dest only (#8076).** Sume lists `POST /v1/remix-media/:verb`,
> `GET /v1/remix-media/:id`, and the `sume_*` tools only where
> `SUME_COM_REMIX_MEDIA_ENABLED` permits them (development auto-on,
> production opt-in). Every name in this lane has the `remix-` prefix. This
> lane does not change
> [Reference ingest](/models/reference-ingest) or `video_inspect`.

Remix media 1.0 returns the artifacts that a director's read of a
reference leaves. The artifacts are a
probe, a `transcript.json` with per-word confidence, and a mechanical
`boundaries` series. They also include time- and word-labeled `tile` grids
and `frames`, a time-labeled `cut` excerpt, and the agent's `ANALYSIS.md` /
`TIMELINE.md` / `PROGRESS.md`.

One job type (`remix_media`) carries one verb in each request. The first
verb, `probe`, mints the **understanding id** (its own job id). Every later
verb binds to that id. The compute runs on its own media workers with a
source cache. Thus, the tenth grid of one understanding never downloads the
source again.

```text
POST /v1/remix-media/probe        video_url                              → understanding_id
POST /v1/remix-media/transcribe   understanding_id, language, engine? (whisperx default | scribe_v2 + allow_billed_stt: true), checkpoint? (small default)
POST /v1/remix-media/boundaries   understanding_id, rate?, threshold?
POST /v1/remix-media/tile         understanding_id, layout?, at[] | around | ranges[] | start/end + every|frames, every_frame?, cell?, transcript_job_id?
POST /v1/remix-media/cut          understanding_id, start/end | keep[], label_time?
POST /v1/remix-media/notes        understanding_id, name (ANALYSIS.md | TIMELINE.md | PROGRESS.md | FORMAT.md | TREATMENT.md), markdown
POST /v1/remix-media/align        understanding_id, take_understanding_id, transcript_job_id, script, segment   → remix.take/1
GET  /v1/remix-media/:id          the job; for a probe id also the bundle (remix.understanding/1, with takes[])
```

The hosted MCP tools are `sume_probe`, `sume_transcribe`, `sume_boundaries`,
`sume_tile`, `sume_cut`, `sume_notes_write`, `sume_align`, and
`sume_understand_get`. The Sume Agent host adds `sume-agent__sume_*`. These
tools run the same verbs and attach `sume_tile`'s pages **inline as images**
(up to six per call). The host also adds the host-only
`sume-agent__sume_compose` (below).

## Verbs

## Order: one probe, then batches

Every verb after `probe` needs only the `understanding_id`. The app runs these
verbs in parallel containers. Thus, we recommend that an agent does not run the verbs one after
another.

In stage 0, run `probe` alone. In stage 1, run one batch of `transcribe` +
`boundaries` + the overview `tile` calls. In stage 2, run dense `tile` calls
and `cut`s that you select from the boundaries and the overview (no transcript
is necessary). In stage 3, when the transcript is available, run `tile` with
`around { phrase }` and any word-labeled close read (`transcript_job_id`). In
stage 4, write the three `notes`, then get the bundle.

Only `around`, word labels, and the notes wait on the transcript. You never
need a second `transcribe`.

| Verb | Produces | Billing |
|---|---|---|
| `probe` | ffprobe facts, 16 kHz wav (`audio.wav_url`), loudness gate (`audio.silent`) | Modal compute |
| `transcribe` | `remix.transcript/1`: `passages[]` → `words[]` with `start_seconds`, `end_seconds`, `score` (0–1). The default is `engine: whisperx`, WhisperX (faster-whisper ASR + wav2vec2 alignment on the media workers). This engine gives a score on every aligned word. `engine: scribe_v2` is Sume STT 1.0, with no per-word score. You must send an explicit `language`. `checkpoint` is optional: `small` is the default (int8 on a CPU container), and `medium`, `large-v3-turbo`, `large-v3` are opt-in on the GPU. This route has no `model` field, because `model` is the public model id | whisperx = its Modal compute, with no STT line. scribe_v2 = the STT per-minute rate on the probe's duration. scribe_v2 needs `allow_billed_stt: true`, and the charge settles to zero when the audio is silent |
| `boundaries` | `remix.boundaries/1`: `candidates[] {at, score}` from a 12 fps 32×32 mean-absolute-difference pass (defaults `rate` 12, `threshold` 0.1) | Modal compute |
| `tile` | `remix.tile/1`: pages of exact frames. Each cell has the label `HH:MM:SS.mmm`. With `transcript_job_id`, each cell also shows the words spoken at that instant, with spans and context. `layout: frames` = one native-size labeled picture per instant. `every_frame` = every native frame of ≤ 4 s | Modal compute |
| `cut` | `remix.cut/1`: an MP4 of `[start, end)` or joined `keep[]` spans (≤ 120 s), with `mapping[]`. When `label_time` is set, the cut burns the source time into each frame | Modal compute |
| `notes` | `remix.notes/1`: `ANALYSIS.md`, `TIMELINE.md`, `PROGRESS.md` (the read of the reference) or `FORMAT.md`, `TREATMENT.md` (the project: why the format works and which spoken words its captions attach to, plus the target treatment) as a text artifact. A light lint applies: evidence links must be `media.sume.com` artifacts, and `FORMAT.md` must have a `## Captions` section | unbilled |
| `align` | `remix.take/1` (#8076 steps 2–5): one `script.svml` segment aligned to a **generated take**'s transcript (the take's own `probe` + `transcribe`). The output has every script token with the measured `start_seconds` / `end_seconds` / `score` of the transcript words that it paired with (an authored eojeol can pair with two heard words). It also has `cues[]` from `\|\|` breaks and role changes, `selections{}` and `moments{}` resolved to word boundaries, and `unmatched[]` for words that the take did not say. No argument carries a time. An unmatched word stays untimed. `take.json` + `script.svml` are artifacts, and the bundle lists `takes[]` | unbilled |

To select samples on `tile`, use one of these: `start`/`end` with `every`
(seconds) or `frames` (evenly spaced midpoints, approximately 1.5 a second by
default, at least 4, at most 9), `at[]` explicit instants,
`around {phrase, occurrence, padding}` resolved on the transcript, or
`ranges[]` for several stretches. The verb sizes the pages so that the agent
host's 1080 px inline downscale does not make a cell smaller. For example, a
480 px portrait cell stays 480 px, with two cells on each page. To trade size
for count, pass `cell` and `columns`.

Every image, clip, and json is a durable `media.sume.com` artifact that the
job's `artifacts[]` lists. Thus, a later step binds it by id. For
idempotency, the host tools derive the key from the verb and its arguments.
On REST, send `Idempotency-Key`.

## Compose (steps 2–5, Sume Agent host only)

`sume-agent__sume_compose` takes `composition.svml`. This file has one
`<take id job src/>` for each aligned segment. It binds
`<video>` / `<image>` / `<audio>` / `<text>` with
`during` / `at` + `for` / `until` + `for` / `start` + `end`. The bind targets
are `story.selection.<name>`, `story.moment.<name>`, `story.segment.<id>`,
`take.<id>`, or `program`. The clock literals are `2s` / `12f` / `250ms` with
`±` offsets, and the tool refuses a literal alone in `during`. The file also
has `<captions>` (a kit face, size, colors, position,
`karaoke: progress | active | off`).

The tool compiles the HyperFrames document. In this document, every cue is a
`data-sume-caption` root with attribute-timed word spans. The tool runs the
document on the current compose engine. `output: "check"` = free lint + a
Modal-metered Chromium audit with a `checked_draft_ref` for
`hyperframes_snapshot`. `output: "mp4"` = the billed `video_caption` bake,
with the thread composition revision on the Studio timeline. The result's
`svml` receipt lists every take, cue, and anchor with the spoken word that it
resolved through.

On dest, `sume/remix` carries the order: project notes →
produce with the current generation tools → measure each take → SVML → align →
check → snapshot → bake.

## Errors that name the fix

| Code | Meaning |
|---|---|
| `source_too_long_for_remix_media` | The clip is longer than 300 s. |
| `source_no_video_stream` | The file has no video stream. |
| `remix_understanding_not_found` | `understanding_id` is not a completed probe of this workspace. |
| `remix_tile_too_many_cells` | The selectors resolve to more than 96 cells. Increase `every`, decrease `frames`, or split `ranges[]`. |
| `remix_transcript_not_found` | `transcript_job_id` is not a completed transcribe of this understanding in this workspace. |
| `remix_tile_transcript_required` | The request has `around` without `transcript_job_id`. |
| `remix_transcribe_checkpoint_conflict` | The request sent `checkpoint` with `engine: scribe_v2`. The checkpoint is a WhisperX knob. |
| `remix_align_script_invalid` | `script.svml` did not parse. The message carries the `remix_svml_*` code and line (a marker inside a word, a selection across segments, a reused name). |
| `remix_align_segment_unknown` | `segment` is not a segment of the script. |
| `remix_svml_*` (compose) | `composition.svml` has a problem: an unknown selection or moment, a take without an align job, a clock literal in `during`, a non-kit font, or a non-`media.sume.com` source. |
| `remix_media_unavailable` | The media workers are not deployed in this environment. |
| `ffmpeg_fields_rejected` | The request has `vf` / `filtergraph` / `codec` / argv fields. Sume compiles every pass. |

## Not this surface

| Need | Use |
|---|---|
| Deterministic shots + OCR text tracks + audio facts in one call | [Reference ingest](/models/reference-ingest) |
| A quick probe or unlabeled stills | [Video inspect](/models/video-inspect) |
| A production trim | [Video trim](/models/video-trim) |
