---
title: hypit Understand
description: Hypit's understand verbs for one reference clip — probe, word-timed transcript with scores, cut candidates, time- and word-labeled grids, labeled excerpts, notes — each a durable artifact with an id. Dest only.
---

> **Dest only (#8076).** `POST /v1/hypit-understand/:verb`, `GET /v1/hypit-understand/:id`
> and the `hypit_*` tools are listed where `SUME_COM_HYPIT_UNDERSTAND_ENABLED`
> allows it (development auto-on, production opt-in). Every name in this lane
> is `hypit-` prefixed and stays off the `sume-*` production packet lane.
> [Reference ingest](/models/reference-ingest) and `video_inspect` are unchanged.

hypit Understand 1.0 returns the artifacts a director's read of a reference
leaves behind, the way a local `/hypit` analysis does: a probe, a
`transcript.json` with per-word confidence, a mechanical `boundaries`
series, time- and word-labeled `tile` grids and `frames`, a time-labeled
`cut` excerpt, and the agent's `ANALYSIS.md` / `TIMELINE.md` / `PROGRESS.md`.
One job type (`hypit_understand`) carries a verb per request; the first
verb, `probe`, mints the **understanding id** (its own job id) and every
later verb binds to it. The compute runs on a separate Modal app
(`hypit-media`) with its own Volume, so the tenth grid of one understanding
never re-downloads the source.

```text
POST /v1/hypit-understand/probe        video_url                              → understanding_id
POST /v1/hypit-understand/transcribe   understanding_id, language, engine? (whisperx default | scribe_v2 + allow_billed_stt: true), checkpoint? (small default)
POST /v1/hypit-understand/boundaries   understanding_id, rate?, threshold?
POST /v1/hypit-understand/tile         understanding_id, layout?, at[] | around | ranges[] | start/end + every|frames, every_frame?, cell?, transcript_job_id?
POST /v1/hypit-understand/cut          understanding_id, start/end | keep[], label_time?
POST /v1/hypit-understand/notes        understanding_id, name (ANALYSIS.md | TIMELINE.md | PROGRESS.md | FORMAT.md | TREATMENT.md), markdown
POST /v1/hypit-understand/align        understanding_id, take_understanding_id, transcript_job_id, script, segment   → hypit.take/1
GET  /v1/hypit-understand/:id          the job; for a probe id also the bundle (hypit.understanding/1, with takes[])
```

Hosted MCP: `hypit_probe`, `hypit_transcribe`, `hypit_boundaries`,
`hypit_tile`, `hypit_cut`, `hypit_notes_write`, `hypit_align`,
`hypit_understand_get`. The Sume Agent host adds `sume-agent__hypit_*`, which
run the same verbs and attach `hypit_tile`'s pages **inline as images** (up
to six per call), plus the host-only `sume-agent__hypit_compose` (below).

## Verbs

## Order: one probe, then batches

Every verb after `probe` needs only the `understanding_id`, and the app runs
them in parallel containers, so an agent should not run the verbs one after
another. Stage 0: `probe` alone. Stage 1, one batch: `transcribe` +
`boundaries` + the overview `tile` calls. Stage 2: dense `tile` calls and
`cut`s chosen from the boundaries and the overview (no transcript needed).
Stage 3, once the transcript is in: `tile` with `around { phrase }` and any
word-labeled close read (`transcript_job_id`). Stage 4: the three `notes`,
then the bundle. Only `around`, word labels and the notes wait on the
transcript; a second `transcribe` is never needed.

| Verb | Produces | Billing |
|---|---|---|
| `probe` | ffprobe facts, 16 kHz wav (`audio.wav_url`), loudness gate (`audio.silent`) | unbilled |
| `transcribe` | `hypit.transcript/1`: `passages[]` → `words[]` with `start_seconds`, `end_seconds`, `score` (0–1). Default `engine: whisperx` — Hypit's own engine upstream (faster-whisper ASR + wav2vec2 alignment on the `hypit-media` box), a score on every aligned word; `engine: scribe_v2` — Sume STT 1.0, no per-word score. Explicit `language`; optional `checkpoint` (`small` default — Hypit's own, int8 on a CPU container; `medium`, `large-v3-turbo`, `large-v3` opt-in on the GPU). There is no `model` field on this route: `model` is the public model id | whisperx unbilled; scribe_v2 = STT per-minute rate on the probe's duration, `allow_billed_stt: true` required, settled to zero when silent |
| `boundaries` | `hypit.boundaries/1`: `candidates[] {at, score}` from a 12 fps 32×32 mean-absolute-difference pass (defaults `rate` 12, `threshold` 0.1) | unbilled |
| `tile` | `hypit.tile/1`: pages of exact frames, each cell labeled `HH:MM:SS.mmm` and, with `transcript_job_id`, the words spoken at that instant with spans and context; `layout: frames` = one native-size labeled picture per instant; `every_frame` = every native frame of ≤ 4 s | unbilled |
| `cut` | `hypit.cut/1`: an MP4 of `[start, end)` or joined `keep[]` spans (≤ 120 s), source time burned per frame when `label_time`, with `mapping[]` | unbilled |
| `notes` | `hypit.notes/1`: `ANALYSIS.md`, `TIMELINE.md`, `PROGRESS.md` (the reference reading) or `FORMAT.md`, `TREATMENT.md` (the project: why the format works and which spoken words its captions stick to; the target treatment) as a text artifact (light lint: evidence links must be `media.sume.com` artifacts; `FORMAT.md` needs a `## Captions` section) | unbilled |
| `align` | `hypit.take/1` (#8076 steps 2–5): one `script.svml` segment aligned to a **generated take**'s transcript (the take's own `probe` + `transcribe`): every script token with the measured `start_seconds` / `end_seconds` / `score` of the transcript words it paired with (an authored eojeol may pair with two heard words), `cues[]` from `\|\|` breaks and role changes, `selections{}` and `moments{}` resolved to word boundaries, `unmatched[]` for words the take did not say. No argument carries a time; an unmatched word stays untimed. `take.json` + `script.svml` are artifacts and the bundle lists `takes[]` | unbilled |

Sample selection on `tile`: `start`/`end` with `every` (seconds) or `frames`
(evenly spaced midpoints; default about 1.5 a second, at least 4, at most 9),
`at[]` explicit instants, `around {phrase, occurrence, padding}` resolved on
the transcript, or `ranges[]` for several stretches. Pages are sized so each
cell survives the agent host's 1080 px inline downscale (a 480 px portrait
cell stays 480 px: two cells per page); pass `cell` and `columns` to trade
size for count.

Every image, clip and json is a durable `media.sume.com` artifact listed on
the job's `artifacts[]`, so a later step binds it by id. Idempotency: the
host tools derive the key from the verb and its arguments; on REST send
`Idempotency-Key`.

## Compose (steps 2–5, Sume Agent host only)

`sume-agent__hypit_compose` takes `composition.svml` — `<take id job src/>`
per aligned segment, `<video>` / `<image>` / `<audio>` / `<text>` bound with
`during` / `at` + `for` / `until` + `for` / `start` + `end` on
`story.selection.<name>`, `story.moment.<name>`, `story.segment.<id>`,
`take.<id>` or `program` (clock literals `2s` / `12f` / `250ms` with `±`
offsets; a literal alone in `during` is refused), and `<captions>` (a kit
face, size, colors, position, `karaoke: progress | active | off`) — compiles
the HyperFrames document (every cue a `data-sume-caption` root with
attribute-timed word spans) and runs it on the existing compose engine:
`output: "check"` = free lint + unbilled Chromium audit with a
`checked_draft_ref` for `hyperframes_snapshot`; `output: "mp4"` = the billed
`video_caption` bake, with the thread composition revision on the Studio
timeline. The result's `svml` receipt lists every take, cue and anchor with
the spoken word it resolved through. The `hypit-compose` packet (dest only)
carries the order: project notes → produce with the existing generation
tools → measure each take → SVML → align → check → snapshot → bake.

## Errors that name the fix

| Code | Meaning |
|---|---|
| `source_too_long_for_hypit_understand` | The clip is over 300 s. |
| `source_no_video_stream` | The file has no video stream. |
| `hypit_understanding_not_found` | `understanding_id` is not a completed probe of this workspace. |
| `hypit_tile_too_many_cells` | The selectors resolve to more than 96 cells; raise `every`, lower `frames` or split `ranges[]`. |
| `hypit_transcript_not_found` | `transcript_job_id` is not a completed transcribe of this understanding in this workspace. |
| `hypit_tile_transcript_required` | `around` without `transcript_job_id`. |
| `hypit_transcribe_checkpoint_conflict` | `checkpoint` sent with `engine: scribe_v2`; the checkpoint is a WhisperX knob. |
| `hypit_align_script_invalid` | `script.svml` did not parse; the message carries the `hypit_svml_*` code and line (a marker inside a word, a selection crossing segments, a reused name). |
| `hypit_align_segment_unknown` | `segment` is not a segment of the script. |
| `hypit_svml_*` (compose) | `composition.svml` problems: an unknown selection or moment, a take without an align job, a clock literal in `during`, a non-kit font, a non-`media.sume.com` source. |
| `hypit_media_unavailable` | The `hypit-media` Modal app is not deployed in this workspace. |
| `ffmpeg_fields_rejected` | `vf` / `filtergraph` / `codec` / argv fields: Sume compiles every pass. |

## Not this surface

| Need | Use |
|---|---|
| Deterministic shots + OCR text tracks + audio facts in one call | [Reference ingest](/models/reference-ingest) |
| A quick probe or unlabeled stills | [Video inspect](/models/video-inspect) |
| A production trim | [Video trim](/models/video-trim) |
