hypit Understand
Dest only (#8076).
POST /v1/hypit-understand/:verb,GET /v1/hypit-understand/:idand thehypit_*tools are listed whereSUME_COM_HYPIT_UNDERSTAND_ENABLEDallows it (development auto-on, production opt-in). Every name in this lane ishypit-prefixed and stays off thesume-*production packet lane. Reference ingest andvideo_inspectare unchanged.
hypit Understand 1.0 returns the artifacts a director's read of a reference
leaves behind, the way a local /hypit analysis does: a probe, a
transcript.json with per-word confidence, a mechanical boundaries
series, time- and word-labeled tile grids and frames, a time-labeled
cut excerpt, and the agent's ANALYSIS.md / TIMELINE.md / PROGRESS.md.
One job type (hypit_understand) carries a verb per request; the first
verb, probe, mints the understanding id (its own job id) and every
later verb binds to it. The compute runs on a separate Modal app
(hypit-media) with its own Volume, so the tenth grid of one understanding
never re-downloads the source.
Hosted MCP: hypit_probe, hypit_transcribe, hypit_boundaries,
hypit_tile, hypit_cut, hypit_notes_write, hypit_align,
hypit_understand_get. The Sume Agent host adds sume-agent__hypit_*, which
run the same verbs and attach hypit_tile's pages inline as images (up
to six per call), plus the host-only sume-agent__hypit_compose (below).
Verbs
Order: one probe, then batches
Every verb after probe needs only the understanding_id, and the app runs
them in parallel containers, so an agent should not run the verbs one after
another. Stage 0: probe alone. Stage 1, one batch: transcribe +
boundaries + the overview tile calls. Stage 2: dense tile calls and
cuts chosen from the boundaries and the overview (no transcript needed).
Stage 3, once the transcript is in: tile with around { phrase } and any
word-labeled close read (transcript_job_id). Stage 4: the three notes,
then the bundle. Only around, word labels and the notes wait on the
transcript; a second transcribe is never needed.
| Verb | Produces | Billing |
|---|---|---|
probe | ffprobe facts, 16 kHz wav (audio.wav_url), loudness gate (audio.silent) | unbilled |
transcribe | hypit.transcript/1: passages[] → words[] with start_seconds, end_seconds, score (0–1). Default engine: whisperx — Hypit's own engine upstream (faster-whisper ASR + wav2vec2 alignment on the hypit-media box), a score on every aligned word; engine: scribe_v2 — Sume STT 1.0, no per-word score. Explicit language; optional checkpoint (small default — Hypit's own, int8 on a CPU container; medium, large-v3-turbo, large-v3 opt-in on the GPU). There is no model field on this route: model is the public model id | whisperx unbilled; scribe_v2 = STT per-minute rate on the probe's duration, allow_billed_stt: true required, settled to zero when silent |
boundaries | hypit.boundaries/1: candidates[] {at, score} from a 12 fps 32×32 mean-absolute-difference pass (defaults rate 12, threshold 0.1) | unbilled |
tile | hypit.tile/1: pages of exact frames, each cell labeled HH:MM:SS.mmm and, with transcript_job_id, the words spoken at that instant with spans and context; layout: frames = one native-size labeled picture per instant; every_frame = every native frame of ≤ 4 s | unbilled |
cut | hypit.cut/1: an MP4 of [start, end) or joined keep[] spans (≤ 120 s), source time burned per frame when label_time, with mapping[] | unbilled |
notes | hypit.notes/1: ANALYSIS.md, TIMELINE.md, PROGRESS.md (the reference reading) or FORMAT.md, TREATMENT.md (the project: why the format works and which spoken words its captions stick to; the target treatment) as a text artifact (light lint: evidence links must be media.sume.com artifacts; FORMAT.md needs a ## Captions section) | unbilled |
align | hypit.take/1 (#8076 steps 2–5): one script.svml segment aligned to a generated take's transcript (the take's own probe + transcribe): every script token with the measured start_seconds / end_seconds / score of the transcript words it paired with (an authored eojeol may pair with two heard words), cues[] from || breaks and role changes, selections{} and moments{} resolved to word boundaries, unmatched[] for words the take did not say. No argument carries a time; an unmatched word stays untimed. take.json + script.svml are artifacts and the bundle lists takes[] | unbilled |
Sample selection on tile: start/end with every (seconds) or frames
(evenly spaced midpoints; default about 1.5 a second, at least 4, at most 9),
at[] explicit instants, around {phrase, occurrence, padding} resolved on
the transcript, or ranges[] for several stretches. Pages are sized so each
cell survives the agent host's 1080 px inline downscale (a 480 px portrait
cell stays 480 px: two cells per page); pass cell and columns to trade
size for count.
Every image, clip and json is a durable media.sume.com artifact listed on
the job's artifacts[], so a later step binds it by id. Idempotency: the
host tools derive the key from the verb and its arguments; on REST send
Idempotency-Key.
Compose (steps 2–5, Sume Agent host only)
sume-agent__hypit_compose takes composition.svml — <take id job src/>
per aligned segment, <video> / <image> / <audio> / <text> bound with
during / at + for / until + for / start + end on
story.selection.<name>, story.moment.<name>, story.segment.<id>,
take.<id> or program (clock literals 2s / 12f / 250ms with ±
offsets; a literal alone in during is refused), and <captions> (a kit
face, size, colors, position, karaoke: progress | active | off) — compiles
the HyperFrames document (every cue a data-sume-caption root with
attribute-timed word spans) and runs it on the existing compose engine:
output: "check" = free lint + unbilled Chromium audit with a
checked_draft_ref for hyperframes_snapshot; output: "mp4" = the billed
video_caption bake, with the thread composition revision on the Studio
timeline. The result's svml receipt lists every take, cue and anchor with
the spoken word it resolved through. The hypit-compose packet (dest only)
carries the order: project notes → produce with the existing generation
tools → measure each take → SVML → align → check → snapshot → bake.
Errors that name the fix
| Code | Meaning |
|---|---|
source_too_long_for_hypit_understand | The clip is over 300 s. |
source_no_video_stream | The file has no video stream. |
hypit_understanding_not_found | understanding_id is not a completed probe of this workspace. |
hypit_tile_too_many_cells | The selectors resolve to more than 96 cells; raise every, lower frames or split ranges[]. |
hypit_transcript_not_found | transcript_job_id is not a completed transcribe of this understanding in this workspace. |
hypit_tile_transcript_required | around without transcript_job_id. |
hypit_transcribe_checkpoint_conflict | checkpoint sent with engine: scribe_v2; the checkpoint is a WhisperX knob. |
hypit_align_script_invalid | script.svml did not parse; the message carries the hypit_svml_* code and line (a marker inside a word, a selection crossing segments, a reused name). |
hypit_align_segment_unknown | segment is not a segment of the script. |
hypit_svml_* (compose) | composition.svml problems: an unknown selection or moment, a take without an align job, a clock literal in during, a non-kit font, a non-media.sume.com source. |
hypit_media_unavailable | The hypit-media Modal app is not deployed in this workspace. |
ffmpeg_fields_rejected | vf / filtergraph / codec / argv fields: Sume compiles every pass. |
Not this surface
| Need | Use |
|---|---|
| Deterministic shots + OCR text tracks + audio facts in one call | Reference ingest |
| A quick probe or unlabeled stills | Video inspect |
| A production trim | Video trim |