hypit Understand

Dest only (#8076). POST /v1/hypit-understand/:verb, GET /v1/hypit-understand/:id and the hypit_* tools are listed where SUME_COM_HYPIT_UNDERSTAND_ENABLED allows it (development auto-on, production opt-in). Every name in this lane is hypit- prefixed and stays off the sume-* production packet lane. Reference ingest and video_inspect are unchanged.

hypit Understand 1.0 returns the artifacts a director's read of a reference leaves behind, the way a local /hypit analysis does: a probe, a transcript.json with per-word confidence, a mechanical boundaries series, time- and word-labeled tile grids and frames, a time-labeled cut excerpt, and the agent's ANALYSIS.md / TIMELINE.md / PROGRESS.md. One job type (hypit_understand) carries a verb per request; the first verb, probe, mints the understanding id (its own job id) and every later verb binds to it. The compute runs on a separate Modal app (hypit-media) with its own Volume, so the tenth grid of one understanding never re-downloads the source.

Hosted MCP: hypit_probe, hypit_transcribe, hypit_boundaries, hypit_tile, hypit_cut, hypit_notes_write, hypit_align, hypit_understand_get. The Sume Agent host adds sume-agent__hypit_*, which run the same verbs and attach hypit_tile's pages inline as images (up to six per call), plus the host-only sume-agent__hypit_compose (below).

Verbs

Order: one probe, then batches

Every verb after probe needs only the understanding_id, and the app runs them in parallel containers, so an agent should not run the verbs one after another. Stage 0: probe alone. Stage 1, one batch: transcribe + boundaries + the overview tile calls. Stage 2: dense tile calls and cuts chosen from the boundaries and the overview (no transcript needed). Stage 3, once the transcript is in: tile with around { phrase } and any word-labeled close read (transcript_job_id). Stage 4: the three notes, then the bundle. Only around, word labels and the notes wait on the transcript; a second transcribe is never needed.

VerbProducesBilling
probeffprobe facts, 16 kHz wav (audio.wav_url), loudness gate (audio.silent)unbilled
transcribehypit.transcript/1: passages[] → words[] with start_seconds, end_seconds, score (0–1). Default engine: whisperx — Hypit's own engine upstream (faster-whisper ASR + wav2vec2 alignment on the hypit-media box), a score on every aligned word; engine: scribe_v2 — Sume STT 1.0, no per-word score. Explicit language; optional checkpoint (small default — Hypit's own, int8 on a CPU container; medium, large-v3-turbo, large-v3 opt-in on the GPU). There is no model field on this route: model is the public model idwhisperx unbilled; scribe_v2 = STT per-minute rate on the probe's duration, allow_billed_stt: true required, settled to zero when silent
boundarieshypit.boundaries/1: candidates[] {at, score} from a 12 fps 32×32 mean-absolute-difference pass (defaults rate 12, threshold 0.1)unbilled
tilehypit.tile/1: pages of exact frames, each cell labeled HH:MM:SS.mmm and, with transcript_job_id, the words spoken at that instant with spans and context; layout: frames = one native-size labeled picture per instant; every_frame = every native frame of ≤ 4 sunbilled
cuthypit.cut/1: an MP4 of [start, end) or joined keep[] spans (≤ 120 s), source time burned per frame when label_time, with mapping[]unbilled
noteshypit.notes/1: ANALYSIS.md, TIMELINE.md, PROGRESS.md (the reference reading) or FORMAT.md, TREATMENT.md (the project: why the format works and which spoken words its captions stick to; the target treatment) as a text artifact (light lint: evidence links must be media.sume.com artifacts; FORMAT.md needs a ## Captions section)unbilled
alignhypit.take/1 (#8076 steps 2–5): one script.svml segment aligned to a generated take's transcript (the take's own probe + transcribe): every script token with the measured start_seconds / end_seconds / score of the transcript words it paired with (an authored eojeol may pair with two heard words), cues[] from || breaks and role changes, selections{} and moments{} resolved to word boundaries, unmatched[] for words the take did not say. No argument carries a time; an unmatched word stays untimed. take.json + script.svml are artifacts and the bundle lists takes[]unbilled

Sample selection on tile: start/end with every (seconds) or frames (evenly spaced midpoints; default about 1.5 a second, at least 4, at most 9), at[] explicit instants, around {phrase, occurrence, padding} resolved on the transcript, or ranges[] for several stretches. Pages are sized so each cell survives the agent host's 1080 px inline downscale (a 480 px portrait cell stays 480 px: two cells per page); pass cell and columns to trade size for count.

Every image, clip and json is a durable media.sume.com artifact listed on the job's artifacts[], so a later step binds it by id. Idempotency: the host tools derive the key from the verb and its arguments; on REST send Idempotency-Key.

Compose (steps 2–5, Sume Agent host only)

sume-agent__hypit_compose takes composition.svml — <take id job src/> per aligned segment, <video> / <image> / <audio> / <text> bound with during / at + for / until + for / start + end on story.selection.<name>, story.moment.<name>, story.segment.<id>, take.<id> or program (clock literals 2s / 12f / 250ms with ± offsets; a literal alone in during is refused), and <captions> (a kit face, size, colors, position, karaoke: progress | active | off) — compiles the HyperFrames document (every cue a data-sume-caption root with attribute-timed word spans) and runs it on the existing compose engine: output: "check" = free lint + unbilled Chromium audit with a checked_draft_ref for hyperframes_snapshot; output: "mp4" = the billed video_caption bake, with the thread composition revision on the Studio timeline. The result's svml receipt lists every take, cue and anchor with the spoken word it resolved through. The hypit-compose packet (dest only) carries the order: project notes → produce with the existing generation tools → measure each take → SVML → align → check → snapshot → bake.

Errors that name the fix

CodeMeaning
source_too_long_for_hypit_understandThe clip is over 300 s.
source_no_video_streamThe file has no video stream.
hypit_understanding_not_foundunderstanding_id is not a completed probe of this workspace.
hypit_tile_too_many_cellsThe selectors resolve to more than 96 cells; raise every, lower frames or split ranges[].
hypit_transcript_not_foundtranscript_job_id is not a completed transcribe of this understanding in this workspace.
hypit_tile_transcript_requiredaround without transcript_job_id.
hypit_transcribe_checkpoint_conflictcheckpoint sent with engine: scribe_v2; the checkpoint is a WhisperX knob.
hypit_align_script_invalidscript.svml did not parse; the message carries the hypit_svml_* code and line (a marker inside a word, a selection crossing segments, a reused name).
hypit_align_segment_unknownsegment is not a segment of the script.
hypit_svml_* (compose)composition.svml problems: an unknown selection or moment, a take without an align job, a clock literal in during, a non-kit font, a non-media.sume.com source.
hypit_media_unavailableThe hypit-media Modal app is not deployed in this workspace.
ffmpeg_fields_rejectedvf / filtergraph / codec / argv fields: Sume compiles every pass.

Not this surface

NeedUse
Deterministic shots + OCR text tracks + audio facts in one callReference ingest
A quick probe or unlabeled stillsVideo inspect
A production trimVideo trim