Video analyses

Current SoT. Do not start new work on this surface.

  • Default on dest and prod: Video inspect — POST /v1/video-inspect (MCP video_inspect) for probe + stills + optional transcript. Probe and stills are free. The default is mode: sync.
  • Dest (api.dev.sume.com): SUME_COM_VIDEO_ANALYSIS_ENABLED=false. POST /v1/video-analyses answers 410 video_analysis_retired. The MCP tools video-analyses_* are not in tools_list. Stored vana_ rows stay readable through GET. For semantic questions / intervals on dest, use video_analyze / video_segment only when those names appear in tools_list (trusted api.dev.sume.com origin + TwelveLabs key, never production).
  • Prod (api.sume.com): create stays on until #5953 PR-C2. Dest 410 is not a production outage.

This page documents the legacy video_analysis / vana_ resource (typed scenes[]). This resource is not trend discovery or virality prediction, and it does not generate video. A reusable Format from clip structure is a remix job, not this analysis.

When create is still enabled, a video-understanding model (TwelveLabs Pegasus) watches the MP4 and segments it into typed scenes. Then ffmpeg extracts at least one JPEG still per second of each scene (keyframes), plus a representative keyframe_url. The current analysis_version is 1.1.

The remote MCP wrappers video-analyses_create / _get / _list obey the same flag. Production lists them until PR-C2, and dest does not list them. Poll the remaining jobs with jobs_wait (usual runtime 3–5 minutes, poll every 30–60s).

Accuracy caveat: longer videos give less accurate results. Short clips give the most reliable results.

Create an analysis job (production until PR-C2)

Dest callers: do not send a POST to this endpoint. Use video inspect, or dest-only video_analyze / video_segment when Sume lists them.

The required field is video_url. The optional fields are max_scenes (2–40, default 24), include_transcript (when true, each scene's audio carries the spoken speech for that scene, and if not, audio is null), plus the usual mode / webhook_url / wait_timeout_seconds communication fields.

A successful submit returns 202 with a vana_… resource id, a video_analysis job (request_id), and status_url / result_url.

If possible, use a durable media.sume.com URL from media imports (POST /v1/media-imports) or a completed Sume generation artifact. This surface has no Higgsfield-style media_id intake gate.

Pricing note

Each accepted analysis reserves and captures $0.30 USD of Sume usage under the current fixed estimate. Confirm the live price in GET /v1/catalog and OpenAPI.

Duration limits

CapBehavior
Soft (~90s)The job succeeds. The response can include the warning low_confidence_long_video.
Hard (300s)Sume rejects the job after the probe with a stable duration error.

Unsupported inputs

InputResult
YouTube URLs422 unsupported_source
Non-HTTPS / private / localhostSume rejects them at admit
Non-video assetsThey fail with a stable, agent-legible code

Poll and read

When the resource is ready, it includes whole-video metadata plus a typed scenes[] array. Each scene has contiguous start_seconds / end_seconds, duration_seconds, summary, visual, optional shot_type / camera_motion, on_screen_text, and confidence (0–1). Each scene also has audio ({ speech, has_speech, has_music } when the request asked for include_transcript, otherwise null).

Each scene also has keyframe_url, the representative JPEG still. This still is the 1s sample nearest 40% of the scene, or the historical 40% pick on sub-1s scenes. The value can also be null with a keyframe_mirror_failed:scene_N warning. The keyframes field is an array of { t, url } stills for every whole second of that scene. A failed second is url: null plus keyframe_mirror_failed:scene_N:t_T, and it does not drop the scene.

The 300s hard duration cap is also the stills ceiling (~300 JPEGs). This surface has no separate 1-second text/context track.

Scopes

ScopeOperations
video_analyses:writePOST /v1/video-analyses
video_analyses:readGET /v1/video-analyses, GET /v1/video-analyses/:id