Models overview

Sume makes job-backed generation families available on the public Developer API. Each family has a primary product URL. Most families also have a model-run alias. Submit a request with an Idempotency-Key. Poll the job, then read the result.

To discover the current production capabilities, use catalog (GET /v1/catalog) and the OpenAPI snapshot. The detailed guides below agree with the production OpenAPI request schemas (api.sume.com / docs snapshot).

Avatar 1.0

Avatar 1.0 is a two-step workflow:

  1. Create a reusable avatar.
  2. Use this avatar to generate talking videos from scripts or multi-scene inputs.

For new integrations, the canonical product routes are preferred:

StepCanonical route
Create avatarPOST /v1/avatar-1.0/generate
List / read avatarsGET /v1/avatar-1.0/avatars, GET /v1/avatar-1.0/avatars/:id
Create talking videoPOST /v1/avatar-1.0/talking-video
List / read videosGET /v1/avatar-videos, GET /v1/avatar-videos/:id

Legacy model-run aliases, for example POST /v1/models/sume/avatar/v1.0/runs and POST /v1/models/sume/avatar-video/v1.0/runs, are still supported for compatibility. Each guide page gives the details.

First, create an avatar from a prompt, a reference photo, or a supported avatar input. Avatar creation is job-backed. Thus, Sume returns a job first. When the job completes, the avatar becomes a reusable resource in your workspace.

When possible, use a stable avatar handle. The handle gives your app or agent a simple name to use again later. Thus, your app or agent does not depend only on a generated id.

After the avatar is ready, send a script (or video_inputs) and the avatar handle to create an avatar video. Each video is also job-backed. Submit the request, and poll or wait for completion. Then read the result URL.

Avatar Video supports quality: "standard" | "plus" | "max". To use the default plus execution path, omit it. Use standard for the fastest path. Use max when quality is more important than turnaround.

GuideUse for
Create your avatarAvatar creation request.
Generate avatar videoTalking video from a ready avatar.
Avatar video previewsFirst-frame stills before a full render. Then generate-video.
Face swap (Beta)Swap a ready avatar face onto a public source video.
Video captionsBurn captions onto a public video URL that you already have.
Video inspectProbe + stills + optional STT of one media.sume.com clip. Default clip inspection on dest and prod.
Reference ingestOne reference clip → ReferenceVideoManifest (frame-exact shots, source-resolution OCR text tracks, audio facts, labeled strip). Sume bills it by its Modal compute (+ STT when it transcribes). Dest first.
Remix mediaOne reference clip → remix media artifacts (probe, word-timed transcript with scores, cut candidates, time- and word-labeled grids, labeled excerpts, notes). Each is a durable artifact with an id. align binds a script segment to the words of a generated take. sume_compose renders word-anchored SVML on HyperFrames. Sume bills it by its Modal compute, WhisperX included (transcribe with engine: scribe_v2 bills STT, notes / align are free, and the final bake bills as a render). Dest only.
Video framesExact stills at at[] / fps from one hosted clip. Source-size artf_ images. Sume bills it by its Modal compute. Always 202.
Video trim[start, end) of one hosted clip → new MP4. $0.02 flat. Material for timeline, not placement.
Audio detachAudio track of one hosted video → durable wav / mp3. $0.01 flat.
Video filterDim / crop / allowlisted pixel graph on one hosted clip → new MP4. $0.02 encode. /check is free.
Timeline 1.0Audio spine + ordered video[] → one MP4. $0.10 / ceil(output minute). The assembly surface.
Timeline composeStill + video in one frame (반배너 / overlay) → one MP4 shot. $0.02 flat.
Timeline audioConcat / split Sume-hosted audio → durable files. $0.01 flat.
Video analysesLegacy vana_ resource. Dest create is 410. Prod still accepts it until #5953 PR-C2.
Trending videosDiscover TikTok trending video metadata for research.

Image, Video, Music, Fabric

These are managed product models. Sume selects the providers. Callers do not send provider queue ids. VEED Fabric 1.0 is public as veed/fabric-1.0.

FamilyPrimary URLGuide
Image 1.0POST /v1/image-1.0/generateImage 1.0
Image APIPOST /v1/imagesImage API
Video 1.0POST /v1/video-1.0/generateVideo 1.0
Video generationPOST /v1/videosVideo generation
Music RouterPOST /v1/music-router/generateMusic Router
Music 1.0 (retiring, resolves through Music Router)POST /v1/music-1.0/generateMusic 1.0
VEED Fabric 1.0POST /v1/veed/fabric-1.0Talking still + audio clips (veed/fabric-1.0)
MiniMax H3 Max Lip SyncPOST /v1/minimax/h3-max/lip-syncSame still + audio body as Fabric, audio 5–14.8 s, list × 1.25 (minimax/h3-max/lip-sync)

Fabric alias migration

Deprecated compatibility aliases that Sume will retire — still supported:

Deprecated routeCorrect call
POST /v1/avatar-1.0/image-to-videoPOST /v1/veed/fabric-1.0
POST /v1/models/sume/avatar-1.0/image-to-video/runsPOST /v1/models/veed/fabric-1.0/runs (or POST /v1/veed/fabric-1.0)

The public model id is veed/fabric-1.0. Both aliases still accept the same body. Thus, current Mobidoo / live-commerce integrations can migrate if they change only the URL.

Send audio_url, measured duration_seconds, and only one visual source. The preferred source is the image_url of the generated, inspected posed still. Use avatar_handle only when the user named that avatar. You cannot send the two sources together. With MCP, the same body is inside payload on the supported avatar-image-to-video_create tool.

Each shot where a person speaks on camera is Fabric with an accepted still + TTS. This rule is applicable to short UGC and presenter ads, testimonials, Recreate beats where a person speaks, and LC / long-form host talk. Video models do not lip-sync to generated TTS or to a later voice-over. Thus, a face that talks is never a video-model clip with narration under it. Wordless beats, B-roll and product motion use Auto image → inspect → Auto video, without Fabric or Stage P.

Model-run aliases use POST /v1/models/sume/<family>-1.0/runs (same request body) for Image 1.0 / Video 1.0 / Music 1.0.

If you do not need an explicit catalog model id, Image 1.0 + routing_preset is the preferred method. Video 1.0 (POST /v1/video-1.0/generate) is deprecated. It maps to sume/auto and ignores routing_preset. Thus, send model: "sume/auto" or an explicit catalog id (for example Seedance). Pin that id on Video generation (POST /v1/videos) or Image API (POST /v1/images). Legacy /v1/video-router/* stays registered as a Sume-envelope alias (refer to Video Router).

Shared job lifecycle

These families return the same job envelope pattern:

  1. Submit → store job.id, status_url, result_url.
  2. Poll the status (or wait with mode: sync / subscribe for a maximum of 30s).
  3. When result_ready is true, fetch /result.
  4. Read Sume-hosted artifacts for media jobs.

Refer to Jobs and results, Generation admission, Media inputs, and API recipes.

Ahead of production OpenAPI

Some more generators (for example STT) can appear on api.dev.sume.com before the production OpenAPI snapshot lists them. Until these paths land in the docs OpenAPI snapshot and the api.sume.com reference, they are not production Developer API surface.

POST /v1/avatar-1.0/fabric

sume/avatar-1.0/fabric is a temporary, test-only route. Sume uses it to compare a different talking-clip backend against POST /v1/veed/fabric-1.0 (legacy POST /v1/avatar-1.0/image-to-video) on identical inputs. It takes the same request body. Thus, a comparison script changes only the path.

Differences from image-to-video:

image-to-videofabric (experimental)
duration_seconds1–3001–15 (rounded up, and the route rejects requests over 15)
speed_tierselects a provider speed tieraccepted and ignored
Price$0.1875/s @720p$0.3024/s @720p

Do not build production integrations on this route:

  • The name fabric is a placeholder test name and will change before general availability.
  • Sume can change the route or remove it completely after the comparison is complete.
  • VEED Fabric 1.0 (veed/fabric-1.0) stays the supported path for talking clips. This includes Live Commerce and the MCP tools.