Generate avatar video

Avatar videos change a ready avatar into a script-driven talking video. The canonical Avatar 1.0 route is preferred:

Launch requests use the top-level avatar_handle (or per-scene character fields inside video_inputs) to identify a ready avatar. Provide one of script or video_inputs, not both.

Sume accepts scripts and multi-scene plans when it estimates the target video duration at 4-60 seconds inclusive. Make longer scripts shorter, or split them into multiple jobs.

Create an avatar video

Create an avatar talking video

POST /v1/avatar-1.0/talking-video

Required

Optional product and scene inputs

  • Omit product_image for a productless avatar video.
  • Use scene: { "type": "prompt", "prompt": "..." } for scene direction.
  • Use scene: { "type": "photo", "image_url": "https://..." } for a photo scene reference.
  • Media fields must be fetchable public HTTPS URLs. Refer to Media inputs.

Quality

Avatar Video accepts quality: "standard" | "plus" | "max".

ValueBehavior
plusDefault when omitted. Balanced quality path.
standardFastest Sume execution path.
maxHighest quality tier. Slower turnaround.

Aspect ratio and resolution

  • aspect_ratio supports 1:1, 3:4, 9:16, 4:3, and 16:9. Default: 9:16.
  • resolution is 720p at this time.

video_inputs

For scene hooks, demos, or silence beats in one composed video, use ordered video_inputs, not a single script. The total planned duration must still be in the 4-60 second window.

voice.type: "silence" is a beat with no speech. duration is required, and script / input_text are not permitted. Spoken scenes still use type: "text" with one of script or input_text, not both.

The current execution supports one resolved avatar for each final video. It also expects that the scene backgrounds resolve to one shared scene.

Inline captions

After generation, the optional captions burns styles into the clean final MP4. It uses the spoken script / video_inputs text. Sume never adds captions to preview stills.

  • captions takes the same four settings as standalone Video captions: style, optional font, a language hint and script_text.
  • Styles: slam (default), punch, tiktok-green, korean-ad (Hangul karaoke for Korean speech), and the Hangul identities weight-shift, black-outline, highlight, pill-karaoke, clip-wipe and editorial-emphasis.
  • If a Korean script uses style: "slam" (or punch / tiktok-green), Sume rejects it with 400 caption_hangul_text_latin_style. Sume does not change the style. Those font faces render Hangul as tofu. For Korean speech, select a Hangul style.
  • For inline captions, Sume rejects an estimated duration of more than 60 seconds.
  • A failure in the caption stage is a soft failure. The avatar job can still succeed with a clean primary video_url and captions.status=failed.
  • Inline captions do not create a separate billed video-caption job. To add captions to a public video URL that you already have, use Video captions.

Poll and recover

Completed results can include public media.sume.com video artifacts. They can also include public-safe preview fields, for example preview_image_url and scene_previews.

Read avatar-video resources

Preview first, then generate

To review first-frame stills before you pay for a full render, create an Avatar video preview. Then call generate-video on the preview id.

Compatibility aliases

AliasNotes
POST /v1/models/sume/avatar-1.0/talking-video/runsCanonical model-run alias. For new integrations, /v1/avatar-1.0/talking-video is preferred.
POST /v1/models/sume/avatar-video/v1.0/runsLegacy launch alias. Same body contract.