Generate avatar video
Avatar videos change a ready avatar into a script-driven talking video. The canonical Avatar 1.0 route is preferred:
Launch requests use the top-level avatar_handle (or per-scene character
fields inside video_inputs) to identify a ready avatar. Provide one of
script or video_inputs, not both.
Sume accepts scripts and multi-scene plans when it estimates the target video duration at 4-60 seconds inclusive. Make longer scripts shorter, or split them into multiple jobs.
Create an avatar video
Create an avatar talking video
POST /v1/avatar-1.0/talking-video
Required
Optional product and scene inputs
- Omit
product_imagefor a productless avatar video. - Use
scene: { "type": "prompt", "prompt": "..." }for scene direction. - Use
scene: { "type": "photo", "image_url": "https://..." }for a photo scene reference. - Media fields must be fetchable public HTTPS URLs. Refer to Media inputs.
Quality
Avatar Video accepts quality: "standard" | "plus" | "max".
| Value | Behavior |
|---|---|
plus | Default when omitted. Balanced quality path. |
standard | Fastest Sume execution path. |
max | Highest quality tier. Slower turnaround. |
Aspect ratio and resolution
aspect_ratiosupports1:1,3:4,9:16,4:3, and16:9. Default:9:16.resolutionis720pat this time.
video_inputs
For scene hooks, demos, or silence beats in one composed video, use ordered
video_inputs, not a single script. The total planned duration must still be
in the 4-60 second window.
voice.type: "silence" is a beat with no speech. duration is required, and
script / input_text are not permitted. Spoken scenes still use
type: "text" with one of script or input_text, not both.
The current execution supports one resolved avatar for each final video. It also expects that the scene backgrounds resolve to one shared scene.
Inline captions
After generation, the optional captions burns styles into the clean final
MP4. It uses the spoken script / video_inputs text. Sume never adds captions
to preview stills.
captionstakes the same four settings as standalone Video captions:style, optionalfont, alanguagehint andscript_text.- Styles:
slam(default),punch,tiktok-green,korean-ad(Hangul karaoke for Korean speech), and the Hangul identitiesweight-shift,black-outline,highlight,pill-karaoke,clip-wipeandeditorial-emphasis. - If a Korean script uses
style: "slam"(orpunch/tiktok-green), Sume rejects it with400 caption_hangul_text_latin_style. Sume does not change the style. Those font faces render Hangul as tofu. For Korean speech, select a Hangul style. - For inline captions, Sume rejects an estimated duration of more than 60 seconds.
- A failure in the caption stage is a soft failure. The avatar job can still
succeed with a clean primary
video_urlandcaptions.status=failed. - Inline captions do not create a separate billed video-caption job. To add captions to a public video URL that you already have, use Video captions.
Poll and recover
Completed results can include public media.sume.com video artifacts. They can
also include public-safe preview fields, for example preview_image_url and
scene_previews.
Read avatar-video resources
Preview first, then generate
To review first-frame stills before you pay for a full render, create an
Avatar video preview. Then call
generate-video on the preview id.
Compatibility aliases
| Alias | Notes |
|---|---|
POST /v1/models/sume/avatar-1.0/talking-video/runs | Canonical model-run alias. For new integrations, /v1/avatar-1.0/talking-video is preferred. |
POST /v1/models/sume/avatar-video/v1.0/runs | Legacy launch alias. Same body contract. |