Avatar video previews

Avatar video previews generate the first-frame still stage. They do not start the full talking-video render. Use them when you want to approve the composition before you pay for a full Avatar Video generation.

When to use

  • Review the scene composition / first frames before a full render.
  • Multi-scene video_inputs where you want one still per scene.
  • Store caption intent on create. Then apply captions only at generate-video time (Sume never burns captions into preview stills).

For a direct full render without the preview stage, use Generate avatar video.

Create a preview

The create body has the same fields as Avatar Video. Provide one of script or video_inputs, not both. You can also provide the optional product_image, scene, quality, aspect_ratio, title, and captions.

If you omit quality, the default is plus (standard | plus | max).

Create an avatar-video preview

POST /v1/avatar-video-previews

Required

The response includes the URLs to poll the job and an avatar_video_preview_id. Poll the job as for any other generation:

Read the preview resource

When the preview is ready, the public-safe fields include:

  • preview_image_url — primary still (scene 0 for multi-scene).
  • scene_previews[] — one still per input scene when available.
  • resource_status / job_status — these are preferred to the legacy status field for readiness vs job status.

For multi-scene previews with a shared scene, later scene stills are pose-anchored continuations of the first frame.

Regenerate stills

Reuse the stored preview request (avatar, script/video_inputs, scene, quality, aspect ratio), and refresh only the first-frame stills:

Regenerate preview stills

POST /v1/avatar-video-previews/{id}/regenerate

Required

The call returns the same avatar_video_preview_id and a new preview-only job.

Generate the final video

When the preview is correct, start the normal Avatar Video workflow from the preview id. When the preview first frame is available, Sume uses it again. The captions that you stored on preview create apply at this step.

An empty body (or {}) keeps the quality that you selected at preview create. The optional quality overrides only the final render tier. Preview stills are tier-independent, and Sume always uses them again. Thus, if you approve and then downgrade/upgrade, you do not need a new preview.

Generate video from preview

POST /v1/avatar-video-previews/{id}/generate-video

Required

Admission, pre-spend, ledger reservation, provider submit, and readback all use the effective (overridden) tier. Changes to structural fields (script, video_inputs, avatar_handle, scene, aspect_ratio) still need a new preview.

Poll the returned job. Then read the avatar-video resource:

Constraints

  • Same duration window as Avatar Video: estimated 4-60 seconds inclusive.
  • Media inputs are URL-first public HTTPS fields (product_image, scene.image_url, and any scene background image URLs).
  • Sume stores inline captions from preview create for generate-video. Sume does not burn them into preview stills.
  • Preview stills are tier-independent. generate-video quality changes only the provider tier of the final video.
  • Exact request/response schemas: live OpenAPI.