Video captions

Standalone video captions use a public HTTPS video URL as input. The API runs a job and returns a captioned video. We recommend this API when you already have a finished clip. For Avatar Video, you can also set inline captions on talking-video. Or you can store the caption intent on previews.

Create a caption job

You must send video_url. These fields are optional: style, font, language, script_text, words, cues / segments (authored overlay text, without speech-to-text), and the usual mode / webhook_url / wait_timeout_seconds communication fields.

Speech-to-captions (script_text, or STT if you do not send it) works only when the clip has audible speech. If the clip is silent, the job fails with caption_no_speech (next_action: use_overlay_captions). This error is not a generic policy rejection. To burn authored overlay text without ASR, send cues (or segments) with text + start + end.

Create a video caption job

POST /v1/video-captions

Required

Style, font and language

FieldValues / default
styleIf you do not send it, the text of the captions sets the style: slam for Latin, black-outline for Korean. Or select slam, punch, tiktok-green, korean-ad (CapCut-style Hangul karaoke for Korean speech: one short phrase at a time, in the lower third, and the spoken word changes to a heavy weight. Use it with language: "ko". Korean script_text is supported), or one of the Hangul identities below
fontOptional Hangul face for the style that you select. If you do not send it, the style keeps its own face. Only for Hangul styles. Refer to Fonts.
designOptional overrides for one request. They change the colors, typography, placement, phrasing, and motion of the style. Refer to Design overrides.
languageSpeech-to-text hint (ko, en, …). If you do not send it, speech-to-text finds the language automatically.

style selects the look and the motion. design changes that look. font selects the face of the text. language only tells speech-to-text which language to expect. language never selects the style or the font.

If you do not send style, the text of the captions sets the style. Korean text resolves to black-outline, and not to the Latin default, which burns Korean text as tofu. But if you select a style, Sume renders that style.

Design overrides

A style is a set of design tokens. design overrides these tokens for one request. Each field is optional. Sume merges each field over the value of the style. Thus, one key changes one thing:

This request burns black-outline with a cyan emphasis. All the other values of the style stay the same.

GroupFields
colorsbase, active, stroke, accent, accent_deep, card (null shows no card)
typographybase_weight, active_weight, active_scale, font_size_ratio, safe_width_ratio, stroke_width_px
placementanchor_ratio, landscape_anchor_ratio: the center of the line as a fraction of the frame height
phrasingmax_words, max_chars, pause_seconds
motionenter_seconds, exit_seconds, emphasis_in_seconds, emphasis_out_seconds

Colors are hex, rgb()/rgba(), or transparent. Sume rejects other CSS syntax and does not put it into the render document. A number outside its documented range gives a 400. Thus, an incorrect look fails at request time, and you do not pay for an incorrect render. punch and tiktok-green do not support design. These styles still render on a path that reads none of these tokens.

Korean text on a Latin style is rejected

slam, punch, and tiktok-green use Latin display faces that have no Hangul glyphs. If you send Korean text to one of these styles, the API returns 400 (caption_hangul_text_latin_style). The API does not change the style of the job. A caption style that you did not select is worse than an error. The alternative is a video full of tofu boxes, and that video has the same price as a good video. Latin text on slam works as before.

This rule is also applicable to font. The faces below are Hangul faces. If you send one of these faces with a Latin style, the API returns 400 (caption_font_requires_hangul_style).

Hangul caption identities

For Korean speech (talking head, UGC, creator voiceover), use one of these styles and not slam. The Latin display face of that style has no Hangul glyphs and renders Korean as tofu. These styles put words into phrase cards and do not show one word at a time. They burn the transcript accurately as written, with no case folding.

StyleLook
black-outlineCapCut white fill on a thick black outline, in the middle of the frame. The safe default.
weight-shiftPhrase cards. The spoken word gets the heavy weight, and the other words go back to a light weight.
highlightAn accent block moves in behind the word that the voice speaks.
pill-karaokeThe card is in a dark pill. The color changes with the voice.
clip-wipeEach word comes in with a wipe from left to right. The most clear style at small phone sizes.
editorial-emphasisLeft-aligned card with two lines. The lead words stay small. The last word of the phrase moves to a second line at approximately two times the size, in a display face. That word moves in from the margin.

korean-ad is the ad karaoke look (weight-shift, with an accent color on the spoken word). If you do not send style, the style does not resolve to this look. It resolves to black-outline. Thus, if you want the ad look, select korean-ad.

If you do not send style, the render also has no accent color. In their original design, the two default styles have a gold tint on the spoken word. Sume does not add a tint to a look that you did not select. Thus, a default render keeps the fill color on the spoken word. If you select black-outline (or slam), that style keeps its own gold tint. design.colors.active sets the tint in the two cases.

Fonts

This field is optional. If you do not send font, the style keeps its own face. The face is Pretendard for korean-ad, weight-shift, highlight, pill-karaoke, and editorial-emphasis. The face is Do Hyeon for black-outline and clip-wipe. If you set it, the style stays the same and the face changes.

editorial-emphasis always shows its emphasis line in Black Han Sans, for all values of font. The look of this identity is the contrast between its two faces, and not one of the two faces. font changes its lead line, as it changes the single line in all the other styles.

fontFamilyFeel
pretendardPretendard VariableClean baseline
do-hyeonDo HyeonThick rounded, the CapCut classic
black-han-sansBlack Han SansImpact
juaJuaSoft cute rounded
dunggeunmoDungGeunMoPixel / retro
bagel-fat-oneBagel Fat OneFat rounded
dongleDongle BoldPlayful rounded display
gasoek-oneGasoek OneUltra-thick impact
yeon-sungYeon SungBrushy
single-daySingle DaySoft cute handwritten
hi-melodyHi MelodySoft rounded cute
nanum-penNanum Pen ScriptHandwritten
gowun-dodumGowun DodumSoft editorial
gmarket-sansGmarket SansGeometric retail display, the CapCut / YouTube title staple
noto-sans-krNoto Sans KRWorkhorse sans
noto-serif-krNoto Serif KRWorkhorse serif
ibm-plex-sans-krIBM Plex Sans KRWorkhorse sans
gothic-a1Gothic A1Workhorse sans
hahmletHahmletDisplay
song-myungSong MyungDisplay
poor-storyPoor StoryScript
gamja-flowerGamja FlowerScript
stylishStylishDisplay
sunflowerSunflowerDisplay
nanum-gothicNanum GothicWorkhorse sans
nanum-myeongjoNanum MyeongjoWorkhorse serif
gaeguGaeguScript
cute-fontCute FontScript
east-sea-dokdoEast Sea DokdoScript

All of these fonts use the SIL Open Font License 1.1, and the renderer includes them. The API rejects all other names and does not use a replacement. Thus, a caption never uses a fallback face that you did not select.

weight-shift and korean-ad animate the wght axis. Only Pretendard has this axis. On a static face, these styles keep their color and scale emphasis, but they do not have the weight travel.

Restyling without transcribing twice

To burn the same video again with a different style, send source_caption_id and not video_url:

Sume uses the source video of that caption again, with the word timings that it already has. Thus, speech-to-text does not run a second time. Send words with it only to correct the text. The price does not change, because a restyle is still a render.

You can also send words (word-level) or cues / segments (phrase-level overlay cards) with text, start, end in seconds. With these forms, Sume does not run speech-to-text. Sume burns that text accurately at those times. This is the path for silent clips. You can send only one of script_text, words, cues, and segments.

script_text

If you send this field, Sume keeps the speech-to-text word timings as the source of truth for time. Sume aligns the burned-in text to your script. The alignment can fail with these typed public job errors:

  • script_alignment_mismatch
  • script_alignment_failed

Recommended next action: simplify_script_text_or_omit. To burn the STT text, do not send script_text.

Pricing note

Each accepted standalone caption job reserves and captures $0.20 USD of Sume usage. This price is for videos of maximum 60 seconds, under the current fixed estimate. For the live price, refer to GET /v1/catalog and OpenAPI.

Inline Avatar Video captions are a separate add-on in the avatar-video estimate. They do not create a video_caption resource.

Poll and read

When the job is complete, the resource returns the public-safe status, the style, and the captioned video_url / artifacts. Raw transcripts, renderer internals, and signed source URLs are not part of the public contract.

Constraints

  • video_url must be a public HTTPS video URL that Sume can fetch.
  • The API rejects localhost, private-network, non-HTTPS, signed/private, and provider task URLs.
  • The API does not support SRT uploads and provider task IDs. To send phrase-level text, use cues / segments.