Video captions
Standalone video captions use a public HTTPS video URL as input. The API runs a
job and returns a captioned video. We recommend this API when you already have a
finished clip. For Avatar Video, you can also set inline captions on
talking-video. Or you can store the caption intent on
previews.
Create a caption job
You must send video_url. These fields are optional: style, font,
language, script_text, words, cues / segments (authored overlay text,
without speech-to-text), and the usual mode / webhook_url /
wait_timeout_seconds communication fields.
Speech-to-captions (script_text, or STT if you do not send it) works only when
the clip has audible speech. If the clip is silent, the job fails with
caption_no_speech (next_action: use_overlay_captions). This error is not a
generic policy rejection. To burn authored overlay text without ASR, send cues
(or segments) with text + start + end.
Create a video caption job
POST /v1/video-captions
Required
Style, font and language
| Field | Values / default |
|---|---|
style | If you do not send it, the text of the captions sets the style: slam for Latin, black-outline for Korean. Or select slam, punch, tiktok-green, korean-ad (CapCut-style Hangul karaoke for Korean speech: one short phrase at a time, in the lower third, and the spoken word changes to a heavy weight. Use it with language: "ko". Korean script_text is supported), or one of the Hangul identities below |
font | Optional Hangul face for the style that you select. If you do not send it, the style keeps its own face. Only for Hangul styles. Refer to Fonts. |
design | Optional overrides for one request. They change the colors, typography, placement, phrasing, and motion of the style. Refer to Design overrides. |
language | Speech-to-text hint (ko, en, …). If you do not send it, speech-to-text finds the language automatically. |
style selects the look and the motion. design changes that look. font
selects the face of the text. language only tells speech-to-text which language
to expect. language never selects the style or the font.
If you do not send style, the text of the captions sets the style. Korean text
resolves to black-outline, and not to the Latin default, which burns Korean
text as tofu. But if you select a style, Sume renders that style.
Design overrides
A style is a set of design tokens. design overrides these tokens for one
request. Each field is optional. Sume merges each field over the value of the
style. Thus, one key changes one thing:
This request burns black-outline with a cyan emphasis. All the other values of
the style stay the same.
| Group | Fields |
|---|---|
colors | base, active, stroke, accent, accent_deep, card (null shows no card) |
typography | base_weight, active_weight, active_scale, font_size_ratio, safe_width_ratio, stroke_width_px |
placement | anchor_ratio, landscape_anchor_ratio: the center of the line as a fraction of the frame height |
phrasing | max_words, max_chars, pause_seconds |
motion | enter_seconds, exit_seconds, emphasis_in_seconds, emphasis_out_seconds |
Colors are hex, rgb()/rgba(), or transparent. Sume rejects other CSS syntax
and does not put it into the render document. A number outside its documented
range gives a 400. Thus, an incorrect look fails at request time, and you do
not pay for an incorrect render. punch and tiktok-green do not support
design. These styles still render on a path that reads none of these tokens.
Korean text on a Latin style is rejected
slam, punch, and tiktok-green use Latin display faces that have no Hangul
glyphs. If you send Korean text to one of these styles, the API returns 400
(caption_hangul_text_latin_style). The API does not change the style of the
job. A caption style that you did not select is worse than an error. The
alternative is a video full of tofu boxes, and that video has the same price as a
good video. Latin text on slam works as before.
This rule is also applicable to font. The faces below are Hangul faces. If you
send one of these faces with a Latin style, the API returns 400
(caption_font_requires_hangul_style).
Hangul caption identities
For Korean speech (talking head, UGC, creator voiceover), use one of these
styles and not slam. The Latin display face of that style has no Hangul glyphs
and renders Korean as tofu. These styles put words into phrase cards and do not show
one word at a time. They burn the transcript accurately as written, with no case
folding.
| Style | Look |
|---|---|
black-outline | CapCut white fill on a thick black outline, in the middle of the frame. The safe default. |
weight-shift | Phrase cards. The spoken word gets the heavy weight, and the other words go back to a light weight. |
highlight | An accent block moves in behind the word that the voice speaks. |
pill-karaoke | The card is in a dark pill. The color changes with the voice. |
clip-wipe | Each word comes in with a wipe from left to right. The most clear style at small phone sizes. |
editorial-emphasis | Left-aligned card with two lines. The lead words stay small. The last word of the phrase moves to a second line at approximately two times the size, in a display face. That word moves in from the margin. |
korean-ad is the ad karaoke look (weight-shift, with an accent color on the
spoken word). If you do not send style, the style does not resolve to
this look. It resolves to black-outline. Thus, if you want the ad look,
select korean-ad.
If you do not send style, the render also has no accent color. In their
original design, the two default styles have a gold tint on the spoken word. Sume
does not add a tint to a look that you did not select. Thus, a default render
keeps the fill color on the spoken word. If you select black-outline (or
slam), that style keeps its own gold tint. design.colors.active sets the tint
in the two cases.
Fonts
This field is optional. If you do not send font, the style keeps its own face.
The face is Pretendard for korean-ad, weight-shift, highlight,
pill-karaoke, and editorial-emphasis. The face is Do Hyeon for
black-outline and clip-wipe. If you set it, the style stays the same and the
face changes.
editorial-emphasis always shows its emphasis line in Black Han Sans, for all
values of font. The look of this identity is the contrast between its two faces,
and not one of the two faces. font changes its lead
line, as it changes the single line in all the other styles.
font | Family | Feel |
|---|---|---|
pretendard | Pretendard Variable | Clean baseline |
do-hyeon | Do Hyeon | Thick rounded, the CapCut classic |
black-han-sans | Black Han Sans | Impact |
jua | Jua | Soft cute rounded |
dunggeunmo | DungGeunMo | Pixel / retro |
bagel-fat-one | Bagel Fat One | Fat rounded |
dongle | Dongle Bold | Playful rounded display |
gasoek-one | Gasoek One | Ultra-thick impact |
yeon-sung | Yeon Sung | Brushy |
single-day | Single Day | Soft cute handwritten |
hi-melody | Hi Melody | Soft rounded cute |
nanum-pen | Nanum Pen Script | Handwritten |
gowun-dodum | Gowun Dodum | Soft editorial |
gmarket-sans | Gmarket Sans | Geometric retail display, the CapCut / YouTube title staple |
noto-sans-kr | Noto Sans KR | Workhorse sans |
noto-serif-kr | Noto Serif KR | Workhorse serif |
ibm-plex-sans-kr | IBM Plex Sans KR | Workhorse sans |
gothic-a1 | Gothic A1 | Workhorse sans |
hahmlet | Hahmlet | Display |
song-myung | Song Myung | Display |
poor-story | Poor Story | Script |
gamja-flower | Gamja Flower | Script |
stylish | Stylish | Display |
sunflower | Sunflower | Display |
nanum-gothic | Nanum Gothic | Workhorse sans |
nanum-myeongjo | Nanum Myeongjo | Workhorse serif |
gaegu | Gaegu | Script |
cute-font | Cute Font | Script |
east-sea-dokdo | East Sea Dokdo | Script |
All of these fonts use the SIL Open Font License 1.1, and the renderer includes them. The API rejects all other names and does not use a replacement. Thus, a caption never uses a fallback face that you did not select.
weight-shift and korean-ad animate the wght axis. Only Pretendard has this
axis. On a static face, these styles keep their color and scale emphasis, but
they do not have the weight travel.
Restyling without transcribing twice
To burn the same video again with a different style, send source_caption_id
and not video_url:
Sume uses the source video of that caption again, with the word timings that it
already has. Thus, speech-to-text does not run a second time. Send words with
it only to correct the text. The price does not change, because a restyle is
still a render.
You can also send words (word-level) or cues / segments (phrase-level
overlay cards) with text, start, end in seconds. With these forms, Sume
does not run speech-to-text. Sume burns that text accurately at those times. This
is the path for silent clips. You can send only one of script_text, words,
cues, and segments.
script_text
If you send this field, Sume keeps the speech-to-text word timings as the source of truth for time. Sume aligns the burned-in text to your script. The alignment can fail with these typed public job errors:
script_alignment_mismatchscript_alignment_failed
Recommended next action: simplify_script_text_or_omit. To burn the STT text,
do not send script_text.
Pricing note
Each accepted standalone caption job reserves and captures $0.20 USD of
Sume usage. This price is for videos of maximum 60 seconds, under the current
fixed estimate. For the live price, refer to GET /v1/catalog and OpenAPI.
Inline Avatar Video captions are a separate add-on in the avatar-video estimate.
They do not create a video_caption resource.
Poll and read
When the job is complete, the resource returns the public-safe status, the
style, and the captioned video_url / artifacts. Raw transcripts, renderer
internals, and signed source URLs are not part of the public contract.
Constraints
video_urlmust be a public HTTPS video URL that Sume can fetch.- The API rejects localhost, private-network, non-HTTPS, signed/private, and provider task URLs.
- The API does not support SRT uploads and provider task IDs. To send
phrase-level text, use
cues/segments.
Related
- Generate avatar video (inline captions)
- Avatar video previews
- Media inputs
- Jobs and results