Video Generation
Sume supports video generation from text prompts (and optional reference
images) through a dedicated asynchronous API. To find the supported models,
their capabilities, and their prices, use GET /v1/videos/models.
This surface agrees field-for-field with the OpenRouter Video Generation API. Thus, a client that you wrote from their docs works here after you change the base URL and the API key. The few differences of Sume are in Sume differences.
Model Discovery
You can find video generation models in these ways:
Via the Video Models API
To get a list of all available video generation models and their supported parameters, use the dedicated video models endpoint:
The response returns a data array. Each model in the array includes these
fields:
| Field | Description |
|---|---|
id | Model slug to use in generation requests |
canonical_slug | Permanent model identifier |
supported_resolutions | List of supported output resolutions (for example, 720p, 1080p) |
supported_aspect_ratios | List of supported aspect ratios (for example, 16:9, 9:16) |
supported_sizes | List of supported pixel dimensions (for example, 1280x720), or null |
supported_durations | Supported video lengths in whole seconds |
supported_frame_images | The frame_type values that the model accepts |
supported_input_references | The input_references types that the model accepts |
generate_audio | Shows if the model can generate an audio track |
seed | Shows if the model accepts a seed |
pricing_skus | Price information for each SKU |
allowed_passthrough_parameters | Provider-specific parameters that you can send through the provider option |
Before you submit a generation request, use this endpoint to find the
resolutions, aspect ratios, and durations that each model supports. Limits are
different for each model. seedance-2.5 accepts 4–30 seconds at
480p/720p/1080p. wan-3.0 accepts 2–30 seconds. higgsfield-genjutsu (Motion
Transfer: one source video plus 1–8 reference images, 480p/720p, in the catalog
only when its provider is configured) accepts 4–30 seconds.
H3 Max Recast (h3-max-recast: one source video plus 1–4 person photos,
768p/1080p, optional prompt) accepts 5–30 seconds. minimax-h3 accepts 5–15
seconds at native 480p/768p (768p is first-class, not 720p). For this model,
Sume bills 2K/4K upscales if your request includes them. Each other catalog
model has a maximum of 15 seconds. seedance-2 also has 1080p.
MiniMax H3 Max (minimax-h3-max) is the faster 768p variant. It does
text-to-video, first/last-frame image-to-video, and reference-to-video at
480p/768p/1080p (1080p is a latent refinement from native 768p) for 5–15
seconds, with native stereo audio. It reports image, video, and audio
supported_input_references. Gemini Omni Flash 1.1 (gemini-omni-flash-1.1)
accepts 3–10 seconds at 360p/720p/1080p/4K in 16:9 or 9:16, with native synced
audio. It accepts image and video input_references (no audio). It makes its
video edit mode available through the Video Router video_url field.
Via the Models API
You can also use the public model catalog to find video generation models:
On the Models Page
Go to the Models page. Find the models that show video as an
output modality.
How It Works
Different from text or image generation, video generation is asynchronous, because a video takes much more time to generate. The workflow has these steps:
- Submit a generation request to
POST /v1/videos - Receive a job ID and a polling URL immediately
- Poll the polling URL (
GET /v1/videos/{jobId}) until the status iscompleted - Download the video from the content URL (
GET /v1/videos/{jobId}/content)
API Usage
Submitting a Video Generation Request
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | The model for video generation (for example, seedance-2.5), or sume/auto to let Sume select the model |
prompt | string | Yes | Text description of the video to generate |
duration | integer | No | Duration of the generated video in seconds |
resolution | string | No | Resolution of the output video (for example, 720p, 1080p) |
aspect_ratio | string | No | Aspect ratio of the output video (for example, 16:9, 9:16, 3:2) |
size | string | No | Accurate pixel dimensions in WIDTHxHEIGHT format (for example, 1280x720). An alternative to resolution + aspect_ratio |
frame_images | array | No | Images for first/last frames (image-to-video) |
input_references | array | No | Reference images for the style (reference-to-video) |
generate_audio | boolean | No | Tells the model to generate audio with the video, or not. The default is the audio capability of the model |
seed | integer | No | Seed for deterministic generation (the result is not deterministic on all providers) |
callback_url | string | No | URL that receives a webhook notification when the job completes. It must be HTTPS |
provider | object | No | Provider-specific passthrough configuration |
Supported Resolutions
480p720p768p1080p1K2K4K
Each model shows the subset that it accepts in supported_resolutions.
Supported Aspect Ratios
16:9— Widescreen landscape9:16— Vertical/portrait1:1— Square4:3— Standard landscape3:4— Standard portrait3:2— Photography landscape2:3— Photography portrait21:9— Ultra-wide9:21— Ultra-tall
Each model shows the subset that it accepts in supported_aspect_ratios.
Using Images
You can send images in two ways. Each way starts a different generation mode:
frame_images— Gives first or last frame images for image-to-video generation. Each entry must include aframe_typeoffirst_frameorlast_frame.input_references— Gives style or content reference images for reference-to-video generation. The model uses these images as visual guidance, not as accurate frames.
If you send the two fields, frame_images controls the mode, and Sume
processes the request as image-to-video.
Image-to-Video (frame_images)
Reference-to-Video (input_references)
A model accepts a reference type only if its supported_input_references
includes that type. The Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3
Max accept audio and video references. Gemini Omni Flash 1.1,
higgsfield-genjutsu, and h3-max-recast accept video references but not
audio.
Provider-Specific Options
You can send provider-specific options in the provider parameter. Each
provider slug is the key for its options. Sume forwards only the options for the
matched provider:
To find the passthrough parameters that each model supports, read the
allowed_passthrough_parameters field in the
Video Models API. In v1, that list is empty for
each model. Thus, Sume rejects provider.options entries with an error, and
does not ignore them. Refer to Sume differences.
Response Format
Submit Response (202 Accepted)
When you submit a video generation request, you immediately receive a response with the job details:
Poll Response
When you poll the job status, the response includes more fields as the job continues:
Job Statuses
| Status | Description |
|---|---|
pending | You submitted the job, and it is in the queue |
in_progress | The video generation is in progress |
completed | You can download the video |
failed | The generation failed (examine the error field) |
cancelled | The job is canceled, and it did not complete |
Downloading the Video
When the job status is completed, the unsigned_urls array contains URLs for
the download of the generated video content. You can also use the content
endpoint directly:
The default of the index query parameter is 0. If the model generates more
than one video output, use this parameter.
Webhooks
If you do not want to poll for the job status, you can receive a webhook
notification when a video generation job completes. Send callback_url in the
request body. When the job gets to a terminal state, Sume POSTs to that URL.
Sume signs the raw JSON body and sends x-sume-webhook-timestamp and
x-sume-webhook-signature headers. The payload is the standard job webhook
envelope of Sume, not the OpenRouter video.generation.* envelope. For the
accurate shape and the verification steps, refer to
Sume differences and the
webhooks guide.
Sume differences
All the information above agrees with the OpenRouter Video Generation API. This table gives the only differences.
| Area | OpenRouter | Sume |
|---|---|---|
| Base path | https://openrouter.ai/api/v1/videos | https://api.sume.com/v1/videos (no /api segment) |
| Auth | Authorization: Bearer $OPENROUTER_API_KEY | Authorization: Bearer $SUME_API_KEY |
| Model ids | org/slug (for example, google/veo-3.1) | bare catalog IDs (for example, seedance-2). The published contract of Sume never has a provider-org prefix |
| Auto routing | no generate-time auto | model: "sume/auto" lets Sume select the family. Responses show sume/auto. Sume never discloses the family that ran |
size | accepted when the model shows supported_sizes | each v1 model reports supported_sizes: null. Thus, size returns 400 unsupported_parameter. Use resolution + aspect_ratio |
provider.options | OpenRouter forwards it to the matched upstream provider | v1 runs one backend for each model. Thus, a non-empty provider.options returns 400 unsupported_parameter |
seed | many models accept it | no v1 model accepts seed. Each model reports seed: false and rejects the field |
| Webhook envelope | video.generation.* events, X-OpenRouter-Signature | Sume's standard job webhook envelope with x-sume-webhook-signature |
| Idempotency | none on this route | send Idempotency-Key to make retries safe. A replay returns the original job |
| Job lifecycle | only the polling URL | you can also see the same job at GET /v1/jobs/{id}/status and GET /v1/jobs/{id}/result |
| Zero Data Retention | video generation is ZDR-ineligible | Sume has no ZDR toggle. Refer to the privacy docs |
| Billing | credits | workspace USD balance. At submit, Sume reserves provider list × 1.25 (each model, minimax-h3-max included). usage.cost is the Sume billable amount |
sume/auto
sume/auto is an addition that only Sume has. If you do not want to pin a
family, send it as model:
Resolution is a pure function of the normalized request and the catalog
version. Thus, an idempotent replay gets the same price and the same route. The
poll response reports "model": "sume/auto". Sume does not disclose which family
served the request. Do not use observable traits of the output to find the
family.
/v1/video-router/*
POST /v1/video-router/generate and GET /v1/video-router/models still work
as before. They create the same jobs, with the same model IDs. The difference is
the wire. Video Router returns the { "data": ... } job envelope of Sume and
accepts the flat image_url / reference_image_urls fields of Sume.
/v1/videos returns the response shape in the sections above.
The two APIs use the same model vocabulary. Thus, a migration is a path-and-body
change, and you do not have to map IDs again. We recommend that new
integrations use /v1/videos.
Best Practices
- Detailed Prompts: For better video quality, write specific prompts with much detail. Include details of motion, camera angles, lighting, and scene composition.
- Appropriate Resolution: A higher resolution takes more time to generate and has a higher price. Select the resolution that is applicable to your use case.
- Polling Interval: Use a moderate polling interval (for example, 30 seconds) to prevent too many API calls. Video generation usually takes from 30 seconds to several minutes, as a function of the model and the parameters.
- Error Handling: Always examine the job status for the
failedstate. Make sure that your code processes theerrorfield correctly. - Reference Images: When you use reference images, make sure that they have high quality and are applicable to the video that you want.
Troubleshooting
Job stays in pending for a long time?
- Video generation can take several minutes, as a function of the model, the resolution, and the server load
- Continue to poll at regular intervals
Generation failed?
- Examine the
errorfield in the poll response for details - Make sure that the model guidelines permit your prompt
- Make sure that all reference images are available over public HTTPS and are in supported formats
Model not found?
- Use the Video Models API to find available video generation models
- Make sure that the model ID is correct (for example,
seedance-2). Sume uses bare catalog IDs, notorg/slug
Workspace fal keys
In the development preview, workspace creators and admins can connect a fal API key in Dashboard → Integrations → Model keys. fal bills video models directly. Sume charges only 5.5% of fal list price, under the workspace fee terms. Image jobs do not change.
For BYOK video jobs, the poll response returns usage.is_byok: true, and usage.cost is only the Sume fee. /v1/usage includes byok: { provider: "fal", basis_usd_micros: ... }. Usage shows the label Billed by provider on the informational estimate.
There is no fallback to Sume credentials. If fal rejects the key, or the key balance is empty, the job fails with provider_byok_rejected, and Sume refunds the fee. If you disconnect the key, the videos that still run on that key fail. But fal can still complete them and bill them. If Sume cannot resolve the key storage, it returns 503 byok_unavailable before the submission. On production, this feature stays disabled by default.
Vercel AI Gateway keys for Agents
In Keys → Bring your own key, immediately below fal, you can connect a Vercel AI Gateway API key. This key is for the agent model calls of your workspace. This development preview uses the same owner/admin controls, encrypted storage, key tests, replacement, and disconnect actions as fal. A Gateway key test uses its credit endpoint and does not generate paid content.
Auto and pinned agent models use your Gateway account. Vercel bills the model. Sume charges only your workspace Agent Fee, which is 5.5% by default.
Usage shows BYOK · Gateway, the Sume fee, and the approximate customer-paid provider basis as separate items. Gateway generation receipts give the actual Gateway debit. An upstream provider-key list cost or a token fallback is an estimate. If a turn fails, Sume refunds the Sume fee. But Vercel can still bill the requests that completed.
In a BYOK turn, Sume never uses its own Gateway key as a fallback. If you remove or replace the key, an in-flight turn fails when its next credential check runs. Requests that Sume already accepted can still complete. After a disconnect, new turns use the usual Sume billing.
This key is for the main agent model. Video/media and auxiliary services keep their current providers and billing. The key does not add upstream provider keys to your Vercel team.