Video Generation

Sume supports video generation from text prompts (and optional reference images) through a dedicated asynchronous API. To find the supported models, their capabilities, and their prices, use GET /v1/videos/models.

This surface agrees field-for-field with the OpenRouter Video Generation API. Thus, a client that you wrote from their docs works here after you change the base URL and the API key. The few differences of Sume are in Sume differences.

Model Discovery

You can find video generation models in these ways:

Via the Video Models API

To get a list of all available video generation models and their supported parameters, use the dedicated video models endpoint:

The response returns a data array. Each model in the array includes these fields:

FieldDescription
idModel slug to use in generation requests
canonical_slugPermanent model identifier
supported_resolutionsList of supported output resolutions (for example, 720p, 1080p)
supported_aspect_ratiosList of supported aspect ratios (for example, 16:9, 9:16)
supported_sizesList of supported pixel dimensions (for example, 1280x720), or null
supported_durationsSupported video lengths in whole seconds
supported_frame_imagesThe frame_type values that the model accepts
supported_input_referencesThe input_references types that the model accepts
generate_audioShows if the model can generate an audio track
seedShows if the model accepts a seed
pricing_skusPrice information for each SKU
allowed_passthrough_parametersProvider-specific parameters that you can send through the provider option

Before you submit a generation request, use this endpoint to find the resolutions, aspect ratios, and durations that each model supports. Limits are different for each model. seedance-2.5 accepts 4–30 seconds at 480p/720p/1080p. wan-3.0 accepts 2–30 seconds. higgsfield-genjutsu (Motion Transfer: one source video plus 1–8 reference images, 480p/720p, in the catalog only when its provider is configured) accepts 4–30 seconds.

H3 Max Recast (h3-max-recast: one source video plus 1–4 person photos, 768p/1080p, optional prompt) accepts 5–30 seconds. minimax-h3 accepts 5–15 seconds at native 480p/768p (768p is first-class, not 720p). For this model, Sume bills 2K/4K upscales if your request includes them. Each other catalog model has a maximum of 15 seconds. seedance-2 also has 1080p.

MiniMax H3 Max (minimax-h3-max) is the faster 768p variant. It does text-to-video, first/last-frame image-to-video, and reference-to-video at 480p/768p/1080p (1080p is a latent refinement from native 768p) for 5–15 seconds, with native stereo audio. It reports image, video, and audio supported_input_references. Gemini Omni Flash 1.1 (gemini-omni-flash-1.1) accepts 3–10 seconds at 360p/720p/1080p/4K in 16:9 or 9:16, with native synced audio. It accepts image and video input_references (no audio). It makes its video edit mode available through the Video Router video_url field.

Via the Models API

You can also use the public model catalog to find video generation models:

On the Models Page

Go to the Models page. Find the models that show video as an output modality.

How It Works

Different from text or image generation, video generation is asynchronous, because a video takes much more time to generate. The workflow has these steps:

  1. Submit a generation request to POST /v1/videos
  2. Receive a job ID and a polling URL immediately
  3. Poll the polling URL (GET /v1/videos/{jobId}) until the status is completed
  4. Download the video from the content URL (GET /v1/videos/{jobId}/content)

API Usage

Submitting a Video Generation Request

Request Parameters

ParameterTypeRequiredDescription
modelstringYesThe model for video generation (for example, seedance-2.5), or sume/auto to let Sume select the model
promptstringYesText description of the video to generate
durationintegerNoDuration of the generated video in seconds
resolutionstringNoResolution of the output video (for example, 720p, 1080p)
aspect_ratiostringNoAspect ratio of the output video (for example, 16:9, 9:16, 3:2)
sizestringNoAccurate pixel dimensions in WIDTHxHEIGHT format (for example, 1280x720). An alternative to resolution + aspect_ratio
frame_imagesarrayNoImages for first/last frames (image-to-video)
input_referencesarrayNoReference images for the style (reference-to-video)
generate_audiobooleanNoTells the model to generate audio with the video, or not. The default is the audio capability of the model
seedintegerNoSeed for deterministic generation (the result is not deterministic on all providers)
callback_urlstringNoURL that receives a webhook notification when the job completes. It must be HTTPS
providerobjectNoProvider-specific passthrough configuration

Supported Resolutions

  • 480p
  • 720p
  • 768p
  • 1080p
  • 1K
  • 2K
  • 4K

Each model shows the subset that it accepts in supported_resolutions.

Supported Aspect Ratios

  • 16:9 — Widescreen landscape
  • 9:16 — Vertical/portrait
  • 1:1 — Square
  • 4:3 — Standard landscape
  • 3:4 — Standard portrait
  • 3:2 — Photography landscape
  • 2:3 — Photography portrait
  • 21:9 — Ultra-wide
  • 9:21 — Ultra-tall

Each model shows the subset that it accepts in supported_aspect_ratios.

Using Images

You can send images in two ways. Each way starts a different generation mode:

  • frame_images — Gives first or last frame images for image-to-video generation. Each entry must include a frame_type of first_frame or last_frame.
  • input_references — Gives style or content reference images for reference-to-video generation. The model uses these images as visual guidance, not as accurate frames.

If you send the two fields, frame_images controls the mode, and Sume processes the request as image-to-video.

Image-to-Video (frame_images)

Reference-to-Video (input_references)

A model accepts a reference type only if its supported_input_references includes that type. The Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max accept audio and video references. Gemini Omni Flash 1.1, higgsfield-genjutsu, and h3-max-recast accept video references but not audio.

Provider-Specific Options

You can send provider-specific options in the provider parameter. Each provider slug is the key for its options. Sume forwards only the options for the matched provider:

To find the passthrough parameters that each model supports, read the allowed_passthrough_parameters field in the Video Models API. In v1, that list is empty for each model. Thus, Sume rejects provider.options entries with an error, and does not ignore them. Refer to Sume differences.

Response Format

Submit Response (202 Accepted)

When you submit a video generation request, you immediately receive a response with the job details:

Poll Response

When you poll the job status, the response includes more fields as the job continues:

Job Statuses

StatusDescription
pendingYou submitted the job, and it is in the queue
in_progressThe video generation is in progress
completedYou can download the video
failedThe generation failed (examine the error field)
cancelledThe job is canceled, and it did not complete

Downloading the Video

When the job status is completed, the unsigned_urls array contains URLs for the download of the generated video content. You can also use the content endpoint directly:

The default of the index query parameter is 0. If the model generates more than one video output, use this parameter.

Webhooks

If you do not want to poll for the job status, you can receive a webhook notification when a video generation job completes. Send callback_url in the request body. When the job gets to a terminal state, Sume POSTs to that URL.

Sume signs the raw JSON body and sends x-sume-webhook-timestamp and x-sume-webhook-signature headers. The payload is the standard job webhook envelope of Sume, not the OpenRouter video.generation.* envelope. For the accurate shape and the verification steps, refer to Sume differences and the webhooks guide.

Sume differences

All the information above agrees with the OpenRouter Video Generation API. This table gives the only differences.

AreaOpenRouterSume
Base pathhttps://openrouter.ai/api/v1/videoshttps://api.sume.com/v1/videos (no /api segment)
AuthAuthorization: Bearer $OPENROUTER_API_KEYAuthorization: Bearer $SUME_API_KEY
Model idsorg/slug (for example, google/veo-3.1)bare catalog IDs (for example, seedance-2). The published contract of Sume never has a provider-org prefix
Auto routingno generate-time automodel: "sume/auto" lets Sume select the family. Responses show sume/auto. Sume never discloses the family that ran
sizeaccepted when the model shows supported_sizeseach v1 model reports supported_sizes: null. Thus, size returns 400 unsupported_parameter. Use resolution + aspect_ratio
provider.optionsOpenRouter forwards it to the matched upstream providerv1 runs one backend for each model. Thus, a non-empty provider.options returns 400 unsupported_parameter
seedmany models accept itno v1 model accepts seed. Each model reports seed: false and rejects the field
Webhook envelopevideo.generation.* events, X-OpenRouter-SignatureSume's standard job webhook envelope with x-sume-webhook-signature
Idempotencynone on this routesend Idempotency-Key to make retries safe. A replay returns the original job
Job lifecycleonly the polling URLyou can also see the same job at GET /v1/jobs/{id}/status and GET /v1/jobs/{id}/result
Zero Data Retentionvideo generation is ZDR-ineligibleSume has no ZDR toggle. Refer to the privacy docs
Billingcreditsworkspace USD balance. At submit, Sume reserves provider list × 1.25 (each model, minimax-h3-max included). usage.cost is the Sume billable amount

sume/auto

sume/auto is an addition that only Sume has. If you do not want to pin a family, send it as model:

Resolution is a pure function of the normalized request and the catalog version. Thus, an idempotent replay gets the same price and the same route. The poll response reports "model": "sume/auto". Sume does not disclose which family served the request. Do not use observable traits of the output to find the family.

/v1/video-router/*

POST /v1/video-router/generate and GET /v1/video-router/models still work as before. They create the same jobs, with the same model IDs. The difference is the wire. Video Router returns the { "data": ... } job envelope of Sume and accepts the flat image_url / reference_image_urls fields of Sume. /v1/videos returns the response shape in the sections above.

The two APIs use the same model vocabulary. Thus, a migration is a path-and-body change, and you do not have to map IDs again. We recommend that new integrations use /v1/videos.

Best Practices

  • Detailed Prompts: For better video quality, write specific prompts with much detail. Include details of motion, camera angles, lighting, and scene composition.
  • Appropriate Resolution: A higher resolution takes more time to generate and has a higher price. Select the resolution that is applicable to your use case.
  • Polling Interval: Use a moderate polling interval (for example, 30 seconds) to prevent too many API calls. Video generation usually takes from 30 seconds to several minutes, as a function of the model and the parameters.
  • Error Handling: Always examine the job status for the failed state. Make sure that your code processes the error field correctly.
  • Reference Images: When you use reference images, make sure that they have high quality and are applicable to the video that you want.

Troubleshooting

Job stays in pending for a long time?

  • Video generation can take several minutes, as a function of the model, the resolution, and the server load
  • Continue to poll at regular intervals

Generation failed?

  • Examine the error field in the poll response for details
  • Make sure that the model guidelines permit your prompt
  • Make sure that all reference images are available over public HTTPS and are in supported formats

Model not found?

  • Use the Video Models API to find available video generation models
  • Make sure that the model ID is correct (for example, seedance-2). Sume uses bare catalog IDs, not org/slug

Workspace fal keys

In the development preview, workspace creators and admins can connect a fal API key in Dashboard → Integrations → Model keys. fal bills video models directly. Sume charges only 5.5% of fal list price, under the workspace fee terms. Image jobs do not change.

For BYOK video jobs, the poll response returns usage.is_byok: true, and usage.cost is only the Sume fee. /v1/usage includes byok: { provider: "fal", basis_usd_micros: ... }. Usage shows the label Billed by provider on the informational estimate.

There is no fallback to Sume credentials. If fal rejects the key, or the key balance is empty, the job fails with provider_byok_rejected, and Sume refunds the fee. If you disconnect the key, the videos that still run on that key fail. But fal can still complete them and bill them. If Sume cannot resolve the key storage, it returns 503 byok_unavailable before the submission. On production, this feature stays disabled by default.

Vercel AI Gateway keys for Agents

In Keys → Bring your own key, immediately below fal, you can connect a Vercel AI Gateway API key. This key is for the agent model calls of your workspace. This development preview uses the same owner/admin controls, encrypted storage, key tests, replacement, and disconnect actions as fal. A Gateway key test uses its credit endpoint and does not generate paid content.

Auto and pinned agent models use your Gateway account. Vercel bills the model. Sume charges only your workspace Agent Fee, which is 5.5% by default.

Usage shows BYOK · Gateway, the Sume fee, and the approximate customer-paid provider basis as separate items. Gateway generation receipts give the actual Gateway debit. An upstream provider-key list cost or a token fallback is an estimate. If a turn fails, Sume refunds the Sume fee. But Vercel can still bill the requests that completed.

In a BYOK turn, Sume never uses its own Gateway key as a fallback. If you remove or replace the key, an in-flight turn fails when its next credential check runs. Requests that Sume already accepted can still complete. After a disconnect, new turns use the usual Sume billing.

This key is for the main agent model. Video/media and auxiliary services keep their current providers and billing. The key does not add upstream provider keys to your Vercel team.