Image API

Sume has a dedicated Image API that generates images from text prompts and optional reference images. The service covers model discovery, per-endpoint capabilities, and generation features. To find the available models and their prices, use GET /v1/images/models.

The catalog publishes the parameters that each model accepts as capability descriptors. Thus, you can find what a model supports before you call it. Sume specifics lists the behavior that is specific to Sume.

To select a family, use a catalog model ID. To let Image Router choose, use model: "sume/auto". Sume will retire Image 1.0 soon. Its public URLs stay compatibility aliases for this Auto pipe.

Model discovery

ChatGPT Image 2.5 is available as openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst. Both support text-to-image, up to 16 image references, an optional mask_url, and background: auto|transparent|opaque. Use quality: auto|low|medium|high|xhigh|max. If you omit quality, the default is high.

image_size accepts named presets, auto, or custom pixels. For custom pixels, both edges must be multiples of 16, and the maximum edge is 3840. The aspect ratio must be at most 3:1, and the image must have 655,360–8,294,400 pixels. You can still select ChatGPT Image 2.

Flare and Sunburst use the same Fal token rates: $30 per million output image tokens, $8 per million input image tokens, and $5 per million input text tokens. Output estimates use OpenAI's ChatGPT Image 2.5 size/quality calculator. At 1024×1024, xhigh output is $0.09366 and max output is $0.21072, before input tokens and Sume pricing. Input token counts are estimates. Fal rounds the total up to $0.0001.

auto quality reserves max. auto size and named presets without a verified GPT-specific pixel mapping reserve the upper bound of output tokens. Auto model routing continues to use Flare. Fal does not advertise a separate Effortless endpoint. Its documented automatic quality option is auto.

Ideogram 4.5 is available as ideogram/ideogram-v4.5. Without input_references, it generates from text. With references, it edits the first image and uses up to 4 more as references (5 total). Use quality: low|medium|high (if you omit it, the default is medium) and resolution: 1K|2K. The Fal list price is $0.03, $0.06, or $0.22 per image by quality, for all sizes. An edit without aspect_ratio keeps the shape of the source image.

Via the Image Models API

To list the available models and their capabilities, use the image models endpoint:

Key response fields include:

  • id: Model slug for generation requests
  • architecture: Supported input/output modalities
  • supported_parameters: Union of capabilities across endpoints
  • supports_streaming: Shows if native SSE streaming is available
  • endpoints: URL for per-endpoint records

Per-endpoint records

To get the definitive capabilities and prices for a model, use this call:

The important fields are:

  • provider_slug: Use this value for provider-specific parameters
  • provider_tag: Use this value to pin requests to specific providers
  • supported_parameters: Definitive parameter set for this endpoint
  • allowed_passthrough_parameters: Provider-specific keys
  • pricing: Billable lines with cost information

In v1, Sume serves every catalog model through a single sume endpoint. Thus, the model-level and endpoint-level supported_parameters are identical.

Capability descriptors

Parameters use typed descriptors:

  • enum: Discrete allowlist of string values
  • range: Any integer in min/max bounds
  • boolean: Supported (present) or unsupported (absent)

If a request sets a parameter that the selected model does not list, Sume rejects the request with 400 unsupported_parameter. Sume does not silently drop the parameter.

API usage

Send a POST request to /v1/images with a model and a prompt:

Python (requests):

TypeScript (fetch):

cURL:

Response format

The response gives the images as Sume-hosted URLs with usage data:

For non-PNG formats, the response is:

model echoes the id that you requested. For example, sume/auto stays sume/auto. cost is the USD amount that Sume bills to your wallet. In v1, token counts are always 0. Sume meters image models per image, and per-token usage data is not available yet.

Long-running requests

POST /v1/images blocks for up to 30 seconds and returns the response above with 200. Most catalog models complete in that budget.

If the generation does not complete before the budget expires, Sume returns 202 with the standard job envelope instead. Sume also returns this envelope if you send mode: "async", or mode: "webhook" with a webhook_url:

Poll GET /v1/jobs/{id}/status. Then fetch GET /v1/jobs/{id}/result to get the generated images. These are the standard Sume job endpoints. They return the standard job result shape, not the image body above. Refer to Jobs and results.

Examine the status code, not the body shape. 200 is the image response, and 202 is the job envelope. Slow configurations (4K, high quality, large n) are the most likely to degrade to 202.

Image configuration options

Resolution and aspect ratio

  • resolution: Normalized tier (512, 1K, 2K, 4K)
  • aspect_ratio: Normalized ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 1:2, 2:1, 1:4, 4:1, 1:8, 8:1, 9:21, 21:9). Use "auto" to let the provider choose
  • size: Shorthand for a resolution tier. Do not put custom pixels on size. Use image_size / aspect_ratio. 4:5 is Instagram portrait (1080×1350), not 4:3. Banana Pro sends aspect_ratio: "4:5" (native ~928×1152 at 1K). Exact 1080×1350 comes from a documented post-step through job target_pixels.

A model accepts only the values that its catalog descriptors list. Thus, read supported_parameters before you pin a tier or ratio.

On edit and image-to-image calls, use aspect_ratio: "auto" to match the reference. If you omit the field, the result is not the same as auto.

Quality and output format

  • quality: auto, low, medium, high, xhigh, or max (catalog-gated)
  • output_format: png, jpeg, webp, or svg
  • background: auto, transparent, or opaque (ChatGPT Image 2.5 supports this field)
  • output_compression: 0–100 for webp/jpeg (Sume does not serve this field in v1)

output_compression and seed are part of the schema, but no model advertises them yet. Thus, a request that sends one of them returns 400 unsupported_parameter. For transparent stills today, use Image 1.0 with transparency: true.

Multiple images

Use n to request up to 10 images per call. Per-model ceilings are lower. Read the n range descriptor from the catalog.

Image-to-image (reference images)

Reference URLs must be public HTTPS. Sume rejects localhost, private-network, and non-HTTPS URLs before submission. If the input_references descriptor of a model is {"min": 0, "max": 0}, the model is text-to-image only and rejects references.

Provider routing

The routing fields are:

  • only: Allow only the listed provider slugs
  • order: Try providers in the listed order
  • ignore: Do not use the listed provider slugs
  • sort: Sort by price, throughput, or latency
  • allow_fallbacks: If false, stop after the primary provider

In v1, Sume publishes a single sume endpoint for each model. Thus, only and order accept only "sume". Sume accepts ignore, sort, and allow_fallbacks, but they have no effect. Any other slug returns 400 provider_not_available.

Provider-specific options

In v1, allowed_passthrough_parameters is empty for every endpoint. Thus, you must omit provider.options or send it empty.

Streaming image generation

In v1, Sume does not serve native SSE streaming. Every catalog row reports supports_streaming: false, and stream: true returns 400 streaming_not_supported. The field is in the schema so that clients can use streams without a code change when Sume ships the feature.

Until then, submit with mode: "async". Then read GET /v1/jobs/:id/events for progress. As an alternative, use a webhook for the terminal event. mode: "subscribe" is not a progress stream. It is an alias of sync and gives you one bounded 30-second wait. Refer to what "subscribe" means.

Billing and cancellation

Sume bills image generation on an all-or-nothing basis. Either a generation completes and Sume bills it in full, or the generation fails and Sume does not bill it.

  • Completed generations get the full charge from the endpoint pricing
  • Failed or cancelled generations get no charge. Failed requests return 502 Bad Gateway
  • Client disconnects: Sume bills a request that ends early as a failed generation (no charge)

Endpoint pricing lines are the amount that Sume charges to your wallet. These lines already include the Sume margin. Thus, you pay cost_usd × n.

Request parameters

ParameterTypeRequiredDescription
modelstringYesModel slug (for example, bytedance-seed/seedream-4.5), or sume/auto
promptstringYesText that describes the image
nintegerNoNumber of images to generate (1–10)
resolutionstringNoResolution tier (512, 1K, 2K, 4K)
aspect_ratiostringNoAspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 1:4, 4:1, etc.)
sizestringNoShorthand for a resolution tier. Sume does not serve explicit pixels in v1
qualitystringNoauto, low, medium, high, xhigh, or max (catalog-gated)
output_formatstringNopng, jpeg, webp, or svg
mask_urlstringNoOptional public HTTPS mask URL for ChatGPT Image 2.5 edits.
backgroundstringNoauto, transparent, or opaque (ChatGPT Image 2.5)
output_compressionintegerNoCompression level (0–100) for webp/jpeg (not served in v1)
seedintegerNoSeed for deterministic generation (not served in v1)
streambooleanNoStream partial images through SSE (not served in v1)
input_referencesarrayNoReference images for image-to-image
provider.onlystring[]NoAllow only these provider slugs
provider.orderstring[]NoTry provider slugs in this order
provider.ignorestring[]NoExclude these provider slugs
provider.sortstring or objectNoSort by price, throughput, or latency
provider.allow_fallbacksbooleanNoAllow fallback provider on failure
provider.optionsobjectNoProvider-specific parameters by slug
metadataobjectNoCaller metadata that Sume stores on the job. Sume does not send it to the provider
modestringNosync (default on this route), async, subscribe, webhook
webhook_urlstringNoPublic HTTPS callback for terminal delivery in webhook mode
wait_timeout_secondsintegerNo0–30, default 30 on this route. Maximum time that sync / subscribe waits before it returns

Sume specifics

AreaBehavior
sume/autoSume-only model value. Sume selects the family for you and never discloses which one ran. GET /v1/images/models does not list it, and job.model stays sume/auto.
Result payloaddata[].url (Sume-hosted, signed), not inline base64. Sume already mirrors generated media, and URLs keep responses small.
AsyncSume limits the wait of a call to 30s. Generations that exceed this limit, and mode: "async" / "webhook", return the Sume job envelope with 202. You then read the images from the standard job result endpoint.
ProviderOne sume endpoint for each model in v1. Sume does not disclose the upstream provider identity. Sume accepts the multi-provider routing fields, but they have no effect.
streamThe schema accepts this field. Until native SSE ships, Sume rejects it at runtime with 400 streaming_not_supported.
Catalog-gated parametersoutput_compression, seed, and explicit pixel size are in the schema, but no model advertises them in v1. Thus, they return 400 unsupported_parameter.
usagecost is the billed USD amount. Token counts are 0 in v1.
Legacy idsSume accepts the bare Image Router ids (gpt-image-2, nano-banana-2.1, …) as aliases for their org/slug equivalents.
Retired modelsNano Banana 2 is retired. google/nano-banana-2 and nano-banana-2 still work and run as Nano Banana 2.1 (google/nano-banana-2.1), and the job stores the 2.1 id. The catalog does not list the retired id.

The legacy POST /v1/image-router/generate and GET /v1/image-router/models routes still work without change. But they are deprecated, and this surface replaces them. These routes will not get new parameters.