Image API
Sume has a dedicated Image API that generates images from text prompts and
optional reference images. The service covers model discovery, per-endpoint
capabilities, and generation features. To find the available models and
their prices, use GET /v1/images/models.
The catalog publishes the parameters that each model accepts as capability descriptors. Thus, you can find what a model supports before you call it. Sume specifics lists the behavior that is specific to Sume.
To select a family, use a catalog model ID. To let Image Router choose, use
model: "sume/auto". Sume will retire Image 1.0 soon. Its
public URLs stay compatibility aliases for this Auto pipe.
Model discovery
ChatGPT Image 2.5 is available as openai/gpt-image-2.5 (Flare) and
openai/gpt-image-2.5-sunburst. Both support text-to-image, up to 16 image
references, an optional mask_url, and background: auto|transparent|opaque.
Use quality: auto|low|medium|high|xhigh|max. If you omit quality, the
default is high.
image_size accepts named presets, auto, or custom pixels. For custom
pixels, both edges must be multiples of 16, and the maximum edge is 3840. The
aspect ratio must be at most 3:1, and the image must have 655,360–8,294,400
pixels. You can still select ChatGPT Image 2.
Flare and Sunburst use the same Fal token rates:
$30 per million output image tokens, $8 per million input image tokens, and $5
per million input text tokens. Output estimates use OpenAI's ChatGPT Image 2.5
size/quality calculator. At 1024×1024, xhigh output is $0.09366 and max
output is $0.21072, before input tokens and Sume pricing. Input token counts
are estimates. Fal rounds the total up to $0.0001.
auto quality reserves max. auto size and named presets without a
verified GPT-specific pixel mapping reserve the upper bound of output tokens.
Auto model routing continues to use Flare. Fal does not advertise a separate
Effortless endpoint. Its documented automatic quality option is auto.
Ideogram 4.5 is available as ideogram/ideogram-v4.5. Without
input_references, it generates from text. With references, it edits the
first image and uses up to 4 more as references (5 total). Use
quality: low|medium|high (if you omit it, the default is medium) and
resolution: 1K|2K. The Fal list price
is $0.03, $0.06, or $0.22 per image by quality, for all sizes. An edit without
aspect_ratio keeps the shape of the source image.
Via the Image Models API
To list the available models and their capabilities, use the image models endpoint:
Key response fields include:
- id: Model slug for generation requests
- architecture: Supported input/output modalities
- supported_parameters: Union of capabilities across endpoints
- supports_streaming: Shows if native SSE streaming is available
- endpoints: URL for per-endpoint records
Per-endpoint records
To get the definitive capabilities and prices for a model, use this call:
The important fields are:
- provider_slug: Use this value for provider-specific parameters
- provider_tag: Use this value to pin requests to specific providers
- supported_parameters: Definitive parameter set for this endpoint
- allowed_passthrough_parameters: Provider-specific keys
- pricing: Billable lines with cost information
In v1, Sume serves every catalog model through a single sume endpoint.
Thus, the model-level and endpoint-level supported_parameters are identical.
Capability descriptors
Parameters use typed descriptors:
- enum: Discrete allowlist of string values
- range: Any integer in min/max bounds
- boolean: Supported (present) or unsupported (absent)
If a request sets a parameter that the selected model does not list, Sume
rejects the request with 400 unsupported_parameter. Sume does not silently
drop the parameter.
API usage
Send a POST request to /v1/images with a model and a prompt:
Python (requests):
TypeScript (fetch):
cURL:
Response format
The response gives the images as Sume-hosted URLs with usage data:
For non-PNG formats, the response is:
model echoes the id that you requested. For example, sume/auto stays
sume/auto. cost is the USD amount that Sume bills to your wallet. In v1,
token counts are always 0. Sume meters image models per image, and per-token
usage data is not available yet.
Long-running requests
POST /v1/images blocks for up to 30 seconds and returns the response above
with 200. Most catalog models complete in that budget.
If the generation does not complete before the budget expires, Sume returns
202 with the standard job envelope instead. Sume also returns this envelope
if you send mode: "async", or mode: "webhook" with a webhook_url:
Poll GET /v1/jobs/{id}/status. Then fetch GET /v1/jobs/{id}/result to get
the generated images. These are the standard Sume job endpoints. They return
the standard job result shape, not the image body above. Refer to
Jobs and results.
Examine the status code, not the body shape. 200 is the image response, and
202 is the job envelope. Slow configurations (4K, high quality, large n)
are the most likely to degrade to 202.
Image configuration options
Resolution and aspect ratio
- resolution: Normalized tier (512, 1K, 2K, 4K)
- aspect_ratio: Normalized ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 1:2, 2:1, 1:4, 4:1, 1:8, 8:1, 9:21, 21:9). Use "auto" to let the provider choose
- size: Shorthand for a resolution tier. Do not put custom pixels on
size. Useimage_size/aspect_ratio.4:5is Instagram portrait (1080×1350), not 4:3. Banana Pro sendsaspect_ratio: "4:5"(native ~928×1152 at 1K). Exact 1080×1350 comes from a documented post-step through jobtarget_pixels.
A model accepts only the values that its catalog descriptors list. Thus, read
supported_parameters before you pin a tier or ratio.
On edit and image-to-image calls, use aspect_ratio: "auto" to match the
reference. If you omit the field, the result is not the same as auto.
Quality and output format
- quality: auto, low, medium, high, xhigh, or max (catalog-gated)
- output_format: png, jpeg, webp, or svg
- background: auto, transparent, or opaque (ChatGPT Image 2.5 supports this field)
- output_compression: 0–100 for webp/jpeg (Sume does not serve this field in v1)
output_compression and seed are part of the schema, but no model
advertises them yet. Thus, a request that sends one of them returns
400 unsupported_parameter.
For transparent stills today, use Image 1.0 with
transparency: true.
Multiple images
Use n to request up to 10 images per call. Per-model ceilings are lower.
Read the n range descriptor from the catalog.
Image-to-image (reference images)
Reference URLs must be public HTTPS. Sume rejects localhost, private-network,
and non-HTTPS URLs before submission. If the input_references descriptor of a
model is {"min": 0, "max": 0}, the model is text-to-image only and rejects
references.
Provider routing
The routing fields are:
- only: Allow only the listed provider slugs
- order: Try providers in the listed order
- ignore: Do not use the listed provider slugs
- sort: Sort by price, throughput, or latency
- allow_fallbacks: If false, stop after the primary provider
In v1, Sume publishes a single sume endpoint for each model. Thus, only and
order accept only "sume". Sume accepts ignore, sort, and
allow_fallbacks, but they have no effect. Any other slug returns
400 provider_not_available.
Provider-specific options
In v1, allowed_passthrough_parameters is empty for every endpoint. Thus, you
must omit provider.options or send it empty.
Streaming image generation
In v1, Sume does not serve native SSE streaming. Every catalog row reports
supports_streaming: false, and stream: true returns
400 streaming_not_supported. The field is in the schema so that clients can
use streams without a code change when Sume ships the feature.
Until then, submit with mode: "async". Then read GET /v1/jobs/:id/events
for progress. As an alternative, use a webhook for the
terminal event. mode: "subscribe" is not a progress stream. It is an
alias of sync and gives you one bounded 30-second wait. Refer to
what "subscribe" means.
Billing and cancellation
Sume bills image generation on an all-or-nothing basis. Either a generation completes and Sume bills it in full, or the generation fails and Sume does not bill it.
- Completed generations get the full charge from the endpoint pricing
- Failed or cancelled generations get no charge. Failed requests return 502 Bad Gateway
- Client disconnects: Sume bills a request that ends early as a failed generation (no charge)
Endpoint pricing lines are the amount that Sume charges to your wallet. These
lines already include the Sume margin. Thus, you pay cost_usd × n.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model slug (for example, bytedance-seed/seedream-4.5), or sume/auto |
prompt | string | Yes | Text that describes the image |
n | integer | No | Number of images to generate (1–10) |
resolution | string | No | Resolution tier (512, 1K, 2K, 4K) |
aspect_ratio | string | No | Aspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 1:4, 4:1, etc.) |
size | string | No | Shorthand for a resolution tier. Sume does not serve explicit pixels in v1 |
quality | string | No | auto, low, medium, high, xhigh, or max (catalog-gated) |
output_format | string | No | png, jpeg, webp, or svg |
mask_url | string | No | Optional public HTTPS mask URL for ChatGPT Image 2.5 edits. |
background | string | No | auto, transparent, or opaque (ChatGPT Image 2.5) |
output_compression | integer | No | Compression level (0–100) for webp/jpeg (not served in v1) |
seed | integer | No | Seed for deterministic generation (not served in v1) |
stream | boolean | No | Stream partial images through SSE (not served in v1) |
input_references | array | No | Reference images for image-to-image |
provider.only | string[] | No | Allow only these provider slugs |
provider.order | string[] | No | Try provider slugs in this order |
provider.ignore | string[] | No | Exclude these provider slugs |
provider.sort | string or object | No | Sort by price, throughput, or latency |
provider.allow_fallbacks | boolean | No | Allow fallback provider on failure |
provider.options | object | No | Provider-specific parameters by slug |
metadata | object | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider |
mode | string | No | sync (default on this route), async, subscribe, webhook |
webhook_url | string | No | Public HTTPS callback for terminal delivery in webhook mode |
wait_timeout_seconds | integer | No | 0–30, default 30 on this route. Maximum time that sync / subscribe waits before it returns |
Sume specifics
| Area | Behavior |
|---|---|
sume/auto | Sume-only model value. Sume selects the family for you and never discloses which one ran. GET /v1/images/models does not list it, and job.model stays sume/auto. |
| Result payload | data[].url (Sume-hosted, signed), not inline base64. Sume already mirrors generated media, and URLs keep responses small. |
| Async | Sume limits the wait of a call to 30s. Generations that exceed this limit, and mode: "async" / "webhook", return the Sume job envelope with 202. You then read the images from the standard job result endpoint. |
| Provider | One sume endpoint for each model in v1. Sume does not disclose the upstream provider identity. Sume accepts the multi-provider routing fields, but they have no effect. |
stream | The schema accepts this field. Until native SSE ships, Sume rejects it at runtime with 400 streaming_not_supported. |
| Catalog-gated parameters | output_compression, seed, and explicit pixel size are in the schema, but no model advertises them in v1. Thus, they return 400 unsupported_parameter. |
usage | cost is the billed USD amount. Token counts are 0 in v1. |
| Legacy ids | Sume accepts the bare Image Router ids (gpt-image-2, nano-banana-2.1, …) as aliases for their org/slug equivalents. |
| Retired models | Nano Banana 2 is retired. google/nano-banana-2 and nano-banana-2 still work and run as Nano Banana 2.1 (google/nano-banana-2.1), and the job stores the 2.1 id. The catalog does not list the retired id. |
The legacy POST /v1/image-router/generate and GET /v1/image-router/models
routes still work without change. But they are deprecated, and this surface
replaces them. These routes will not get new parameters.