Music 1.0
The retirement of Music 1.0 is gradual. Its routes continue to work and keep
job.model = sume/music-1.0. But every request now resolves through the
Music Router (sume/music-auto). For new integrations,
we recommend POST /v1/music-router/generate.
Use Music 1.0 for prompt-driven music generation. The request is text-first,
and the model also supports optional image conditioning. Provider model ids
stay internal. Music 1.0 runs on Google Lyria 3.5. It makes full-length
structured songs of up to a few minutes, and the prompt controls them. For
example, write "a 2-minute track" or use section markers such as
[0:00-0:30] Intro: ....
Primary invoke URL:
Model-run alias (same body):
The public model id is sume/music-1.0.
When to use
| Goal | Approach |
|---|---|
| Text → music | prompt only |
| Visual conditioning | prompt + optional image_url |
| Avoid specific styles | put exclusions in the positive prompt (for example, “no vocals, no spoken word”) |
Writing the prompt (music brief)
Music 1.0 accepts a text prompt and optional image conditioning. It has no seed, temperature, guidance, or duration parameter. To prevent generic musical choices, write a scene-specific brief with seven axes. These axes are creative directions, not guaranteed output values. Examine the generated audio.
| Axis | Example |
|---|---|
| Emotion (precise, darker permitted) | “hushed, slightly melancholic”, “proud and nostalgic”, “cocky, restless” |
| Genre / lineage | neo-soul, drill, bossa nova, gugak fusion, synthwave, chamber folk |
| Tempo as a number | “72 BPM”, “142 BPM half-time” |
| Key and mode | “D minor”, “E phrygian”, “G major with a lydian lift” |
| Instruments with texture (2–4) | “Rhodes through tape wow”, “gayageum plucks”, “808 with long glide” |
| Arc with one named moment | “breakdown to bass and claps at 0:20, full return at 0:28” |
| Era / production | “1998 production, dry and close”, “2024 hyper-clean” |
End the prompt with one clause: “Instrumental, no vocals.” Add “no spoken
word” only under narration. For scenes in one project that contrast, change
the broad genre family, the tempo (at least 12 BPM apart), and the lead
instrument. If the user requests one consistent score, keep continuity. When it
is applicable, pass the accepted scene still as image_url.
After a policy rejection, change the flagged content, but keep the musical
brief. Try again only in the authorized budget. Do not reduce the request to a
generic bed. Provider lyrics can describe tempo and structure. These lyrics
are model-reported metadata, not an audio measurement.
Hard constraints
- Do not send
durationorduration_seconds. Music 1.0 does not recognize these fields and rejects them. - Do not send a non-empty
negative_prompt. Music 1.0 / Lyria does not support negative prompts. A non-empty value returns HTTP 400 withpublic_reason=negative_prompt_unsupported. Omit the field or send"". - The maximum prompt length is 5000 characters.
- Image URLs must be public HTTPS. Send
nullforimage_urlonly when you intentionally clear an image input on a client that reuses request objects.
Request fields
| Field | Required | Notes |
|---|---|---|
prompt | Yes | 1–5000 characters. Include exclusions in the positive prompt. |
negative_prompt | No | Sume does not support a non-empty value. Omit the field or send "". |
image_url | No | Optional public HTTPS image URL, or null to clear. |
metadata | No | Caller metadata that Sume stores on the job. Sume does not send it to the provider. |
mode | No | async, sync, subscribe, webhook. |
webhook_url | No | Public HTTPS callback for webhook mode. |
wait_timeout_seconds | No | 0–30 for sync / subscribe. |
Create a music job
Generate a Music 1.0 job
POST /v1/music-1.0/generate
Required
Image-conditioned example:
Poll and fetch the result
If the job is successful, read the audio artifact from result.artifacts[]
where type is audio (usually audio/mpeg on media.sume.com).
Artifacts
Completed Music 1.0 jobs return Sume-hosted audio artifacts:
Use Sume media URLs from the result. Raw provider URLs are not public outputs.
Pricing
The price is a fixed $0.125 USD for each accepted Music 1.0 generation. The price does not change with prompt length or optional image conditioning.
Next
- Jobs and results for polls and webhooks
- Recipes for short copy-paste flows