Mobidoo

Music 1.0

Music 1.0 is retiring gradually. Its routes keep working and keep job.model = sume/music-1.0, but every request now resolves through the Music Router (sume/music-auto). New integrations should call POST /v1/music-router/generate.

Use Music 1.0 for prompt-driven music generation. The request is text-first; optional image conditioning is supported. Provider model ids stay internal. Music 1.0 runs on Google Lyria 3.5: full-length structured songs up to a few minutes, steered by the prompt (say "a 2-minute track" or use section markers like [0:00-0:30] Intro: ...).

Primary invoke URL:

Model-run alias (same body):

Public model id: sume/music-1.0.

When to use

GoalApproach
Text → musicprompt only
Visual conditioningprompt + optional image_url
Steer away from stylesput exclusions in the positive prompt (e.g. “no vocals, no spoken word”)

Writing the prompt (music brief)

Music 1.0 accepts a text prompt and optional image conditioning, with no seed, temperature, guidance or duration parameter. To avoid generic musical choices, write a scene-specific brief with seven axes. These are creative directions, not guaranteed output settings; verify the generated audio.

AxisExample
Emotion (precise; darker allowed)“hushed, slightly melancholic”, “proud and nostalgic”, “cocky, restless”
Genre / lineageneo-soul, drill, bossa nova, gugak fusion, synthwave, chamber folk
Tempo as a number“72 BPM”, “142 BPM half-time”
Key and mode“D minor”, “E phrygian”, “G major with a lydian lift”
Instruments with texture (2–4)“Rhodes through tape wow”, “gayageum plucks”, “808 with long glide”
Arc with one named moment“breakdown to bass and claps at 0:20, full return at 0:28”
Era / production“1998 production, dry and close”, “2024 hyper-clean”

Close with one clause: “Instrumental, no vocals.” (add “no spoken word” only under narration). For contrasting scenes in one project, vary broad genre family, tempo (at least 12 BPM apart) and lead instrument. Preserve continuity when the user requests one consistent score. Pass the accepted scene still as image_url when appropriate. On a policy rejection, revise the flagged content while retaining the musical brief; retry only within the authorized budget. Do not strip the request to a generic bed. Provider lyrics may describe tempo and structure; this is model-reported metadata, not an audio measurement.

Hard constraints

  • Do not send duration or duration_seconds. They are unrecognized and rejected on Music 1.0.
  • Do not send a non-empty negative_prompt. Music 1.0 / Lyria does not support negative prompting; non-empty values return HTTP 400 with public_reason=negative_prompt_unsupported. Omit the field or send "".
  • Prompt max length is 5000 characters.
  • Image URLs must be public HTTPS. Send null for image_url only when you intentionally clear an image input on a client that reuses request objects.

Request fields

FieldRequiredNotes
promptYes1–5000 characters. Include exclusions in the positive prompt.
negative_promptNoUnsupported when non-empty. Omit or send "".
image_urlNoOptional public HTTPS image URL, or null to clear.
metadataNoCaller metadata stored on the job; not sent to the provider.
modeNoasync, sync, subscribe, webhook.
webhook_urlNoPublic HTTPS callback for webhook mode.
wait_timeout_secondsNo0–30 for sync / subscribe.

Create a music job

Generate a Music 1.0 job

POST /v1/music-1.0/generate

Required

Image-conditioned example:

Poll and fetch the result

On success, read the audio artifact from result.artifacts[] where type is audio (typically audio/mpeg on media.sume.com).

Artifacts

Completed Music 1.0 jobs return Sume-hosted audio artifacts:

Use Sume media URLs from the result. Raw provider URLs are not public outputs.

Pricing

Fixed $0.10 USD per accepted Music 1.0 generation. Price does not vary by prompt length or optional image conditioning.

Next