Describe a song, hear it back: text-to-song in practice

TL;DR

Text-to-song turns a sentence into a finished track — arrangement, instrumentation, and vocals if you want them. The sentence has to carry genre, mood, and a concrete subject; "make it happy" gives the model nothing to build on. Leave the musical controls on Auto for the first attempt, then use them to move one thing. In Musy AI the description is up to 150 characters and most songs are ready in about a minute.

The sentence is the whole instrument

Every good result we've had contains three ingredients, and most bad results are missing one:

Assembled: "A lo-fi beat for a rainy night in Tokyo" — genre, mood, scene, twelve words. That's the shape to copy.

Musy AI's create screen on iPhone with a song description field and optional genre, mood, BPM and duration controls.
Description first. Every musical control below it is optional and defaults to Auto.

Leave the controls alone at first

Genre, mood, BPM, and duration can all be set explicitly, and the instinct is to set all of them immediately. Resist it. If your description says "a rainy night in Tokyo" and you also set the genre to Metal, you've asked for two different songs and you'll get the average of them.

The controls earn their place on the second pass, when you have a result that's close: same description, BPM nudged down, mood set explicitly, regenerate. That's a controlled experiment. Setting six things at once isn't.

ControlLeave on Auto when…Set it when…
GenreYour description already names a soundYou want a specific one of the 29 and the text is ambiguous
MoodThe description carries emotionThe result came back emotionally flat
BPMAlmost alwaysThe track is for something time-bound — a workout, a video edit
DurationNever — pick this deliberatelyAlways. It changes the structure, not just the length.

Instrumental or vocal

Instrumental is the safer default for anything the song sits underneath — a video, a podcast intro, background listening. Vocals are the point when the song is the thing itself, or when it's about someone.

With lyrics switched on you can write your own (up to 500 characters) or have them generated from the same description. Writing your own is worth it when specifics matter — a name, an in-joke, a place. Generated lyrics are better than most people expect at genre pastiche and worse at meaning anything in particular.

Vocals can be set to Auto, male, female, or child. You can also record your own voice once in the app and have songs sung in it. That's deliberately scoped to your voice — we don't build cloning of other people's voices, and we'd encourage the same caution wherever you find the feature.

Why the first result isn't the keeper

Generation has real variance: the same description twice gives you two different songs, both valid. That's not a defect to prompt around, it's the medium. The workflow that works:

  1. Write the three-ingredient sentence. Pick a duration. Generate.
  2. Listen once, all the way through. Decide what specifically is wrong — tempo, energy, instrumentation, vocal.
  3. If nothing is specifically wrong but it's dull, regenerate unchanged. Variance alone often fixes it.
  4. If something specific is wrong, change that one thing and regenerate.
  5. Keep the good ones. Generation is cheap; deciding is the slow part.

When to start from a template instead

For occasion songs — birthdays, love songs, songs about a pet — the Lab's templates set genre and mood for you and ask only what that occasion needs to know. It's a faster path than writing a description from scratch, and it's usually a better one, because the template already knows what makes that kind of song work. See the birthday song walkthrough for how that flow goes end to end.

Once you've got a track you like, the prompt-craft rabbit hole is in song prompts that work.

FAQ

How does AI text-to-song work?

You describe the song; the model composes and performs it, returning finished audio. In Musy AI the description is up to 150 characters, generation runs in the background, and most songs take about a minute.

What should the description say?

Genre or reference sound, mood, and a concrete subject. "A lo-fi beat for a rainy night in Tokyo" carries all three.

Should I set genre, mood, and BPM manually?

Not on the first attempt — controls that contradict your text pull the result two ways. Use them to move one specific thing on the second pass.

How long can an AI song be?

30 seconds, 60, 90, or 2 minutes. Shorter is often better: one strong idea beats two minutes of repetition.