# Describe a song, hear it back: text-to-song in practice

By the GO AI Team · August 21, 2026

> **TL;DR:** Text-to-song turns a sentence into a finished track — arrangement, instrumentation, vocals if you want them. The sentence must carry **genre, mood, and a concrete subject**; "make it happy" gives the model nothing. Leave musical controls on Auto first, then move *one* thing. In [Musy AI](https://goaichat.app/musy) the description is up to 150 characters and most songs are ready in about a minute.

## The sentence is the instrument

- **A sound to aim at** — genre, era, or instrument. "Lo-fi," "90s alt-rock," "solo piano," "Afrobeat."
- **A mood** — "wistful," "triumphant," "menacing," "warm."
- **A subject or scene** — the ingredient people skip, and the one that matters. "A rainy night in Tokyo" is a scene; "chill vibes" isn't.

Assembled: *"A lo-fi beat for a rainy night in Tokyo."* Twelve words, all three ingredients.

## Leave the controls alone at first

If your text says "rainy night in Tokyo" and you also set genre to Metal, you asked for two songs and got the average.

| Control | Auto when… | Set it when… |
|---|---|---|
| Genre | Description names a sound | You want a specific one of the 29 |
| Mood | Description carries emotion | The result came back flat |
| BPM | Almost always | The track is time-bound — workout, video edit |
| Duration | Never | Always — it changes structure, not just length |

## Instrumental or vocal

Instrumental for anything the song sits *underneath* — video, podcast intro, background. Vocals when the song is the thing itself.

With lyrics on, write your own (up to 500 characters) or generate them from the description. Write your own when specifics matter — a name, a place, an in-joke.

Vocals: Auto, male, female, or child. You can record your own voice once and have songs sung in it. That's deliberately scoped to **your** voice — we don't build cloning of other people's.

## Why the first result isn't the keeper

Same description twice gives two different songs, both valid. That's the medium, not a defect.

1. Three-ingredient sentence, pick a duration, generate.
2. Listen all the way through. Name what's specifically wrong.
3. Nothing specific but dull? Regenerate unchanged — variance often fixes it.
4. Something specific? Change that one thing.
5. Keep the good ones. Generation is cheap; deciding is slow.

## Or start from a template

For birthdays, love songs, songs about a pet, the Lab's templates set genre and mood and ask only what that occasion needs. Usually faster and better — see the birthday walkthrough.

Prompt craft: song prompts that work.

---

We build [Musy AI](https://goaichat.app/musy) — describe a song, hear it back, with cover art generated to match.

More: [GO AI Blog](https://goaichat.app/blog) · [support@goaichats.com](mailto:support@goaichats.com)
