I spent the last few weeks reading every AI music paper and release note I could find. The numbers are absurd. But the feeling is empty. This is my attempt to explain why — and what I’d build instead.


The Numbers Are Real

Let me start with what’s actually true, because the hype is not wrong about the facts.

Suno v6 writes eight-minute songs and was co-trained with Warner and BMG. Mureka’s O3 does reflective music reasoning — it writes a section, reviews it, and revises. ACE-Step 1.5 generates a full song on a 4GB graphics card. Google’s Lyria RealTime responds to your input in under two seconds. Open models like YuE, DiffRhythm 2, and HeartMuLa have closed most of the quality gap in two years. In a blind test across eight countries, 97% of 9,000 people couldn’t tell AI music from human music.

By every metric, this is a solved problem.

So why does it feel so empty?

I use AI to write code every single day, and it feels like a tool that belongs to me. AI music feels like a vending machine. You put in a prompt, you get a finished song, and you feel… nothing.

I couldn’t stop asking why two fields with the same technology ended up in such different places.


The Missing Compiler

Here’s the workflow of AI coding, broken down honestly:

  1. You say what you want — in plain language.
  2. The AI assembles context — it reads your repo, your errors, your existing code.
  3. It generates a scaffold — a big, low-risk pile of structure.
  4. You iterate — change a line, rerun, see what breaks.
  5. You verify — it compiles, it runs, the test passes.
  6. You version it — git records every step, and your tools accumulate.

The engine of this entire loop is steps 4 and 5, working together. And step 5 only exists because code has a right answer. Run it and you know. No argument, no debate.

Now map the same six steps onto music. Steps 1–4 transfer almost perfectly: hum a melody instead of describing a requirement, feed it reference tracks and your taste instead of a repo, iterate on one section instead of debugging.

But step 5 has nothing to stand on. Music doesn’t compile. “Good” has no error message.

This is the part where I need to be more precise than I was in my first draft. It’s not that music has no verification — it’s that verification is distributed, delayed, and subjective. You don’t run a song and get a pass/fail. You listen, you wait, you listen again the next morning, you send it to a friend, you try it on different speakers. Your ears lie when they’re tired. The verification loop exists, but it’s slow, social, and never fully closes.

That missing compiler — or more accurately, that slow, distributed, human compiler — is the whole story.


Why the Factories Optimize for the Lowest Common Denominator

When you can’t verify quickly, you outsource the decision to the safest possible thing: templates, hooks, the pop formula, the structure that has worked ten thousand times.

This isn’t a flaw in the companies. It’s the only reasonable move when your objective function is “a song most people won’t skip.” If you have no fast verifier, you optimize for the average. The average is safe. The average is boring.

That’s why generated music so often sounds correct and forgettable at the same time. It nailed the assignment, but nobody ever asked it to mean something.

There’s a second problem, and it’s worse. The evaluators themselves are broken in a specific, measurable way. The 2026 ICASSP ASAE challenge showed that all current models are terrible at distinguishing truly top-tier songs from merely good ones. And reward learning depends entirely on comparisons in that top range. So the generator gets nudged toward “average good,” never toward “great.”

Static evaluators also go stale. The generator keeps improving, but the evaluator keeps scoring the old distribution. The reward signal drifts out of touch. It’s a treadmill.


The Thing I Couldn’t Stop Thinking About: FL Studio Already Got the Interface Right

Here’s what most AI music companies misunderstand about how music actually gets made.

Real composition doesn’t start with structure. It starts with material. A melody you hum in the shower. A drum pattern you stumbled into while testing a sample pack. A chord that surprised you. Structure — verse, chorus, bridge — is discovered afterward, by listening back to what you accidentally made.

FL Studio’s entire design philosophy is built around this. The step sequencer, the pattern-first workflow, the fact that you can stack loops before you ever think about arrangement — all of it says one thing: material comes before structure.

This is exactly where AI music has it backwards. Every platform wants to hand you a finished song. Almost none of them want to hand you a great four bars and let you play.

The first real AI music product won’t be a better song generator. It’ll be a better supplier of material — and I think that’s actually the near-term future the whole field is heading toward without admitting it.

But supplying material isn’t enough. The deeper problem is the loop itself. In music, creation is bidirectional. A human proposes a direction, the AI responds, and the AI’s response becomes the human’s new input. An unexpected harmony, a misaligned sample, a weird timbre — these aren’t errors to be corrected. They’re fuel.

My first draft of this essay modeled music creation as a one-way pipeline: human intent → AI execution → human verification. That was wrong. The real loop is:

human intuition → AI material → human surprise → AI refinement → human intuition…

The AI isn’t just executing. It’s provoking. That’s the difference between a tool and an instrument.


What I’d Actually Build: An Ear That Grows With You

If the missing piece is verification, then the research problem is obvious: build a verifier for taste.

The current approach is broken in a specific, measurable way. Researchers train a static evaluator — something like SongEval — and then use it as a reward signal to fine-tune music models. Two things go wrong:

One: static evaluators go stale. The generator keeps improving, but the evaluator keeps scoring the old distribution. The reward signal drifts out of touch.

Two — and this is the one that gets me — the evaluators are worst exactly where they matter most. As I mentioned above, the 2026 ICASSP ASAE challenge showed that all these models are terrible at distinguishing truly top-tier songs from merely good ones. And reward learning depends entirely on comparisons in that top range. So the generator gets nudged toward “average good,” never toward “great.”

My idea is to stop treating the ear as external and start treating it as something that grows with the generator. I call it EarLoop:

This mirrors how real producers work. Nobody writes a song in one shot. You listen, you change one part, you listen again — often the next morning, because your ears lie when they’re tired.

An evaluator that can point at where something is wrong, instead of only how wrong it is, is the difference between a grade and a note from a teacher.

But there’s a trap here. If the critic and the generator evolve together without an anchor, they can collude. They can drift into a private language of “good” that has nothing to do with human taste. The fix is to anchor the loop in real human signals — not just preference ratings, but behavior: what do people replay? What do they skip? What do they export and finish? What do they come back to the next day?

The ear has to grow with you. But it also has to stay honest.


The Part That’s Probably Too Weird to Work (Yet)

There’s a deeper question underneath all of this, and it’s the one I can’t let go of.

What if the right way to describe music to a machine isn’t text, and isn’t notes, and isn’t audio tokens?

Text is a translation, and translation loses things. I’ve hummed a melody I couldn’t describe in words — that’s not a vocabulary failure, it’s that the thing isn’t verbal. MIDI is the written language of music theory, but theory is what we use to talk about music after the fact, not how it appears in your head. Audio tokens only capture what a sound looks like, not where it’s going.

I think the real unit is movement.

When a chorus “lifts,” a bridge “sinks,” a riff “has a certain drive” — you feel that in your body. Music cognition research calls this embodied cognition, and it’s not a metaphor; it’s how we actually process it. What if we built an AI on that level, on little gestures — a few seconds each — described not by names but by four continuous curves?

I call these 动元 (gestures) — not discrete tokens but a continuous field, where distance means “these feel like the same kind of motion.” Then “understanding” your hum doesn’t mean transcribing it into words. It means catching the direction and force of your movement.

It might be too hard to build. Extracting gestures from raw audio is one problem; generating in that space is another; and quantifying tension and release — which is what truly makes music work — is a third that nobody has really cracked.

But it’s the only framing I’ve found where a machine could learn your taste instead of the average of everyone’s.


The Honest Part

I should be clear about three things, because I hate reading research hype.

First, the technical stuff might not work at the scale I’m imagining. It might end up as a two-page paper saying “we tried, here’s the boundary.” Honestly, that would still be worth publishing — negative results are useful, and I’ve written one before.

Second, and more uncomfortable: personal tools don’t scale as businesses. That’s the exact tradeoff that made coding AI feel personal and music AI feel industrial. A tool tuned to one person’s taste is worth everything to that person and almost nothing to a venture fund. So the money keeps flowing to the factories.

Third, there’s a cultural risk I haven’t fully solved. If everyone uses the same default taste model, music diversity collapses. Personalization isn’t just a feature — it’s a survival requirement for the entire category. The factories optimize for the average because they have to. An instrument optimizes for you because it can.

But I keep coming back to the same conclusion. The question everyone is asking — “can AI make a hit song?” — is boring. The interesting question is: can it learn the shape of your particular taste, and get a little more like you every time you use it?

That’s not a factory. That’s an instrument.

And I’d rather have an instrument that’s 60% as good but 100% mine.


Background reading: the 2026 AI music landscape is fascinating and moving fast — Suno v6, Mureka’s reasoning-based O3, Google’s real-time Lyria, and the open-source wave of YuE, ACE-Step, DiffRhythm, and HeartMuLa. Most of the numbers above come from model papers and the ICASSP 2026 ASAE challenge. The EarLoop and 动元 ideas are mine, and still mostly fiction. The FL Studio observation is not original to me — every producer who has ever stacked four bars of drums before knowing what the song is about already knows it.