How AI Makes Music: A Middle School Lesson Plan
A student drops a pair of earbuds on your desk after lunch and says, “I made that during passing period.” Thirty seconds in Suno. Vocals, beat, chorus, done. That’s the moment this lesson exists for — because when you ask what’s happening inside the tool, most curricula go quiet. You don’t need a music background to teach this. The explanation is a computer science and media-literacy question, not a music theory question, and this lesson plan works in ELA, social studies, science, or any advisory period that has 40 minutes.
TL;DR: AI music tools are trained on massive collections of existing recordings. They compress audio into small units called tokens, learn the statistical patterns, then predict or gradually refine a new token sequence and convert it back into a playable waveform. The core idea — predict the next piece based on everything that came before — is the same one behind text AI like ChatGPT. Beyond the mechanism, two live questions belong in your lesson: a federal copyright lawsuit that is still active as of August 2026, and a documented tendency to flatten non-Western musical styles toward Western pop. Everything that follows gives you the words, the table, the 40-minute plan, and the standards crosswalk to walk in Monday.
How does AI actually make music?

The short answer: AI music is not composed. It is predicted — or, in some systems, gradually denoised from static into sound. Four steps get you from “train on recordings” to a playable audio file.
| Step | What happens | Classroom analogy |
|---|---|---|
| 1. Train on audio | The model studies a large collection of recordings and learns which patterns of sound tend to follow which. Meta’s MusicGen trained on roughly 20,000 hours across ~400,000 licensed recordings. Suno stated in an August 2024 court filing it trained on “tens of millions” of recordings scraped from the internet. | Studying every meal in a restaurant’s history to learn what flavors the chef tends to combine. |
| 2. Tokenize audio | Raw audio runs at ~44,100 data points per second — far too large to model directly. A neural audio codec (Meta’s EnCodec, for example) compresses sound into discrete tokens: a compact vocabulary of small sound-chunks. | Turning a full paragraph into a series of LEGO blocks so a pattern-matcher can work with them. |
| 3. Generate | Either an autoregressive transformer predicts the next token (same idea as ChatGPT predicting the next word), or a diffusion model starts from noise and refines it step by step toward a target. Google’s Lyria 3 uses latent diffusion. | Autocomplete finishing your sentence — or an artist erasing a scribble in stages until a portrait appears. |
| 4. Decode | The codec turns the token sequence back into a waveform the speaker can play. | Running the LEGO blocks back through a translator to get a readable paragraph. |
The tokenize-then-predict loop is covered in depth in our post on how generative AI works and why it matters for grade 7 — the next-token idea is identical whether the model is writing text or composing a beat. Connecting the two lessons takes about five minutes of direct comparison.
For a more technical treatment of audio codec design, Sound On Sound’s explainer on how AI music works is readable without a signal-processing background.
Which AI music tools are your students using?
By the time a school year starts, students have typically already encountered at least one of these — often without a teacher ever naming what they were using.
- Suno — free tier, text-to-full-song with vocals. The most common one students mention by name. Trained on internet audio (see the copyright note in the next section).
- Udio — similar text-to-song interface; also offers stem export, which lets students isolate vocals or instruments — useful for critical listening activities.
- Google MusicLM / Lyria (available in Gemini) — Google’s system embeds a SynthID audio watermark in every output, an invisible signal detectable by Google’s own tools. Lyria 3 uses licensed training data.
- Meta MusicGen — open-source, with publicly documented training data. The transparency makes it the most classroom-friendly option when students are comparing “what did the model learn from?”
For a lesson that puts these tools in context of the broader AI landscape — narrow AI vs general AI, generative vs retrieval — the Types of AI Deep-Dive lesson gives students the vocabulary to compare Suno and MusicGen on their own terms.
Where AI music breaks down — and what to listen for

Three detection cues students can hear without any music theory background. Each one is also a research task:
-
Long-form structure collapse. AI-generated tracks tend to hold together for roughly 30–60 seconds before harmonic patterns drift or repeat in ways a human arranger would catch. Peer-reviewed analysis of AI-generated audio confirms structural coherence degrades beyond that window. Classroom move: map a 3-minute AI track on paper — mark where the verse/chorus pattern changes or disappears around the 90-second point.
-
Lyric and phoneme mismatch. In tracks with vocals, the model is predicting both the melody token sequence and the voice-sound token sequence simultaneously. The two sometimes drift: a word that sounds like one thing turns out to be another when students try to transcribe it. Have students write down what they hear in a 30-second clip, then compare to the original prompt text.
-
Homogenization toward the training data average. Because most training datasets skew toward Western popular music, non-Western styles get flattened toward pop conventions. ISMIR 2025 peer-reviewed research documented this pattern. Classroom move: prompt “traditional West African kora music” and “pop ballad” — compare what the same model produces, then discuss what the training data probably looked like.
The copyright question students should be asking
On June 24, 2024, Sony Music, Universal Music Group, and Warner Music Group filed suit against Suno and Udio in federal court, alleging large-scale unlicensed copying during training. The RIAA announcement described it as a “landmark” case for how copyright applies to AI training data. Suno’s defense: the training constitutes fair use — “like a person learning by listening.” Warner Music reached a licensing deal with Suno in November 2025. Sony and Universal’s cases are continuing as of August 2026.
On the legislative side, the NO FAKES Act — a bipartisan bill addressing AI-generated voice and likeness — passed a Senate Judiciary Committee vote in June 2026 but has not been signed into law. Full bill text.
This is the gap. Students open Suno in the hallway and make a song in 30 seconds. Nobody handed them the words for what that action means legally, who the training data belonged to, or what a “licensing deal” actually settles. That is a literacy gap, not a moral failing. The who-owns-AI-generated-art lesson covers the copyright side in full, including the 2025 US Copyright Office ruling and how to run a four-corners classroom debate on who should own the output.
A no-device lesson: “Be the Music Model”

Many schools block Suno entirely. This 40-minute plan works with zero devices and no music theory knowledge — it models the same four-step mechanism using printed cards and group discussion.
Materials: printed “sound-chunk” cards (a set of ~20 short melodic fragments written in simple notation or described in words like “low-low-high-pause”), a timer, an exit ticket slip.
| Time | Activity | What it models |
|---|---|---|
| 0–5 min | Play or describe a 30-second AI-generated clip. Ask: “How did a computer make this?” Collect answers without confirming or correcting. | Activating prior knowledge / surfacing misconceptions |
| 5–15 min | Distribute the sound-chunk cards. Groups arrange them into a sequence that sounds like a complete musical phrase. | Tokenize → sequence (Step 2 and 3 above) |
| 15–25 min | ”Predict the next chunk” round: each group extends the sequence by adding two more cards, choosing the one that “fits best” by statistical feel rather than music rules. | Autoregressive generation — next-token prediction |
| 25–35 min | Swap to “denoise from scramble”: give groups a randomized version of another group’s sequence; they must sort it back toward coherence. Discuss where the result got weird. | Diffusion — refining from noise toward signal |
| 35–40 min | Exit ticket: “Name one thing the model does NOT understand about the music it made.” | Metacognitive close — surfaces limitations |
A typical class finishes the predict-round with at least one group producing a sequence that sounds plausibly musical and at least one group producing something that falls apart at the 8th card — exactly what happens in AI output around the 90-second mark. That asymmetry is the discussion.
For a full set of no-device AI lesson structures, the Unplugged AI Activities pack includes 10 lessons in the same format — card sorts, role-plays, and paper investigations for grades 6–8.
Standards this lesson covers
| Standard code | What it says | Which lesson step |
|---|---|---|
| AI4K12 Big Idea #3 | Learning — computers learn from data | Training-data step (Step 1) |
| AI4K12 Big Idea #2 | Representation & Reasoning | Tokenize step (Step 2) |
| AI4K12 Big Idea #5 | Societal Impact | Copyright discussion |
| ISTE 1.3.d | Build knowledge by actively exploring real-world issues and problems | Unplugged investigation |
| ISTE 1.2.c | Demonstrate an understanding of and respect for the rights and obligations of using and sharing intellectual property | Copyright strand |
| CCSS.ELA-LITERACY.SL.7.1 | Engage effectively in collaborative discussions | Discussion and exit ticket (35–40 min) |
| CCSS.ELA-LITERACY.W.7.8 | Gather relevant information from multiple sources; assess credibility and accuracy of each | Research task on the lawsuit and NO FAKES Act |
Standards source: ISTE Student Standards. ISTE is a registered trademark of the International Society for Technology in Education. These resources are not affiliated with or endorsed by ISTE.
Ready to teach it
The gap this lesson fills is not a music gap. It is a vocabulary gap — the words for what is happening inside the tool, the rights question the tool sidesteps, and the critical listening habit that lets students hear where the model loses the plot. You now have all three. For a 16-lesson sequence that builds from the generative AI mechanism through hallucination, bias, copyright, and societal impact, the How AI Works MEGA bundle covers the full arc — printable, standards-documented, and built for teachers who do not have a computer science background any more than they have a music one.
This post was drafted with AI assistance and human-finalized.
Quick questions
AI music tools are trained on large collections of existing recordings. They compress the audio into small units called tokens, learn the statistical patterns, then either predict the next token (like ChatGPT predicting the next word) or refine a token sequence out of noise, and finally decode the tokens back into a playable waveform. The music isn't composed — it's predicted.
That's unsettled. In June 2024, major record labels sued Suno and Udio for copying recordings during training. Suno's defense is fair use — that learning from music is like a person learning by listening. Warner reached a licensing deal with Suno in late 2025, but the Sony and Universal cases were still active as of August 2026, so ownership of AI music output remains a live legal question.
Get the free AI-Proof Assignment Toolkit
10 ways to redesign any assignment so an AI chatbot structurally can't do it — plus a redesign worksheet, a 45-minute lesson, the “Spot AI Work” card, and parent templates. One email, all 5 pieces.
Prefer the full breakdown? See everything inside the toolkit →