Podcast brief
Can AI Music Be Creative?
Barry Lam explores AI vocal emulation, AI composition, and AI improvisation with Robin James, Theodore Gracyk, and two guitarists to ask what counts as human creativity in music.
A Stoa brief of an episode from Hi-Phi Nation
Rise of the Music Machines
Listen to the original episode
The episode belongs to its makers. This page summarizes its arguments in Stoa's words and points to the moments where they are made. Brief updated .
The brief
This conversation asks what AI music technologies reveal about human creativity, ownership, and the nature of musical experience. Host Barry Lam works through three cases: vocal emulation, generative composition, and improvised solos, each pressing a different philosophical question.
Lam starts with vocal emulators, tools that let a producer keep their own lyrics, rhythm, and performance while swapping in a celebrity's vocal timbre. He argues this is a mistake to call AI-generated music at all, since a human is doing the actual performing and simply playing the emulated voice like an instrument. He places this practice between sampling, which reuses an actual recording, and impersonation, which is a human skill with no ownership claim attached, using Marc Martel's Freddie Mercury vocals as a case where no estate holds rights. Lam predicts that voice timbre will follow the path of song catalogs bought by companies like Hipgnosis, eventually becoming a licensed corporate asset, which makes the current legal gray zone a narrow window of creative freedom for amateur producers.
The second case turns to generative composition models that mimic Bach or Mozart. Theodore Gracyk argues that because AI can convincingly reproduce these styles, it exposes something about human composers themselves: most, even celebrated ones, settle into a pattern early and spend a career varying it rather than continually inventing. He adds that attentive concert-hall listening, often treated as the serious way to engage music, is historically unusual, a nineteenth-century habit, while most people use music as background to other activities. On his view, AI-generated background music is not degraded music but representative of how music is normally used.
Robin James pushes back. She argues that treating composition as pattern generation on a page assumes an isolated, heroic model of creativity, when music is actually a participatory and social practice. A privately owned AI system cannot take part in that collaborative process, so it cannot replace human music-making regardless of how convincing its output sounds.
The final case compares AI-generated and human-improvised blues and metal solos. Guitarists Fabrizio and Keshav judge the AI solos technically fluent but lacking narrative arc, the development and resolution that make a human solo feel like it is telling a story. The episode leaves open whether that storytelling quality is simply a pattern AI has not yet learned, or something that stays distinctly human.
The conversation does not resolve the disagreement between Gracyk and James, but it gives you a way to ask, next time a machine-made track sounds convincing, what exactly you are responding to: a pattern, a performance, or a story.
Strongest arguments
Vocal emulation is human performance, not AI generation
6:00Lam argues that when a producer raps or sings and a vocal emulator swaps in a celebrity's timbre, the words, rhythm, and performance remain fully human. He concludes it is a mistake to call the resulting track AI-generated, since the human is playing the deepfaked voice like an instrument.
Vocal emulation sits between sampling and impersonation
8:30Lam argues that vocal emulators are legally and metaphysically distinct from sampling a recording, since no actual recorded audio is reused, and distinct from ordinary impersonation, since a machine trained on samples produces the mimicry. He illustrates the impersonation side with Marc Martel's Freddie Mercury vocals, which carry no ownership claim from Mercury's estate.
Voice timbre will become a corporate-owned asset
11:40Lam predicts that as with song catalogs bought by firms such as Hipgnosis, celebrity voice timbre will eventually be licensed and monetized, likely ending up owned by corporations rather than the artists themselves. He says the current period, before lawsuits and licensing regimes exist, is the most creative window for amateur producers.
AI composition reveals that human composition is mostly pattern-following
15:14Gracyk argues that because deep learning models can easily generate convincing Bach or Mozart style compositions, this shows that even celebrated composers largely work by finding a pattern early in their careers and then reproducing variations of it, rather than continuously innovating.
Attentive concert listening is a historical anomaly, not the norm for music
23:29Gracyk argues that most people use music as background to another activity such as exercise, work, or study, and that treating focused concert-hall listening as the standard way to relate to music is a 19th century invention. He concludes that AI-generated background music is not degraded music but representative of typical human musical experience.
Music is participatory, so AI cannot replace human musical creativity
25:41Robin James argues that treating music as just notes on a page wrongly centers a heroic individual model of creativity, when in fact music-making is a social and collaborative practice. She argues that privately owned AI systems cannot participate in that collaborative process, so AI is not poised to replace human musical creativity even if it can generate interesting sounds.
Human improvised solos have narrative structure that AI solos lack
46:53Guitarists Fabrizio and Keshav, comparing their own improvised blues and metal solos to machine-generated ones, argue that human solos build a musical idea, develop it, and resolve it like a story, while the AI solos generate technically fluent note patterns with no development, repetition with variation, or emotional arc.
Disagreements
Whether AI composition undermines the value of human composition
15:22Gracyk holds that AI's ease in mimicking Bach and Mozart shows that most human composition, even celebrated composition, is low-creativity pattern reproduction after an initial breakthrough. Robin James instead treats the comparison as revealing an ideological bias in AI research and criticism, arguing that framing composition as isolated pattern-generation ignores that music is a participatory, conversational practice among people.
Philosophers and works discussed
- Robin James
- Theodore Gracyk
Questions this episode answers
-
Is a track made with an AI vocal emulator actually AI-generated music?
6:00Barry Lam argues no: the words, rhythm, and performance remain human, and the AI only swaps in the timbre of another person's voice, so the human is playing the emulated voice like an instrument.
-
Is AI vocal emulation more like sampling or more like impersonation?
8:30Lam argues it is exactly in between: like sampling, it depends on a machine learning from an artist's recordings, but like impersonation, the resulting mimicry is newly generated rather than a reuse of an actual recording, leaving unresolved ownership questions.
-
Does AI's ability to mimic Bach and Mozart show that human musical creativity is mostly pattern-following?
15:14Theodore Gracyk argues yes, because a deep learning model can easily generate convincing Bach or Mozart style pieces, which suggests composers themselves settle into a pattern early in their careers and mostly vary it afterward rather than continually innovating.
-
Is attentive concert-hall listening the normal way humans relate to music?
23:29Gracyk argues it is not: most people use music as background to another activity like exercise or work, and treating focused listening as the norm is a 19th century cultural anomaly, which is why AI-generated background music fits how most people actually use music.
-
Can AI ever replace human musical creativity?
25:41Robin James argues no, because music is a participatory and social practice rather than an individual production of notes, and privately owned AI systems cannot take part in the collaborative process that makes music meaningful.
-
What is missing from AI-generated guitar solo improvisation compared to human improvisation?
46:53Guitarists comparing their own solos to machine-generated ones argue that AI solos are technically fluent but lack narrative development, contrast, repetition with variation, and an emotional arc, the qualities that make a human solo feel like it is telling a story.
Sources
- Hi-Phi Nation, Rise of the Music Machines Podcast episode, original episode