
The tool moment
A study published 15 January 2025 in Empirical Studies of the Arts, titled “AI Performer Bias: Listeners Like Music Less When They Think it was Performed by an AI,” tested whether an already-documented bias against AI-composed music extends to performance. In an online cross-over experiment with 120 participants, listeners watched three videos of classical piano pieces in two versions built from identical audio: one showing “a professional pianist who pretended to play,” the other showing the piano “playing automatically, allegedly thanks to an AI.”
What the documents show
The paper states participants rated the performances “as more likeable, engaging, higher in emotional valence, and of higher quality” when attributed to the pianist, an effect the authors report was insensitive to listeners' own musical expertise but moderated by their attitudes toward AI. The authors are explicit about what the design does and does not test: “the current work does not directly compare real and AI-driven musical performances; instead, it simply aims to analyse the impact of leading an audience to believe that a performance is AI-driven.” A publication record confirms the venue and date; this is a single published study, not a synthesis of the field.
What stays with the musician
The study measured a labeling effect, not a listener's ability to tell two actual performances apart, and it says so directly. A musician deciding whether to disclose an AI-assisted process is not choosing between two audibly different products in this data; they are choosing how attribution itself will shape an audience's judgment before a note is heard. What the study cannot tell a musician is whether their own performance, genre, or audience would show the same size effect, since the paper used one performer and three classical piano excerpts.
Judge it by listening
The authors' own limitations note that participants, when asked what differences they heard between the two labeled versions, “confabulated about differences in rhythm, tempo variations, dynamics, and dissonances” that were not actually there, since the audio was identical. That is the paper's finding, not an invitation to distrust every listener's ear; it is a caution specific to blind-versus-labeled comparisons. A reader running their own comparison should keep the audio identical across conditions and vary only the stated attribution, exactly as this design did, before drawing conclusions about what attribution changes.
- Would this same liking gap appear with a different genre or performer than the classical piano excerpts used here?
- Is a listener's stated preference about performance quality, or about the label attached to it?
- If the audio is identical, what does a listener's confident description of “differences” actually reveal?
A single, well-designed study on attribution is not evidence that listeners can detect AI performance in real listening conditions; it is evidence that the label alone, independent of the sound, moves judgment. The two claims are not the same, and the paper does not conflate them.
Sources & reading trail
The published study's abstract, method (N=120 cross-over online experiment, identical audio in two attributed conditions) and findings that human-attributed performances rated more likeable, engaging, higher in emotional valence and higher quality; states the study does not directly compare real and AI-driven performances.
Source published: 15 January 2025 · Retrieved: 16 September 2026
Confirms the paper's journal (Empirical Studies of the Arts), seven authors, and 15 January 2025 publication date.
Source published: Not established · Retrieved: 16 September 2026
Documentation, papers and the makers' own records establish the note; the judgment about what stays with the musician is Mix & Meaning editorial analysis. This retrospective draft does not imply the site published on the event date.