
The tool moment
A paper titled “Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations” was first posted to arXiv on 25 December 2024 and was later accepted to the ACM CHI Conference on Human Factors in Computing Systems, held in Yokohama in April and May 2025. Amuse is a research system that accepts an image, a short narrative, or an existing piece of music as inspiration and, through a two-stage process, a multimodal language model followed by a dedicated chord model, converts that input into chord-progression suggestions a songwriter can bring into a session alongside Aria, an existing chord-suggestion tool the paper says is already built into the notation app Hookpad.
What the documents show
The full paper describes a formative study with eight songwriters that shaped Amuse's design, followed by a within-subjects evaluation in which ten songwriters wrote eight-bar choruses from prompts, once using Amuse alongside Aria and once using Aria alone. The authors report participants describing the multimodal condition as leaving them “feeling more guided and aligned with their creative” intentions and reporting “enhanced agency and creativity.” The paper is explicit that no paired dataset of multimodal inputs and matching chords existed before this work, so the chord model was built without one. Neither document states that Amuse itself is available outside this research setting; the paper links only to sound examples and code.
What stays with the musician
Ten songwriters writing short prompted choruses under lab conditions is a study of a design idea, not a claim about how any working songwriter would use the tool across a full song, a catalog, or a career. The paper's own agency and creativity language describes participants' self-reported experience during the study, not an independent measure of the resulting music's quality. A songwriter still has to judge whether an image- or audio-derived chord suggestion actually fits the lyric, the vocal range, and the emotional arc of the song being written, none of which the system evaluates for them.
Judge it by listening
Because the underlying study measured self-reported agency and process, not the finished music, a reader curious about Amuse's actual musical output should treat the paper's sound examples as illustrations of the method rather than as evidence of consistent quality. Judging any chord-suggestion tool by ear means listening for whether the suggested progression serves the song's own melody and lyric, a judgment this ten-person study was not designed to make and does not claim to have made.
- Does a suggested chord progression actually serve this specific melody and lyric, or just sound plausible on its own?
- Would ten songwriters' reported sense of agency in an eight-bar lab task hold up over a full song written across multiple sessions?
- Is this tool, or anything like it, available outside the research setting described in the paper?
A CHI-reviewed study with ten participants is a credible signal about a design direction, not a verdict on the tool's musical output. The paper's own scope, a short prompted task and self-reported agency, is worth restating alongside its more quotable findings.
Sources & reading trail
The arXiv abstract page gives the paper's title, authors, submission date and short abstract describing Amuse's multimodal chord-suggestion design and a user study with songwriters.
Source published: 25 December 2024 · Retrieved: 16 September 2026
The full paper describes the within-subjects user study (N=10 songwriters comparing Amuse-plus-Aria against Aria alone) and an earlier formative study (N=8 songwriters), plus confirmation of CHI 2025 acceptance and that only sound examples and code are shared publicly.
Source published: 25 December 2024 · Retrieved: 16 September 2026
Documentation, papers and the makers' own records establish the note; the judgment about what stays with the musician is Mix & Meaning editorial analysis. This retrospective draft does not imply the site published on the event date.