
The tool moment
A published comparison stating that listeners rated an AI stem-separation output well against the original is very likely reporting a MUSHRA score. The ITU publishes Recommendation ITU-R BS.1534, Method for the subjective assessment of intermediate quality level of audio systems, known by the acronym MUSHRA: MUltiple Stimuli with Hidden Reference and Anchor. Its version record shows the current revision, BS.1534-3, was approved in October 2015. Where BS.1116 is built for barely audible differences, MUSHRA exists for the more common case in a musician's world: comparing several processed versions of a track against each other and a known reference when the differences are clearly audible but hard to rank consistently by ear alone.
What the documents show
The recommendation's own text describes a single screen on which a listener can switch at will between an open reference, several test conditions, a hidden copy of the reference, and one or more hidden low-quality anchors, scoring each on a 0-100 scale. This design lets a listener directly compare impaired versions against each other, which a one-at-a-time method struggles to do reliably. Tellingly, the recommendation cautions that the MUSHRA name is "often misused" for tests that skip a real hidden reference and anchor, and that scores are only meaningful when most conditions land roughly in the 20-80 range the method was validated for. That is the ITU's own caution, not this outlet's opinion. As a living standard, itself revised across BS.1534-1 in 2001 and BS.1534-3 in 2015, this description reflects the document as retrieved on 16 September 2026.
What stays with the musician
Nothing in the standard tells a musician which stem separator, mastering chain, or plugin to trust; it only describes how a rigorous comparison should be structured if someone runs one. Whether a given published MUSHRA test actually used a hidden anchor, a real hidden reference, and a properly calibrated scale is a question the musician has to ask of whoever ran it, including this site whenever it cites one.
Judge it by listening
Editorially, a chart labeled MUSHRA results is worth treating with skepticism until the anchor and hidden reference are visible in the reported method, since an informal online poll with sliders is not the instrument the ITU standardized.
- Does the reported test include a hidden reference and a hidden anchor, or only labeled options?
- Do the scores fall in the range the method is meant to produce, or cluster near the extremes?
- Were listeners free to switch instantly between all versions, or did they hear them once in sequence?
MUSHRA is a tool for producing a defensible comparison, not a badge a company can attach to any listening survey. The distinction the ITU draws between its formal method and casual reuse of the acronym is worth carrying into any claim built on the name.
Sources & reading trail
States MUSHRA's multi-stimulus, hidden-reference-and-anchor design, its 0-100 scoring range, and its own caution against misuse of the MUSHRA name.
Source published: Not established · Retrieved: 16 September 2026
Confirms the recommendation's version history (BS.1534-1 through BS.1534-3) and current in-force status.
Source published: Not established · Retrieved: 16 September 2026
Provides the contrasting small-impairment method that BS.1534 cites when explaining why the two methods differ in scope.
Source published: Not established · Retrieved: 16 September 2026
Documentation, papers and the makers' own records establish the note; the judgment about what stays with the musician is Mix & Meaning editorial analysis. This retrospective draft does not imply the site published on the event date.