A raw-audio milestone with visible limits

OpenAI published Jukebox on April 30, 2020, with code, weights, a paper and a sample explorer. It described a system that generated music including rudimentary singing in raw audio, conditioned by genre, artist and lyrics. The release also stated its own limits: local coherence did not reliably become familiar large-scale form, upsampling could introduce noise, and sampling was slow enough that OpenAI did not position it as interactive. That candor is part of the historical record, not an embarrassment to edit out.

Why raw audio changes the producer’s question

Symbolic systems leave a musician with notes to orchestrate. Raw-audio generation arrives already sounding like a record: voice color, room impression, drum texture and mix balance arrive together. That can be compelling, but it makes revision harder. If the pre-chorus needs one fewer chord, the producer cannot assume a clean piano-roll edit is available. The apparent completeness of the result can conceal the absence of handles.

Jukebox’s useful lesson is to ask whether a workflow has a revision surface before admiring its surface fidelity. A song needs places to change tempo, rewrite text, retune a performance, clear space for dialogue, or revoice a harmony. When all those decisions are fused into generated audio, the human role shifts toward selection, framing and reconstruction. That may be appropriate for a reference mood board; it is a fragile basis for a session that needs precise client revisions or collaborators’ performances.

Listen for form, not just texture

In a retrospective listening pass, ignore the novelty of the first ten seconds. Mark where a phrase establishes a promise, where the next section changes the promise, and whether a later return earns recognition. If the output has a timbral idea worth keeping, translate that observation into production language: dry close vocal over distant drums; low-register guitar doubled by a noisy synth; an uneven turnaround. Rebuild the relationship rather than treating a generated sample as a composition already solved.

This article documents OpenAI’s April 2020 research release and its stated constraints. It makes no assertion about current models, samples, access, policies or use rights.

Make a two-column note while listening: in one column, write the audible quality you want; in the other, write controllable means by which a band, programmer or mixer could make it. An uneasy lift might become a delayed crash, a held bass note, and a late vocal entrance. The translation preserves a listening discovery while returning authorship to choices a session can rehearse and revise.

That reconstruction can reveal whether the original attraction was harmony, orchestration, phrasing, or the illusion of finished sound.

Arrangement test

Can you name a change to make at bar 17? If not, you may have a recording to admire rather than material to produce with.

The original record

Original source review: 16 September 2026; individual records retain later checks. The workflow analysis is editorial interpretation; it does not report an in-house product test.

Continue the reading path

Trace the idea

Next: MusicLM put the brief inside the generation loop. ↗See the whole path →