Somewhere in an imaginary conservatory, a chatbot has arrived for its composition lesson. It has no pencil, no impressive scarf, and no tragic story about the railway journey. It does, however, have a four-part chorale in G minor. The teacher prepares the red pen. This is usually the enjoyable part.
Then things become awkward. The homework appears to be rather good.
In Auggie’s Bach Benchmark example, OpenAI’s GPT-6 Astra was asked to write a chorale in Bach’s style, in 3/4 time, using LilyPond notation. Auggie reports his best result so far: no voice-leading errors, a Neapolitan sixth chord, and passing tones, which he says previous models had never produced in this test. The experiment used Extra High reasoning effort. Apparently, even artificial intelligence benefits from being told to concentrate.
These details deserve some attention. Voice leading concerns how individual musical lines move while combining into harmony. Imagine four people carrying a sofa downstairs, each with a different opinion about the destination. Keeping everything coordinated takes work. Passing tones connect chord tones through intervening notes, giving a line movement between its harmonic resting places. The Neapolitan sixth adds a chromatic colour and has the considerable advantage of sounding like something you could order after dinner.
What interests me is the coordination behind the result. A chorale asks for decisions that work both across time and among simultaneous voices. A pleasing chord can create trouble in the next bar. A graceful melody can inconvenience everybody downstairs. If Auggie’s assessment holds up, Astra is showing useful competence in managing these relationships. That is a substantial achievement for a system many people still principally associate with drafting emails.
There is a small, revealing wrinkle. Auggie clarified that the older comparison used ordinary MIDI playback, while Astra’s example used MuseSounds. The instrumental presentation therefore differed. Richer playback can influence our impression of a composition. A fair listening comparison would give both scores the same instruments. Even a musical revolution should occasionally check whether someone has simply bought better speakers.
The mechanism is also worth understanding. OpenAI’s documentation lists Astra’s output modality as text. In this experiment, its contribution is written musical notation, which other software turns into a score and sound. That gives us something unusually interesting: musical decisions available for inspection. You can point at a bar, examine a voice, change a note, and investigate why the whole thing suddenly sounds as though the organist has received disappointing news.
LilyPond’s text-based approach makes that inspectability practical. The composition exists in a form that can be edited and engraved. For a musician, access to the working material matters. You can keep the phrase you like, rescue the bass line, or throw away everything except two bars. Two useful bars can justify an afternoon. Composers have certainly spent longer acquiring less.
The second post brings this possibility into professional life. Musician Michael Wall describes Astra as a major advance for his own work and points to seven advanced musical use cases. His public enthusiasm is striking because it comes from someone speaking about his practice. He sees room to build and explore. There is something infectious about a working musician encountering new possibilities and immediately wondering what to try.
Imagine the resulting rehearsal. A composer sketches a melody, asks for several harmonisations, listens, rejects most of them, and keeps one unexpected turn. Then come more specific requests: leave space for the singer; make the inner parts easier; let the ending remain unsettled. These are illustrative possibilities, of course. Their value would depend on how reliably the system follows musical intentions. But the prospect of conducting such experiments through ordinary language is enormously appealing.
Experience may become more valuable in this arrangement. A beginner can ask for “something sad”; a composer can hear that the answer needs a delayed resolution, a less predictable bass, or simply silence. Knowing what to ask, and recognising what deserves to survive, are musical skills. A faster pencil gives those skills more opportunities to matter.
It could also make learning more active. A student could compare alternative settings of the same tune, hear how a moving bass changes the atmosphere, and test a rule by breaking it. A teacher would still need to catch confident nonsense. Teachers already possess considerable experience in that department. The opportunity is to make explanation audible and revision immediate, while keeping judgement firmly in the room.
Musical correctness has limits. A piece can obey every rule and leave you checking the programme for the interval. A departure from convention can supply its most memorable moment. The interesting test will be sustained musical purpose: whether an idea develops, whether surprise feels earned, whether the ending changes how we remember the beginning. A chorale exercise gives us evidence about craft. Larger artistic claims require larger encounters.
Still, I find these examples encouraging. Useful assistance could let musicians spend more time trying ideas that previously seemed too laborious to pursue. That freedom invites curiosity, including the sort that produces peculiar failures. An abundance of competent drafts would also make selection more demanding. Taste gets extra homework.
Back in our imaginary conservatory, the teacher finally puts down the red pen and sits at the piano. There are questions about phrasing, a doubtful cadence, and one passage worth trying again. The chatbot has brought enough musical substance to make the discussion interesting.
That seems an excellent reason to pull up another chair.




No comments yet