I think this is a great project and very HN. Not sure why the comments are so focused on the deliverable- you learned way more and had a much more interesting experience.
One think I didn't see mentioned in the post- maybe I missed it- how large was the data? How many samples did you use to pretrain and post-train
You can also listen to the transcript of four Russian composers, including Rachmaninoff, playing this pattern recognition and generation game at a dinner party in the late 1800’s: https://youtu.be/PlFPOWuwBHI?is=EKBK7QQkJs4MsTCU
Composers at the time could do this just by looking at sheet music and audiating, without using a piano.
Interesting. I do feel letting a machine generating notes is taking the joy out of improvisation, is it not?
I think one of the great, early, joys of learning a piano is gaining the following intuitions: The seemingly harder path of learning sheet is actually faster. Your mind _should_ learn to think in two dimensions Spatial — where fingers go — and Time — pitch and tempo – ** when learning. The _internalization_ of Space and Time queues guide the fingers in a dance that is vastly satisfying. This skill leads you to the final part of the journey that is improvisation and the one more exciting than what i am on now.
---
**
Space: Your finger placement on keys right, e.g. knowing how to go from landmark/anchor notes(mid-c, G, F etc) and then go to the others above and below it. Crudely this is some what like typing from your landmark f and j qwerty keyboard
Time: The out singing/verbalizing of the notes/beats on a time measure as you play them(per the time measure). e.g. you can say out loud 1-2-3-4 for 4/4 measure, if the measure has quarter notes say out loud. And `1-e-and-a-2-e-and-a-3-e-and-a-4-e-and-a` for a 4/4 with 1/16th note granularity. Do this as you play the notes and you get a sense of tempo.
Not the right word; thought it was broken when I played the song samples. Of course, your model is much more interesting than what would be implemented with a prefix tree.
I would love something like that, except that I play the melody, and it produces proper 3-4 part accompaniment, preferably in good baroque style. Extra bonus if it could also write it into a file in a format suitable for music editing programs.
That is an incredibly hard challenge though. Creating the whole backing track (in any meaningful way other than just basic chords) from just melody will require an amazingly high number of highly subjective choices and random gen will not lead to good outcomes since our ears like intentional and artistical creativity in general.
An early attempt at this was Microsoft Songsmith [1] all the way back in 2009, which would take a melody (usually recorded by mic) and try to scaffold an accompaniment around it though obviously not realtime in any sense of the word.
The closest we've had to realtime orchestration around a melody in the "real world" is probably arranger keyboards though your left hand is still responsible for the chord progression itself.
This feels like a natural next step. Starting with a simple melody and having the system fill in the rest while still following your playing style could make it much more useful for experimentation than generating a complete piece from scratch.
This is honestly astonishing and the first "AI music" I've heard that has the potential to sound beautiful. I always thought that MIDI would be a perfect format for this. Glad to see this person make it happen!
Ah, MIDI files. The only type of music you could realistically download from the internet back in the day, and you had to wake up at ungodly hours so that your dialup modem would not rack up a massive phone bill.
I remember at one point RuneScape switched from the built in Microsoft midi whatever to their own sound engine, and from that day, everything sounded wrong, even the frogs, because to me, the crappy midi sounds were the whole personality and feel of the game.
This is really fun. Scaler 3 starts with a chord progression and lets you break it down into musical performances and parts. Useful for ideation when producing.
Would be fun to get a midi clock going and play some chords on my piano and have my synth start jamming along with the bass and my keyboard doing some performance. Or any combination of the above.
That would be a really interesting direction. At that point it starts feeling less like autocomplete and more like having another musician reacting to what you're playing in real time.
Oh, a cool idea! I just tried it, works pretty well. Kudos!
One feature request:
Instead of playing the AI-generated audio solely through the iPhone's speakers, add an option to send the audio as midi notes to a device (probably the same one you received the mini notes from).
How would you expand this to support elements like attack ("velocity of the key-down" in piano speak), grace notes, timing etc. Would each of those be part of this model or another model? How would you model an arbitrary element (pedal, duration, etc...)
The idea is awesome! :) However there's definitely much room for improvement, first of all rythm and composition (so there's some sense of musical form).
Yes, I think I’ve gotten it to roughly a GPT-2 level: good enough to share, but with a lot of room left to improve. I think adding some kind of bar/measure token might help with rhythm, and perhaps some form of longer-term planning for the overall composition.
The biggest speed improvement came from changing the note representation when I switched to compound note events: roughly 5× fewer autoregressive passes per note.
For the current model I’m using Core ML, which optimizes the kernels the first time you run it. I haven’t actually spent that much time tuning performance beyond that.
The answer about changing the note representation was interesting. Sometimes a change in how the problem is represented ends up giving a much bigger improvement than trying to optimize the model itself.
This is so amazing, can you improve the quality of generation at the cost of notes per seconds ? No one can play 108 notes/sec anyways, maybe you can train the model to do CoT for better quality
Yes, some kind of planning step is on my TODO list. Another thing I want to try is generating a few continuations in parallel, picking the one that looks best, and then continuing from there. Maybe the picking could be automatic.
I can probably squeeze out quite a bit more than 100 notes/sec as well. I haven’t spent much time optimizing inference yet.
Luckily I have access to 4x RTX 4090s, so I didn’t have to pay cloud GPU prices directly. If I had, it probably would have added up quite a bit given how many training runs and experiments I ended up doing.
Is this HN? Aren't people supposed to tinker with tech for no reason other than seeing if they can? Is everything that involves a transformer now just AI BAD? Is that what the world has devolved into? Each side screaming "Orange Man Bad" and endless variations at each other?
I’m plenty calm. I was just stating an objection. I don’t know where you’re getting the idea that I am not. I’m actually having a very nice day overall.
So you’re throwing out intentionally charged examples and telling people to calm down, the one thing we all pretty much know is guaranteed to engender the exact opposite reaction? And you believe this is helpful?
While it is a neat parlor trick, a lot of people have specific grevience against the application to art. AI has only served to further disempower artists broadly, and arguably it pushes "art" to a lower common denominator. Try to actually situate yourself in "why" people get rankled instead of making it a thought terminating cliche.
Given the AI crowd is very loudly telling us since years how humans will be replaced by LLMs and we will all be poor and left behind if we don’t join their cult, I think it’s a perfectly reasonable reaction. Anything with AI mixed with Art or other human experience is suspicious.
Also, yes, the orange fascist who attempted to coup his way to power, raped women, and is destroying democratic institutions is indeed bad.
No they haven’t. It’s like two guys that nobody really believes. You specifically seek that information out to get mad, then proceed to see it where it’s not. Like when someone makes a cool personal project about their hobby and it happens to involve AI.
> Given the AI crowd is very loudly telling us since years how humans will be replaced by LLMs and we will all be poor
Oh no, AI will take our jobs??? FUCKING LET IT!
Who even wants to do all these jobs if we don't HAVE to??
Ask politicians to give us UBI.
You want to attack the shit that could make shit easier instead of attacking the 200 year old institutions in place that ensure class divisions and perpetual debt and wage slavery? Smart buggers
> You want to attack the shit that makes shit easier instead of attacking the 200 year old institutions in place that ensure class divisions and perpetual debt and wage slavery
I have bad news for you if you think that will improve with the current deployment of AI. You will also notice I didn’t mention anything about jobs
One think I didn't see mentioned in the post- maybe I missed it- how large was the data? How many samples did you use to pretrain and post-train
For anyone interested, I’d highly recommend reading Robert Gjerdingen’s article Gebrauchs-Formulas. https://www.researchgate.net/publication/259731561_Gebrauchs...
You can also listen to the transcript of four Russian composers, including Rachmaninoff, playing this pattern recognition and generation game at a dinner party in the late 1800’s: https://youtu.be/PlFPOWuwBHI?is=EKBK7QQkJs4MsTCU
Composers at the time could do this just by looking at sheet music and audiating, without using a piano.
I think one of the great, early, joys of learning a piano is gaining the following intuitions: The seemingly harder path of learning sheet is actually faster. Your mind _should_ learn to think in two dimensions Spatial — where fingers go — and Time — pitch and tempo – ** when learning. The _internalization_ of Space and Time queues guide the fingers in a dance that is vastly satisfying. This skill leads you to the final part of the journey that is improvisation and the one more exciting than what i am on now.
---
**
Space: Your finger placement on keys right, e.g. knowing how to go from landmark/anchor notes(mid-c, G, F etc) and then go to the others above and below it. Crudely this is some what like typing from your landmark f and j qwerty keyboard
Time: The out singing/verbalizing of the notes/beats on a time measure as you play them(per the time measure). e.g. you can say out loud 1-2-3-4 for 4/4 measure, if the measure has quarter notes say out loud. And `1-e-and-a-2-e-and-a-3-e-and-a-4-e-and-a` for a 4/4 with 1/16th note granularity. Do this as you play the notes and you get a sense of tempo.
https://www.francoispachet.fr/continuator/
Not the right word; thought it was broken when I played the song samples. Of course, your model is much more interesting than what would be implemented with a prefix tree.
The closest we've had to realtime orchestration around a melody in the "real world" is probably arranger keyboards though your left hand is still responsible for the chord progression itself.
[1] - https://en.wikipedia.org/wiki/Microsoft_Research_Songsmith
Would be fun to get a midi clock going and play some chords on my piano and have my synth start jamming along with the bass and my keyboard doing some performance. Or any combination of the above.
One feature request:
Instead of playing the AI-generated audio solely through the iPhone's speakers, add an option to send the audio as midi notes to a device (probably the same one you received the mini notes from).
Yes, I think I’ve gotten it to roughly a GPT-2 level: good enough to share, but with a lot of room left to improve. I think adding some kind of bar/measure token might help with rhythm, and perhaps some form of longer-term planning for the overall composition.
For the current model I’m using Core ML, which optimizes the kernels the first time you run it. I haven’t actually spent that much time tuning performance beyond that.
I can probably squeeze out quite a bit more than 100 notes/sec as well. I haven’t spent much time optimizing inference yet.
https://magenta.withgoogle.com/magenta-realtime-2
Pretraining was obviously a a lot slower, the 125M model took roughly half a day.
But, but… wouldn't that be… (gasp) DISTILLATION?
Fun project!
You know, I was nodding along until you shoehorned that one in.
So you’re throwing out intentionally charged examples and telling people to calm down, the one thing we all pretty much know is guaranteed to engender the exact opposite reaction? And you believe this is helpful?
Also, yes, the orange fascist who attempted to coup his way to power, raped women, and is destroying democratic institutions is indeed bad.
> No they haven’t. It’s like two guys that nobody really believes.
It’s the entire leadership of the AI industry, in case you haven’t noticed
Oh no, AI will take our jobs??? FUCKING LET IT!
Who even wants to do all these jobs if we don't HAVE to??
Ask politicians to give us UBI.
You want to attack the shit that could make shit easier instead of attacking the 200 year old institutions in place that ensure class divisions and perpetual debt and wage slavery? Smart buggers
I have bad news for you if you think that will improve with the current deployment of AI. You will also notice I didn’t mention anything about jobs
thing go plink
machine hear plink
machine make many more plink
man happy for plink is fun
man not hit machine with club or scream on orange site
man leave cave and touch plant
Or maybe there is some other "spark" deeper within?
When everyone can make anything as soon as they think of it, what will set us apart?