Live data from Hacker News

MusicLM: Generating music from text

arxiv.org

61–70 of 114 posts

Re: MusicLM: Generating music from text

#61
I wonder if this technology will eventually revolutionize music the same way synthesizers did. Or at least lead to music and effects/filters that are simply not possible with current DAWs and plugins.

Custom generation of samples from text alone seems revolutionary.

Re: MusicLM: Generating music from text

#62
I'm absolutely flabbergasted by rapid progression of music related machine learning models. I've always seen music as the "final frontier". With images you can have tiny errors and noise that isn't that noticeable by the human eye, but with music everything has to be impeccable and if a note is slightly off, you instantly hear it. Machine learning models that will be able to create outstanding music, imo, will mark the end-game of creative AI.

Anyone want to share their experiments with MusicLM, feel free to join community-fan discord and subreddit:

https://reddit.com/r/MusicLM/

https://discord.gg/pjVcsyfCJR

BTW, I really hope this gets integrated in software like Ableton similarly how image generators get slowly integrated into Adobe Creative Suite.

Re: MusicLM: Generating music from text

#64

I want to do the reverse: pass in some music and have it describe what I'm hearing

I believe parts of this model could be adapted to accomplish this. The "mulan" network generates embeddings from music and text so that they "match" each other, i.e. audio segments and text captions that go well together would be mapped to the same embedding. At run-time, they use embeddings from the text to generate the music. So conceivably the opposite could be done, using the "mulan" embeddings from audio only to generate text, by replacing the "soundstream" decoder with a text LLM. That said, there's probably simpler methods that could generate decent results, and I'd be surprised if there isn't already some out there. It's worth noting that along with this work they released a dataset of music with captions that could also help training new models for this task.

Re: MusicLM: Generating music from text

#65

Earlier quoted context omitted.

I understand what you're saying, although it could be argued that at least for some types of image tasks one would prefer something like an SVG output with layers to make it easier to edit. For music, I think it's partly an academic question of "can we do it" rather than trying to maximize immediate practical usefulness. There's already quite a bit of work on symbolic music generation (mostly MIDI), a lot of it quite…

I would offer a counter point: Chat GPT uses text, which is analogous to the MIDI of spoken language. Now, what's to prevent a speech to text and text to speech AI slapped on both ends to make it audibly conversational? Or train the franken-model model end to end in one monolithic model in hopes that it might make other interesting connections. That being said I have seen very little in the AI transcription space (Li…

Oh, there's already all sorts of work mashing together GPT text with text-to-speech and speech-to-text. There were some viral videos featuring "Steve Jobs" talking about AI and one being interviewed by "Joe Rogan". I also saw somewhere here on HN a website featuring an "infinite" conversation between "Zizek" and "Herzog", in a similar vein.

Though something end-to-end would probably make more sense; text is a lossy way to encode speech, as it doesn't capture a lot of nuance that goes into spoken words. Interestingly, the same could be said of sheet music and even MIDI in the music realm.

As for transcription for multi-instrument music, there's been a lot of work on that front, and a lot of progress that's already finding its way into commercial products. The latest I've seen in this domain is https://samplab.com/, which has some compelling videos of not only transcribing multi-instrument audio but editing it on a note-by-note basis.

Re: MusicLM: Generating music from text

#66
post #44

Earlier quoted context omitted.

I hope you also consider getting some music from an actual producer!

Why is there this need for humans to be involved at all stages of production? Do we still need humans to knit and handwash our clothes? To live in a caboose or lighthouse? To operate our elevators? To deliver the milk? Why should anyone have to learn to draw in the future when machines promise to do a better job? There are so many better uses for our time, and too many things to do for our short lives to handle. I wa…

Just keep in mind, that movie will come out the same time when possibly 6 billion others will release theirs.

Re: MusicLM: Generating music from text

#67

I'm absolutely flabbergasted by rapid progression of music related machine learning models. I've always seen music as the "final frontier". With images you can have tiny errors and noise that isn't that noticeable by the human eye, but with music everything has to be impeccable and if a note is slightly off, you instantly hear it. Machine learning models that will be able to create outstanding music, imo, will mark t…

> if a note is slightly off, you instantly hear it

I don't know about this. I know professional musicians and they say that they make little mistakes all the time, but part of being a professional is being able to pretend those mistakes didn't exist because the audience doesn't know what it's supposed to sound like.

Re: MusicLM: Generating music from text

#68

Demo: https://google-research.github.io/seanet/musiclm/examples/

Holy hell those are insanely good. Best generated music I've heard so far.

I was born too soon to explore the stars, born too late to be a 17th century pirate... but born just in time to explore these incredible neural networks come to life.

Post reply on HN