Live data from Hacker News

Artificial Intelligence Generates Christmas Song from Holiday Image

news.developer.nvidia.com

51–60 of 115 posts

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#51
post #34

Auto-synthesis of music has been a topic of academic interest since the 1950s, when the first mainframe scribbled out code on paper tape to be translated into sheet music and performed. UToronto's work here is the latest expression of this desire. The huge gap between our cultures' actual music and these synthetic projects can to an extent be described through "receptivity" or the phenomenology of music, in other wor…

That video is fantastic I watched it a week or so ago. Your explanation also explains why computer performed music is so off. It still has that uncanny valley effect. So when Sony had a computer generate a "Beatles-esque pop song", they still had a human perform and produce it. But at the point there's so much creativity and human-added value on top of it that I don't think its fair to call it computer generated imho…

yes. I can tell you a little more about that, too, since I used to research this stuff and think about it a lot still.

One of my models of music is an external model of a regulated system that parallels and trains our own habits and responses. E.g. a song demonstrates tension and release similar to our own lives. The level of tension in a song before release occurs can inform us how much tension which should accept before performing some release activity.

Music's rhythms also inform the pace of our work. E.g. verse-chorus-verse represents switching between two different activities. Even the pitch of a single note acts as a reference for the amount of intensity of a sensation we should use in our own lives. E.g. thrash metal listeners enjoy sudden shifts into massive intensity and hold it there. Dub step listeners are training themselves for unusual, but rather intense aesthetics leading up to disproportionate release. Classical music tends to be for "long-chain thinkers" tumbling ideas over from various perspectives, e.g. writers and politicans, doctors, not factory workers.

With that as a background, consider that a live instrument is also a physical system with a human controlling it interactively. The live system is a bit different every time. Here's the critical part: the human must listen and provide instantaneous feedback to a varying system in order to present the piece of music as a proper response model of a regulated system. If the player fails to do this, the model communicated by the performance is different.

In open-loop systems, such as a sequencer, there is no (or limited) interaction between the player and the sound, so an incidental model emerges. That incidental model represents an unintended and therefore most likely irrelevant model of how to interact with reality. e.g. it relieves tension where no relief was needed. It lingers too long on an idea, long after a human novelty-seeking circuit has starved.

Some people, e.g. in discussions of unstable filters like the TB-303, chalk up the variations as being different at every performance because the instrument is random... However, they're missing the closed loop portion of the performance, in which the performer reacts to the unpredictability of the instrument in order to maintain the model. In other words, the score and notes are not the music, but the performer's response to the environment the score sets up is the music.

To revivify your uncanny valley observation, the "unstable filter creates variations" crowd has a parallel in Perlin noise used to subtly animate human models to make them not look so dead. However, it's incomplete because they don't use (short-term) feedback to determine when the movement suffices to be convincing. That feedback is the essence of performance.

In theory, computer scientists could implement these feedback models in performance to make the sounds more realistic. They could be used in synthesis, but the playback would still require observation of the listener! Which is possible. Personally, I just prefer playing electronic instruments live over using sequencers. It's only the sounds of electronic music I like, the zaps, peowms, zizzes, pews, and poonshes, etc. I don't care for electronics/computers to perform for me.

If you like this hypothesis, you can find more references on my wiki at: http://www.diydsp.com/index.php?title=Computer_Music_Isolati...

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#52

Earlier quoted context omitted.

>hide... inside of your compiler. Not to mention the rest of the ecosystem. Look what happens to humans when they grow up without being properly embedded in the family, with its hundreds of thousands of years of historical contingency: https://en.wikipedia.org/wiki/Genie_(feral_child) We can't even begin to describe the amount of information encoded there.

uh, you can begin to describe the amount of information. if someone grew up with sensory input limited to HD video meaning something like 6-10 GB/hour they would be handicapped versus other humans but not drastically so. In 6 years there are 52,560 hours. A large library, sure, but we already have all this digitized anyway. I'm cutting a lot of corners, and the human substrate in which our brains are embedded is fini…

>50,000 blue ray discs' worth oughta do it.

By this logic, the handful of megabytes of unicode making up War and Peace in the original russian should be enough for a non-russian speaker to fully grasp it and all its meaning and implications. It isn't even enough for a native russian speaker to do so.

Humans aren't raised by simply looking at their surroundings, they're raised by interacting with people, who were in turn raised by interacting with people, going back for the whole history of humanity, or arguably mammals. That information isn't all in the genome, although natural selection has put some of it in there. The bit that we don't know how to describe information-theoretically is the bit that isn't in the genome, because we don't know how it's encoded.

I don't see how you'd even try to put bounds on it: this is essentially the problem posed by post-modernism/post-structuralism/literary theory, once you strip away the marxism, and science's response has understandably been to reject it but it can't do so forever if it wants to create AGI.

Or maybe I've misunderstood your point.

I might be pursuaded that a 2 year old could come sooner than we think, via brute computational force as you describe, but I'd argue that a 2 year old with the capacity to become anything more than a 2 year old is much farther away than we think.

EDIT: if, as you claim, we are close to having the computational power to simulate human children, then why aren't we already successfully simulating much simpler animals? IIRC the best we can do is a tiny chunk of a rat, or the whole of various kinds of microscopic worms, and those are just computational models, not turing test-passing replicants.

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#53
post #30

Show me the same algorithm generate songs in a different genre from different images and I'll be impressed.

I'm from the project team. This is a very interesting point. While it is easy to crawl many songs from the internet, it is a little harder to gather the same amount but with proper genre/style/etc labels, although it is not impossible. For now there's only one genre, which we call it "the genre of whatever is on the internet". So whatever music files on there, many of them quite "crappy", were used to train the model…

I mean, where does the Christmas element comes from? The image alone, the music it was trained with, or is it somehow hardcoded in the algorithm?

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#54
post #37

Earlier quoted context omitted.

I think that your "700MB uncompressed" fails to take into account that the construction, development, and maturation of the brain relies heavily on cellular and molecular mechanisms. I think it is a little disingenuous to hide the enormous wealth of information necessary to create a brain, much less understand and utilize one, inside of your compiler.

I don't think you are correct from an information computation point of view. When looking at the computation done by neurons in the brain it is sufficient to abstract away the lower-level substrate in which it occurs. You and other posters are all correct regarding the huge volume of information on which human minds are trained. it's hardly unsupervised learning either :) --------- EDIT: In response to your comment,…

I will confess to not being an expert, but I disagree: I don't think it's sufficient to abstract away the lower-level substrate when the OP was referring to DNA as source code, which absolutely depends on that level of detail to both construct the system (the brain) and to enable the continued development and maturation of that system (a physical, real brain).

I was not referring to the huge volume of information necessary, as I acknowledge that as being "outside the system" for purposes of this discussion, so my apologies for any confusion I might have caused.

It may be possible that (and it is my belief that) there is a higher-level abstraction for the computations taking place in the brain, even if it is on the neuron-level, but at that point I don't think you can claim that the source code for that is going to fit under 700MB by using DNA as a baseline.

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#55
post #30

Earlier quoted context omitted.

I'm from the project team. This is a very interesting point. While it is easy to crawl many songs from the internet, it is a little harder to gather the same amount but with proper genre/style/etc labels, although it is not impossible. For now there's only one genre, which we call it "the genre of whatever is on the internet". So whatever music files on there, many of them quite "crappy", were used to train the model…

I mean, where does the Christmas element comes from? The image alone, the music it was trained with, or is it somehow hardcoded in the algorithm?

The Christmas element comes from 1. the image, and 2. a 4800-dimensional RNN sentence encoding bias generated from ~30 Christmas songs.

Not sure how to hardcode this.

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#56
post #40

Earlier quoted context omitted.

> This is like heavier-than-air flight. Yes, this is like heavier-than-air flight. It would be like if we didn't really understand how flight works so we are just imitating birds by gluing feathers to cardboard wings, flapping them, and hoping flight just sort of happens. In a similar way, we don't really understand how the human brain works, so we are just imitating neurons and hoping sentience will just sort of hap…

How long did it take to go from really silly and obviously impractical flying machine designs to working, practical ones?

Thousands of years? Who knows how long before the Wright Brothers people have been attempting flight?

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#57
post #40

Earlier quoted context omitted.

> This is like heavier-than-air flight. Yes, this is like heavier-than-air flight. It would be like if we didn't really understand how flight works so we are just imitating birds by gluing feathers to cardboard wings, flapping them, and hoping flight just sort of happens. In a similar way, we don't really understand how the human brain works, so we are just imitating neurons and hoping sentience will just sort of hap…

How long did it take to go from really silly and obviously impractical flying machine designs to working, practical ones?

[deleted]

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#58

Earlier quoted context omitted.

Instead of five, don't you mean "two to three"? And under that comparison, isn't it scary how much it does sound that way? Like a kid who hears words together but don't know what they mean yet? (and doesn't really get the cultural things around it.) To me it sounds like a quite musical 2-3 year old stringing words together. Doesn't it strike anyone else that way? These things are going to grow up very, very soon! We…

> uses some 20 watts Daniel Lemire has an observation here: our brains don't use all the connections, all the time (when it does it's called epilepsy); it would be interesting if we built some computers that also worked like that - with only small parts being active at a certain time, but switching quickly from one to another. That would help with the heating problem and it might allow processors a lot more complex t…

if you're trying to achieve parity with human brain calculation the most important thing is how slow neural connections are. they're snail-paced compared with gigabit links and CPU's @ 3 Ghz+. There's quite a lot of memory that is needed but if you're allowed to make a million round-trips across the room before you miss your realtime deadline, it's kind of silly to think that it won't happen. The fact you point out (about not all parts firing at once) could be useful because it might result in these links not being saturated by data anyway.

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#59
post #15

It kind of feels like a five-year-old trying to make a song. Which seems good - now they only need to improve the mechanisms and who know, maybe we'll get to the point when it's a ten-year-old?

Instead of five, don't you mean "two to three"? And under that comparison, isn't it scary how much it does sound that way? Like a kid who hears words together but don't know what they mean yet? (and doesn't really get the cultural things around it.) To me it sounds like a quite musical 2-3 year old stringing words together. Doesn't it strike anyone else that way? These things are going to grow up very, very soon! We…

By your own argument, any program that a single human being can write can also contain only 700MB of source code, including graphics and videos etc, which is obviously not true.

Re: Artificial Intelligence Generates Christmas Song from Holiday Image

#60
post #30

Show me the same algorithm generate songs in a different genre from different images and I'll be impressed.

I'm from the project team. This is a very interesting point. While it is easy to crawl many songs from the internet, it is a little harder to gather the same amount but with proper genre/style/etc labels, although it is not impossible. For now there's only one genre, which we call it "the genre of whatever is on the internet". So whatever music files on there, many of them quite "crappy", were used to train the model…

Great job! What sort of resources did you use (Time, Processing power etc) to train this?

Edit: Found it in the article.

Post reply on HN