Live data from Hacker News

Make-A-Video: AI system that generates videos from text

makeavideo.studio

31–40 of 399 posts

Re: Make-A-Video: AI system that generates videos from text

#31

Earlier quoted context omitted.

> I don't think it will improve much from here That's a bold prediction. Why do you think that? The first thing I thought was the exact opposite. This isn't very good, but it's only version 1. Motion pictures are less than 150 years old. In another 150 years I bet virtual filmmaking will progress a lot.

I agree, it feels like this technology is progressing by leaps and bounds almost by the day.

Case in point, just compare the results of Stable Diffusion / Dall-E / midjourney to papers from 5-10 years ago (here's a random one I found http://proceedings.mlr.press/v48/reed16.pdf). Remarkable progress in that time, and as I understand it, a good amount (but certainly not all) of the improvement comes “free” from being able to scale up to more weights.

Re: Make-A-Video: AI system that generates videos from text

#32

A lot of people saying it's over for traditional movie-making - lmao. I look at these and see nothing but uncanny valley artifacts, and I don't think it will improve much from here. It's like self-driving cars. They use almost very effective statistical models, certainly better than our previous models, but they never seem to shake off that "almost" and become truly effective.

It’s been like a month since people were saying this tech would be too hard to apply to video.

The leap from "generating a frame" to "generating a video" is not as big as the leap from where we are now to a 'perfect' image synthesis engine(which would be required to get rid of artifacts).

The much more likely scenario IMO is that people get used to the artifacts and notice them less.

Re: Make-A-Video: AI system that generates videos from text

#33

A lot of people saying it's over for traditional movie-making - lmao. I look at these and see nothing but uncanny valley artifacts, and I don't think it will improve much from here. It's like self-driving cars. They use almost very effective statistical models, certainly better than our previous models, but they never seem to shake off that "almost" and become truly effective.

> I look at these and see nothing but uncanny valley artifacts

Ok? It's a nascent technology. Look at the original DALL-E blog post from last year [1]. Now compare it to DALL-E 2 and Stable Diffusion.

[1]: https://openai.com/blog/dall-e/

Re: Make-A-Video: AI system that generates videos from text

#35
post #18

I don’t want to think reality is a simulation, but wouldn’t our everyday experience being generated by a neural network be far far simpler to achieve than a universe of infinite minuscule atoms all interacting.. like ozcam’s razor is pointing towards our lives being a realistic dream. Like in 10 years you could plug this tech into high end VR and get a prompted reality dynamically generated that would be indistinguis…

That what we're seeing is really reality is the simpler explanation. Because the alternative is that what we see is a simulation ... inside another reality that has full complexity. Which increases the overall complexity. Thus, Occam's Razor says what we see is likely real and not a simulation.

Check out the original simulation argument paper[0]. The issue is, if we think we are heading to a world in which we can do simulations, it becomes increasingly likely that we are in one of those (presumably very many) worlds, rather than in the one world that existed before the advent of such simulations.

[0] https://www.simulation-argument.com/simulation.pdf

Re: Make-A-Video: AI system that generates videos from text

#37

Earlier quoted context omitted.

The neural net itself needs something to run in (some universe). If the argument is that it needs a smaller or simpler universe than ours: first of all it is simulating our universe so all complexity of ours is included anyway, but also, maybe a neural net in a big chaotic universe is more likely than a tiny universe designed to perfectly fit one neural net

Neural nets as an abstract concept don’t necessarily need to be built with atoms, just the substance of the external universe which may be far simpler.

I also didn't imply this neural network's universe needs atoms, but it needs to run the computation somehow, the computation itself is the complexity, no matter with what physics or other manifestation it shows up

And if it is simulating our universe's atoms, then that part of it basically is our atoms. But is a neural net doing that really simpler than the atoms just running with a more direct mathematical model?

Re: Make-A-Video: AI system that generates videos from text

#38

A lot of people saying it's over for traditional movie-making - lmao. I look at these and see nothing but uncanny valley artifacts, and I don't think it will improve much from here. It's like self-driving cars. They use almost very effective statistical models, certainly better than our previous models, but they never seem to shake off that "almost" and become truly effective.

CogVideo was released only 4 months ago, you can see some samples on this page https://models.aminer.cn/cogvideo/. Statistical models can be used more efficiently (see KerasCV, https://news.ycombinator.com/item?id=33005585), and computing power can increase.

I don't think AI will be able to create a movie anytime soon, but I think it will become "good enough" to serve as inspiration for creatives, or to replace simple stock footage (Much like SD and DALLE-2 is now).

Re: Make-A-Video: AI system that generates videos from text

#39
post #18

Earlier quoted context omitted.

That what we're seeing is really reality is the simpler explanation. Because the alternative is that what we see is a simulation ... inside another reality that has full complexity. Which increases the overall complexity. Thus, Occam's Razor says what we see is likely real and not a simulation.

The ‘outside’ reality only needs to be the size of a data center, not the entire universe. And maybe the outside reality only needs to be a trained 100gb model. A much simpler outside reality than we have now. You don’t even need to sim 8 billion people, just 1, you.

Occans razor in this case applies to the complexity giving rise to a situation here though, not the complexity of a described system.

Conventional physics giving birth to our universe is currently the model with the fewest assumptions.

What would have to take place to give rise to a universe the size of a data center, running an AI model of a human? It feels like we have to bake in assumptions of stable physics, a rise of a stable system for that data center, and some path towards creating it and modelling a human.

That said, if we believe we're capable of running billions of believable simulations, then we're more likely to be in such a simulation than ground reality. But a datacenter pocket dimension bakes in a lot of assumptions that make it less likely than our own universe.

Post reply on HN