Live data from Hacker News

Sora: Creating video from text

openai.com

141–150 of 1001 posts

Re: Sora: Creating video from text

#141
People here seem mostly impressed by the high resolution of these examples.

Based on my experience doing research on Stable Diffusion, scaling up the resolution is the conceptually easy part that only requires larger models and more high-resolution training data.

The hard part is semantic alignment with the prompt. Attempts to scale Stable Diffusion, like SDXL, have resulted only in marginally better prompt understanding (likely due to the continued reliance on CLIP prompt embeddings).

So, the key question here is how well Sora does prompt alignment.

Re: Sora: Creating video from text

#143

OpenAI demonstrating the size of their moat. How many multi-million-dollar funded startups did this just absolutely obsolete? This is so, so, so much better than every other generative video AI we've seen. Most of those were basically a still image with a very slowly moving background. This is not that. Sam is probably going to get his $7T if he keeps this up, and when he does everybody else will be locked out foreve…

seems like a significant chunk of the population may opt in to the Matrix voluntarily.

on another note I find it funny they released this right after Google announced their new model. Bad luck for Google or did OpenAI just decide to move up their announcement date to steal their thunder?

Re: Sora: Creating video from text

#144
Looked at the first clip and immediately noticed the woman's feet swap at ~15 seconds in. My eyes were drawn to the feet because of the extreme supination in her steps.

Looks like a dramatic improvement in video generation but still a miss in terms of realism unless one can apply pose control to the generated videos.

Re: Sora: Creating video from text

#145
"so far ahead" "leaps and bounds beyond anything out there" "This is insane"

Let's temper the emotions for a second. Sora is great, but it's not ready for prime time. Many people are working on this problem that haven't shared their results yet. The speed of refinement is what's more interesting to me.

Re: Sora: Creating video from text

#146

Did anyone else feel motion sickness or nausea watching some of these videos? In some of the videos with some panning or rotating motion, i felt some nausea like sickness effect. I guess its because some details were changing while in motion and I was unable to keep track or focus anything in particular. Effect was stronger in some videos.

I do. My hypothesis is that there isn't really good bokeh yet in the videos, and our brains get motion sick trying to decide what to focus on. I.e. too much movement and *too much detail* spread out throughout the frame. Add motion to that and you have a recipe for nausea (at least for now)

Re: Sora: Creating video from text

#147
It's interesting how a lot of the higher frequency detail is obviously quantized. The motion of humans in the drone shots for example is very 'low frequency' or 'low framerate', and things like flowing ocean water also appears to be quantized. I assume this is because of the internal precision of these models not being very high?

Re: Sora: Creating video from text

#148

OpenAI demonstrating the size of their moat. How many multi-million-dollar funded startups did this just absolutely obsolete? This is so, so, so much better than every other generative video AI we've seen. Most of those were basically a still image with a very slowly moving background. This is not that. Sam is probably going to get his $7T if he keeps this up, and when he does everybody else will be locked out foreve…

seems like a significant chunk of the population may opt in to the Matrix voluntarily. on another note I find it funny they released this right after Google announced their new model. Bad luck for Google or did OpenAI just decide to move up their announcement date to steal their thunder?

Only those that can afford it. The rest will be forced to live in the real world, like 20th century peasants.

Re: Sora: Creating video from text

#149
Imagine a movie script, but with more detail of the scenes and actors, plugged into this.

The killer app for this is being able to give a prompt of a detailed description of a scene, with actor movements and all detail of environment, structure, furniture, etc. Add to that camera views/angles/movement specified in the prompt along with text for actors.

Re: Sora: Creating video from text

#150
post #49

https://openai.com/sora?video=big-sur In this video, there's extremely consistent geometry as the camera moves, but the texture of the trees/shrubs on the top of the cliff on the left seems to remain very flat, reminiscent of low-poly geometry in games. I wonder if this is an artifact of the way videos are generated. Is the model separating scene geometry from camera? Maybe some sort of video-NeRF or Gaussian Splatti…

The model is essentially doing nothing but dreaming.

I suspect that anything that looks like familiar 3D-rendering limitations is probably a result of the training dataset simply containing a lot of actual 3D-rendered content.

We can't tell a model to dream everything except extra fingers, false perspective, and 3D-rendering compromises.

Post reply on HN