Live data from Hacker News

Sora: Creating video from text

openai.com

111–120 of 1001 posts

Re: Sora: Creating video from text

#111
post #49

https://openai.com/sora?video=big-sur In this video, there's extremely consistent geometry as the camera moves, but the texture of the trees/shrubs on the top of the cliff on the left seems to remain very flat, reminiscent of low-poly geometry in games. I wonder if this is an artifact of the way videos are generated. Is the model separating scene geometry from camera? Maybe some sort of video-NeRF or Gaussian Splatti…

Curious about what current SotA is on physics-infusing generation. Anyone have paper links?

OpenAi has a few details:

>> The current model has weaknesses. It may struggle with accurately simulating the physics of a complex scene, and may not understand specific instances of cause and effect. For example, a person might take a bite out of a cookie, but afterward, the cookie may not have a bite mark.

>> Similar to GPT models, Sora uses a transformer architecture, unlocking superior scaling performance.

>> We represent videos and images as collections of smaller units of data called patches, each of which is akin to a token in GPT. By unifying how we represent data, we can train diffusion transformers on a wider range of visual data than was possible before, spanning different durations, resolutions and aspect ratios.

>> Sora builds on past research in DALL·E and GPT models. It uses the recaptioning technique from DALL·E 3, which involves generating highly descriptive captions for the visual training data. As a result, the model is able to follow the user’s text instructions in the generated video more faithfully.

The implied facts that it understands physics of simple scenes and any instances of cause and effect are impressive!

Although I assume that's been SotA-possible for awhile, and I just hadn't heard?

Re: Sora: Creating video from text

#112

This is insane. Even though there are open-source models, I think this is too dangerous to release to the public. If someone would've uploaded that Tokyo video to youtube, and told me it was a drone.. I would've believed them. All "proof" we have can be contested or fabricated.

"Proof" for thousands of years was whatever was written down, and that was even easier to forge.

There was a brief time (maybe 100 years at the most) where photos and videos were practically proof of something happening; that is coming to an end now, but that's just a regression to the mean, not new territory.

Re: Sora: Creating video from text

#113

I do wonder why OpenAI chose the name "Sora" for this model. AI is now going to have intersectionality with Kingdom Hearts. (Atleast you don't need a PhD to understand AI.)

Sora means sky in Japanese, their reasoning is akin to "the sky's the limit".

> The team behind the technology, including the researchers Tim Brooks and Bill Peebles, chose the name because it “evokes the idea of limitless creative potential.”

Re: Sora: Creating video from text

#114
Those samples are incredibly impressive. It blows RunwayML out of the water.

As a layman watching the space, I didn't expect this level of quality for two or three more years. Pretty blown away, the puppies in the snow were really impressive.

Re: Sora: Creating video from text

#115

OpenAI demonstrating the size of their moat. How many multi-million-dollar funded startups did this just absolutely obsolete? This is so, so, so much better than every other generative video AI we've seen. Most of those were basically a still image with a very slowly moving background. This is not that. Sam is probably going to get his $7T if he keeps this up, and when he does everybody else will be locked out foreve…

That's an interesting take - podcasts have become a replacement for companionship and conversation.

Re: Sora: Creating video from text

#116
post #41

Countdown to when studios licensing this for "unlimited" episodes of your favorite series. There was Seinfeld "Nothing, Forever" AI parody, but once the models improve enough and are cheap enough to deploy, studios will license their content for real and just have endless seasons. Or even custom episodes. Imagine if every episode of a TV show was unique to the viewer.

I imagine it's not long before we see hyper-targeted commercials where the actors look like us, live in our city, etc.

Nothing stopped doing so before AI - just slam a photo of your friends to the ad.

Re: Sora: Creating video from text

#117

Countdown to when studios licensing this for "unlimited" episodes of your favorite series. There was Seinfeld "Nothing, Forever" AI parody, but once the models improve enough and are cheap enough to deploy, studios will license their content for real and just have endless seasons. Or even custom episodes. Imagine if every episode of a TV show was unique to the viewer.

One understated aspect of AI Seinfeld is that it took many steps to differentiate it from the actual Seinfeld and create its own identity, such as the 144p visual filter and the random microwave. Those tweaks added to its charm.

If someone tried to do AI Seinfeld again in 2024, many would criticze it for not being realistic enough now that the tools to do so are now available.

Re: Sora: Creating video from text

#118

This is insane. Even though there are open-source models, I think this is too dangerous to release to the public. If someone would've uploaded that Tokyo video to youtube, and told me it was a drone.. I would've believed them. All "proof" we have can be contested or fabricated.

I guess you can't read Japanese.

Re: Sora: Creating video from text

#119
post #40

Visual sharpness at the expense of wider-scale coherence (see: sliding/floating walking woman in Tokyo demo or tiny people next to giant people in Lagos demo) seems to be a local optimum consistently achieved by today's SOTA models in all domains. This is neat and all but mostly just a toy. Everything I've seen has me convinced either we are optimizing the wrong loss functions or the architectures we have today are f…

Did you just recreate the infamous DropBox comment? https://news.ycombinator.com/item?id=9224

Re: Sora: Creating video from text

#120
post #24

Yeah, you just can't let all media, all the cost and hard work of millions of photographers, animators, filmmakers, etc be completely consumed and devalued by one company just because it's a very cool technical trick. The more powerful these services become the more obvious that will be. What OpenAI does is amazing, but they obviously cannot be allowed to capture the value of every piece of media ever created — it'll…

I am getting sick of these "people can't be allowed to make their own nice things easily, because of a pugnacious (and very online) interest group that wants to keep getting money" takes.
Post reply on HN