Live data from Hacker News

Sora: Creating video from text

openai.com

951–960 of 1001 posts

Re: Sora: Creating video from text

#951
I hope this doesn't get buried...

As several others have pointed out, realism of these models will continue to improve, and will soon be economically useful for producing beautiful or functional artifacts - however prompt adherence (getting what you want or intend) of the models is growing much more slowly.

However I think we have a long ways to go before we'll see a decent "AI Film" that tells a compelling story - and this has nothing to do with some sort of naturalistic fallacy that appeals to some innate nature of humans!

It comes down to the dataset and the limits of human creators in their ability to communicate their process. Image-Text and Video-Text pairs are mostly labeled by semi-skilled humans who describe what they see in detail. They are, for the most part, very good at capturing the obvious salient features of an image or a video. "reflections of the neon lights glisten in the sidewalk". However, what you see in a movie scene is the sum total of dozens if not hundreds of influences, large and subtle. Choices made by the actors, camera operators, lighting designers, sound designers, costuming, makeup, editors, etc... Most people are not trained to recognize these choices at all, or might not even be aware that there are choices to make. We (simply) see "Joaqin Phoenix is making awkward small-talk in the elevator with other office workers".

So much of what we experience processes on subconscious and emotional and purely sensory levels, we don't elevate those lower-level qualia to our higher-brain's awareness and label them with vocabulary without intentional training (such as tasting wine, coffee, beer, etc - developing a palate is an act of sensory-vocabulary alignment).

However, despite not raising these things to our intentional awareness, it has an influence on us -- often the desired impact of the person who made that choice in the first place. The overall effect of all of these intentional choices makes things 'feel right'.

There's no fundamental reason AI can't produce an output that has the same effect as those choices, however finding each little choice is like a needle in a haystack. Accurate labeling of the training data tells the AI where to look -- but the people labeling the data are probably not well-versed in all of the little intentional choices that can be made when creating a piece of video-media.

Beyond the issue of the labeling folks being trained in the art itself, there's the problem too of the artists themselves not being able to fully articulate their (numerous, little, snowflake-into-avalanche) choices - or simply not articulating it even if they could. Ask Jackson Pollock about paint viscosity and you'll learn a great deal, but ask about abstract painting composition and there's this ineffable gap that language seems ill-suited to cross. The painter paints what they feel, and they hope that feeling is conveyed to the viewer - but you'd be hard pressed to recreate "Autumn Rhythm (Number 30)" if you had to transmit the information via language and hope they interpreted it correctly. Art is simultaneously vague and specific!

So, to sum up the problem of conveying your intent to the model:

- The training data labels capture obvious or salient features, but not choices only visible to the trained eye

- The material itself is created by human artists who might not even be able to explain all of their choices in words

- You the prompter might not have the vocabulary that captures succinctly and specifically the intended effect

- The end result will necessarily be not quite what you imagined in your mind's eye as a result of all of this missing information

You can still get good results if you tell it to copy something, because the label "Tarantino" captures a lot of detail, even all the little things you and the training data would never have labeled in words. But it won't be yours and - until we have an army of trained artists providing precise descriptions for training data in their area of expertise, and you know how to speak those artists' language - it can't be yours.

Re: Sora: Creating video from text

#952
post #873

This is both amazing and saddening to me. All our cultural legacy is being fed into a monstrous machine that gives no attribution to the original content with which it was fed, and so the creative industry seems to be in great danger. Creativity being automated while humans are forced to perform menial tasks for minimum wage doesn't seem like a great future and the geriatric political class has absolutely no clue how…

Your argument is similar to the classic hand vs. power tools argument in crafting, which eventually boils down to "did you mine the ore and forge the tools yourself?" Nowadays the argument is about CNC vs. hand crafting.

This is just a point in our overall evolution. It's an exciting time. We are here to learn and adapt.

Humans can still be creative all they want. There's still the stamp of "created by a human" that will never go away. You can choose to respect it or ignore it.

Re: Sora: Creating video from text

#953

Does anyone else feel a sense of doom from these advancements? I'm definitely not a Luddite, I've been working professionally as a programmer for quite some time now, but I just can't shake this feeling. And this is not in the "I might lose my job to this" kind of feeling, that's obviously there, but it's something deeper, more sinister. I don't think I can explain it properly. Anyway, videos look incredible. I genui…

I have felt the same since Stable Diffusion came out. The thing is, things have value in society partly because human efforts were involved in its making. It's not just about the end result; people still go to concert on top of listening to studio recordings for example, and people still watch humans play chess even though it's clear that good enough algorithms can beat the best humans easily. Technology like these w…

I think you're right. A large part of the joy from creative endevours is actually getting good at something, and having other people enjoy your work. In the face of instant high quality generative AI placating the entertainment needs of the masses, we are creating a society where most people are unable to enjoy human creative expression, in part because human artists are just too slow. Attention spans are already shrinking, and after getting used to generative AI, few people will have the patience to wait for an author to write the second part of his magnum opus.

Re: Sora: Creating video from text

#954
post #406

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.

The left and right side of her face are almost... a different person.

Re: Sora: Creating video from text

#955
post #873

This is both amazing and saddening to me. All our cultural legacy is being fed into a monstrous machine that gives no attribution to the original content with which it was fed, and so the creative industry seems to be in great danger. Creativity being automated while humans are forced to perform menial tasks for minimum wage doesn't seem like a great future and the geriatric political class has absolutely no clue how…

We are all standing on the shoulders of giants, whose existence and names we will never know or acknowledge. The way these models are creative is the same way humans are. The artist that painted Mona Lisa didn't credit any of the influences and inspirations that they had. Just as cameras made many artists redundant, so too will every other new tool, and not just artist but pretty much every job. But there are still p…

> The world doesn't work that way.

The human world works that way humans make it work. Pretty much what Jody Foster's character in the movie Contact told that asshole trying to steal all the credit from her, and take her place in the mission to go visit alien dad in Pensacola.

Re: Sora: Creating video from text

#956
post #873

This is both amazing and saddening to me. All our cultural legacy is being fed into a monstrous machine that gives no attribution to the original content with which it was fed, and so the creative industry seems to be in great danger. Creativity being automated while humans are forced to perform menial tasks for minimum wage doesn't seem like a great future and the geriatric political class has absolutely no clue how…

> This is both amazing and saddening to me. All our cultural legacy is being fed into a monstrous machine that gives no attribution to the original content with which it was fed, and so the creative industry seems to be in great danger.

It is the same as what every human being is doing. We consume and we create. Sometimes creations are very good, but most of the time they are just mediocre. If the machines can create better average results, it will be due to the genius of the humans who invented those machines.

So we can be happy, that we have such beings among us and should cherish, that we will have better content to consume in the future. When you look at the world, you will see, that there are still plenty of problems to be solved for humans.

Re: Sora: Creating video from text

#957
post #406

This is insane. But I'm impressed most of all by the quality of motion . I've quite simply never seen convincing computer-generated motion before . Just look at the way the wooly mammoths connect with the ground, and their lumbering mass feels real. Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obvio…

I disagree, just look at the legs of the woman in the first video. First she seems to be limping, than the legs rotate. The mammoth are totally uncanny for me as its both running and walking at the same time. Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.

> But I think many people will be very uncomfortable with such motion very quickly.

Given the momentum in this space, I think you will have get very uncomfortable super quick about any of the shortcomings of any particular model.

Re: Sora: Creating video from text

#960
One one side, we have people who are upset because the creators of the videos in the dataset used for teaching this language model were not compensated.

On the other hand, people find the tech very impressive and there are a lot of mind blowing use-cases.

Personally, this opens up the world for me to create video ads for software projects I create, since I have no financial resources or time to actually make videos, I only know how to code. So I find it pretty exciting. It's great for solo entrepreneurs.

Post reply on HN