Live data from Hacker News

Veo

deepmind.google

371–380 of 539 posts

Re: Veo

#371

Earlier quoted context omitted.

I can see using these video generators to create video storyboards. Especially if you can drop in a scribbled sketch and a prompt for each tile.

That sounds actively harmful. Often we want story boards to be less specific so as not to have some non artist decision maker ask why it doesn't look like the storyboard. And when we want it to match exactly in an animatic or whatever, it needs to be far more precise than this, matching real locations etc.

I hadn't thought about that in movie context before, but it totally makes sense.

I've worked with other developers that want to build high fidelity wire frames, sometimes in the actual UI framework, probably because they can (and it's "easy"). I always push back against that, in favor of using whiteboard or Sharpies. The low-fidelity brings better feedback and discussion: focused on layout and flow, not spacing and colors. Psychologically it also feels temporary, giving permission for others to suggest a completely different approach without thinking they're tossing out more than a few minutes of work.

I think in the artistic context it extends further, too: if you show something too detailed it can anchor it in people's minds and stifle their creativity. Most people experience this in an ironically similar way: consider how you picture the characters of a book differently depending on if you watched the movie first or not.

Re: Veo

#372
post #324

Earlier quoted context omitted.

That sounds actively harmful. Often we want story boards to be less specific so as not to have some non artist decision maker ask why it doesn't look like the storyboard. And when we want it to match exactly in an animatic or whatever, it needs to be far more precise than this, matching real locations etc.

I know you weren't implying this, but not every storyboard is for sharing with (or seeking approval from) decision makers. I could see this being really useful for exploring tone, movement, shot sequences or cut timing, etc.. Right now you scrape together "kinda close enough" stock footage for this kind of exploration, and this could get you "much closer enough" footage..

I think of it in terms of the anchoring bias. Imagine that your most important decisions are anchored for you by what a 10 year old kid heard and understood. Your ideas don’t come to life without first being rendered as a terrible approximation that is convincing to others but deeply wrong to you, and now you get to react to that instead of going through your own method.

So if it’s an optional tool, great, but some people would be fine with it, some would not.

Re: Veo

#373
OpenAI has the model advantage.

Google and Apple have the ecosystem advantage.

Apple in particular has the deeper stack integration advantage.

Both Apple and Google have a somewhat poor software innovation reputation.

How does it all net out? I suspect ecosystem play wins in this case because they can personalize more deeply.

Re: Veo

#374

OpenAI has the model advantage. Google and Apple have the ecosystem advantage. Apple in particular has the deeper stack integration advantage. Both Apple and Google have a somewhat poor software innovation reputation. How does it all net out? I suspect ecosystem play wins in this case because they can personalize more deeply.

Google has a deep addiction to AdWords revenue which makes for a significant disadvantage. Nomatter how good their technology, they will struggle internally with deploying it at scale because that would risk their cash cow. Innovator’s dilemma.

Re: Veo

#375

Earlier quoted context omitted.

That not true, they look grey when they aren't bright enough, but they can look green or red to the naked eyes if they are bright. I have seen it myself and yes I was disappointed to see only grey ones last week. see: https://theconversation.com/what-causes-the-different-colour...

> [Aurora] only appear to us in shades of gray because the light is too faint to be sensed by our color-detecting cone cells." > Thus, the human eye primarily views the Northern Lights in faint colors and shades of gray and white. DSLR camera sensors don't have that limitation. Couple that fact with the long exposure times and high ISO settings of modern cameras and it becomes clear that the camera sensor has a much…

[deleted]

Re: Veo

#376

With so much recent focus by OpenAI/Google on AI's visual capabilities, does anyone know when we might see an OCR product as good as Whisper for voice transcription? (Or has that already happened?) I had to convert some PDFs and MP3s to text recently and was struck by the vast difference in output quality. Whisper's transcription was near-flawless, all the OCR softwares I tried struggled with formatting, missed words…

You might enjoy this breakdown of the lengths one person went through to take advantage of the iOS vision API and creating a local web service for transcribing some very challenging memes: https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip... discussed on HN: https://news.ycombinator.com/item?id=34315782

This is a work of fucking art.

Re: Veo

#377
Oddly enough, I predict the final destination for this train will be for moving images to fade into the background. Everything will have a dazzling sameness to it. It's not unlike the weird place that action movies and pop music have arrived. What would have been considered unbelievable a short time ago has become bland. It's probably more than just novelty that's driving the comeback of vinyl.

Re: Veo

#378
post #374

OpenAI has the model advantage. Google and Apple have the ecosystem advantage. Apple in particular has the deeper stack integration advantage. Both Apple and Google have a somewhat poor software innovation reputation. How does it all net out? I suspect ecosystem play wins in this case because they can personalize more deeply.

Google has a deep addiction to AdWords revenue which makes for a significant disadvantage. Nomatter how good their technology, they will struggle internally with deploying it at scale because that would risk their cash cow. Innovator’s dilemma.

Google Cloud and cloud services generated almost 9.57 billion. That's up 28% from prior:

https://www.crn.com/news/networking/2024/google-cloud-posts-...

They are embedding their models not only widely across their platforms suite of internal products and devices, but also computationally via API for 3rd party development.

Those are all free from any perceived golden handcuffs that AdWords would impose.

Re: Veo

#379
post #353

The first thing I will do when I get access to this is ask it to generate a realistic chess board. I have never gotten a decent looking chessboard with any image generator that doesn't have deformed pieces, the correct number of squares, squares properly in a checkerboard pattern, pieces placed in the correct position, board oriented properly (white on the right!) and not an otherwise illegal position. It seems to be…

Similarly the Veo example of the northern lights is a really interesting one. That's not what the northern lights look like to the naked eye - they're actually pretty grey. The really bright greens and even the reds really only come out when you take a photo of them with a camera. Of course the model couldn't know that because, well, it only gets trained on photos. Gets really existential - simulacra energy - maybe a…

For decades, game engines have been working on realistic rendering. Bumping quality here and there.

The golden standard for rendering has always been cameras. It’s always photo-realistic rendering. Maybe this won’t be true for VR, but so far most effort is to be as good as video, not as good as the human eye.

Any sort of video generation AI is likely to have the same goal. Be as good as top notch cameras, not as eyes.

Re: Veo

#380
post #219

From a filmmaking standpoint I still don't think this is impactful. For that it needs a "director" to say: "turn the horse's head 90˚ the other way, trot 20 feet, and dismount the rider" and "give me additional camera angles" of the same scene. Otherwise this is mostly b-roll content. I'm sure this is coming.

I dont think "turn the horse's head 90˚" is the right path forward. What I think is more likely and more useful is: here is a start keyframe and here is a stop keyframe (generated by text to image using other things like controlnet to control positioning etc.) and then having the AI generate the frames in between. Dont like the way it generated the in between? Choose a keyframe, adjust it, and rerun with the segment…

So a declarative keyframe of "the horses head is pointed forward" and a second one of "the horse is looking left"

And let the robot tween?

Vs an imperative for "tween this by turning the horse's head left"

Post reply on HN