Live data from Hacker News

Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

blog.google

311–320 of 570 posts

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#311
Google hit the jackpot with their acquisition of YouTube and it's now paying dividend. YouTube is the largest single source of data and traffic on the Internet, and it's still growing fast. I think this data will prove incredibly important to robotics as well. It's a shame they sold Boston Dynamics in one of their dumbest ever moves because of bad PR.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#312
post #305

I'm sure by this point, and if not, pretty soon, everyone will have seen a clip of AI generated video and not thought twice about it. Its something that is only obvious when it is obvious. And the more obvious examples you see, the more non-obvious examples slip by.

I saw a video today [1]. Millions of views, ten thousand comments, not a single commenter mentioned that it's AI generated. If you look at the shadows in the background, you can see how they appear and disappear, how things float in the air, and have all the AI artifacts. The video is also slowed down (lower FPS) to overcome the length limit of AI video generator. But the point is not how we can spot these, because i…

the detail around the eyes is a dead-giveaway for AI generated video

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#313

Earlier quoted context omitted.

Definately plausible. All this is in line with my prediction for the first entirely AI generated film (with Sora or other AI video tools) to win an Oscar being less than 5 years away. And we're only 5 months in. https://news.ycombinator.com/item?id=42368951

Oscar in what category? We are about six years into transformer models. By now we can get transformers to write coherent short stories, and you can get to novel lengths with very careful iterative prompting (e.g. let the AI generate an outline, then chapter summaries, consistency notes, world building, then generate the chapters). But to get anything approaching a good story you still need a lot of manual interventio…

I just think the entire framing is wrong.

I have done quite a bit with AI generated audio/sound/music.

At some point in the process, the end result feels like your own and the models were used to create material for the end work.

At some point, using AI in the creative process will be such a given that it is left unsaid.

I would assume the screen play next year that wins the Oscar will have been helped with the aid of a language model. I can't imagine a writer not using a language model to riff on ideas. The delusional idea here is the prompt "write an Oscar winning screenplay" and that somehow that is all there is going to the creative process.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#314

Earlier quoted context omitted.

> burying the creatives among us in a pile of AI generated content. Isn't the creativity in what you put in the prompt? Isn't spending hundreds of hours manually creating and rigging models based on existing sketch the non-creative work that is being automated here?

How does a prompt describe creativity? It's a vision so far off that it's so frustrating because greater creativity came from limited tools, greater creativity came from imperfections, a different point of view, love, a slightly off touch of a painter or a guitar player, the wood of the instrument and the humidity affecting. I can go on and on, prompts are a reduction to the minimum term of everything you'd want to d…

Huh? is poem not creative? If I write a poem and tell AI to create a painting that expresses that poem visually, is that not creative?

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#315

Earlier quoted context omitted.

Not close to the way they wanted, and at too much sacrifice to the other things they were interested in or supported their family with

so they where never interested in the first place... but now they can call them self's artists after prompting a AI to make a image....

the level of discipline needed in a trade has gone down in almost every trade, including all mediums of art, for centuries

its not really anyone's problem, and generally limited to the people that made way too much of their identity to be based on a single field, that they feel they have to gatekeep it

its great that people can express themselves closer to their vision now

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#316

Who is doing all the work of making physical agents that can behave as good as a UBI generator? Something that can not just create videos, but go get groceries(hell grow my food), help a construction worker lay down tiling, help a nurse fetch supplies. https://www.figure.ai/ does not exist yet, at least not for the masses. Why are Meta and Google just building the next coder and not the next robot? Its because those…

I think I have a similar distaste for Google as you, but it's just due to limitations in the (bleeding edge...) technology. There's not like a conspiracy to _not_ make a "UBI generator" - which is surely not possible with current technology and won't be for awhile however hard Google might try.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#317

Earlier quoted context omitted.

And that's exactly how your brain work. What you call "creativity" is nothing more than exactly that: mixing ideas and thoughts you were exposed to. We're all building on others' work. The only difference is that computers do it on a much larger scale. But it's the very same process.

This is completely absurd and reductive point of view, which I always assume is a cop out. Just because it's called "machine learning" doesn't mean it actually has anything to do with how human learning or human brain works, and it's certainly not "exactly how" or "very same". There's much more going on on in human creative process, aside from mere "mixing": personal experience, understanding of the creative process,…

> personal experience, understanding of the creative process, technique and style development, subtext, hidden ideas and nuances

All of these are just human being exposed more to life and learning new skills, in other words -- having more data. LLM already learns those skills and encounters endless experience of people in its training data.

> I hate this argument

That's very subjective. You don't know how the brain works.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#318
post #298

After doing some testing, Imagen 4 doesn't score any higher than Imagen 3 on my comparison chart, approximately ~60% prompt adherence accuracy. https://genai-showdown.specr.net

The website is broken

That's unusual - I don't see anything in the logs and perf tests / website speed tests show everything is good. Maybe Cloudflare had a hiccup.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#319

After doing some testing, Imagen 4 doesn't score any higher than Imagen 3 on my comparison chart, approximately ~60% prompt adherence accuracy. https://genai-showdown.specr.net

I'm curious why you decide to declare victory after one successful attempt, but try many times for unsuccessful models. Are you trying to measure whether a model _can_ get it right, or whether it frequently _does_ get it right? I feel like success rate is a better metric here, or at least a fixed number of trials with some success rate threshold to determine model success.

It's hard to nail down a good objective metric on something that is always going to be marginally qualitative in nature but it's a good call out - I should probably add a FAQ to the site.

To clarify this test is purely a PASS/FAIL - unsuccessful means that the model NEVER managed to generate an image adhering to the prompt. So as an example, Midjourney 7 did not manage to generate the correct vertical stack of translucent cubes ordered by color in 64 gen attempts.

It's a little beyond the scope of my site but I do like the idea of maintaining a more granular metric for the models that were successful to see how often they were successful.

Re: Veo 3 and Imagen 4, and a new tool for filmmaking called Flow

#320

Earlier quoted context omitted.

In my own testing between the two this is what I’ve noticed. Imagen will follow the instructions, and 4o will often not, but produces aesthetically more pleasing images. I don’t know which is more important, but I would say that people mostly won’t pay for fun but disposable images, and I think people will pay for art but there will be an increased emphasis on the human artist. However users might pay for reliable to…

People pay for digital sticker packs so their memoji in iMessage are customized. How much money they make on sticker packs is unknown to me, but image generation platform Midjourney seems to be doing alright.

Midjourney got in REALLY early in the GenAI game despite only allowing image generation through Discord for at least a year. I heard that it was one of the largest Discord channels ever having something absurd like 20+ million members.

I'd love to see some financials but I'd tend to agree they're probably doing pretty well.

Post reply on HN