Live data from Hacker News

Generate videos in Gemini and Whisk with Veo 2

blog.google

41–50 of 142 posts

Re: Generate videos in Gemini and Whisk with Veo 2

#41
post #4

I think I would buy "yes" shares in a Polymarket event that predicts a motion picture created by a single person grossing more than $100M by 2027.

Everyone keeps ignoring supply and demand when talking about the impacts of AI. Let's just assume it really gets so good you can do this and it doesn't suck. Yes the costs will get so low that there will be almost no barrier to making content but if there is no barrier to making content, the ROI will be massive, and so everyone will be doing it, you can more or less have the exact movie you want in your head on deman…

That was the story with CGI too, that there would be overwhelming supply that drives prices and value toward zero.

And yet Marvel exists.

Turns out in a world of infinite supply, value comes from story, character, branding, marketing and celebrity. Those factors in combination have very limited supply and people still pay up.

I don't see any reason why AI-gen video is any different.

Re: Generate videos in Gemini and Whisk with Veo 2

#42
post #15

Earlier quoted context omitted.

I came to the same sort of conclusion when watching Kitsune, which I think was one person and VEO https://vimeo.com/1047370252 Granted, 5 minutes isn't 1h30 but it's not a million miles away either.

This actually so amateurish and cliche it's painful. The fact people like this shows that art never had a chance when the masses have no taste. This makes me depressed for artists and the future.

It won't be novel the 100th or 1000th or millionth time, and standards will rise accordingly. But for now it is, or at least 2 months ago it was.

Someone created that relatively coherent 5min animated story largely by communicating with a computer in natural language.

The masses have had plenty worse

Re: Generate videos in Gemini and Whisk with Veo 2

#43
this is semi-relevant -- and I do love how technically amazing this all is, but a massive caveat for someone who's been dabbling hard in this space, (images+video) -- I cannot emphasize enough how draining text-2- is. even when a result comes out that's kind of cool, I feel nothing because it wasn't really me who did it.

I would say 97% of the time, the results are not what I want (and of course that's the case, it's just textual input) and so I change the text slightly, and a whole new thing comes out that is once again incorrect, and then I sit there for 5minutes while some new slop churns out of the slop factory. All of this back and forth drains not only my wallet/credits, but my patience and my soul. I really don't know how these "tools" are ever supposed to help creatives, short of generating short form ad content that few people really only want to work on anyway. So far the only products spawning from these tools are tiktok/general internet spam companies.

The closest thing that I've bumped into that actually feels like it empowers artists is https://github.com/Acly/krita-ai-diffusion that plugs into Krita and uses a combination of img2img with masking and txt2img. A slightly more rewarding feedback loop

Re: Generate videos in Gemini and Whisk with Veo 2

#44

Earlier quoted context omitted.

My prediction is on track to this and this was made only 4 months ago. https://news.ycombinator.com/item?id=42368951

There may be a solo (not Han) movie good enough to compete in five years, but I doubt that Academy voters will be that welcoming of the tech that can obliterate most of their jobs by then.

Based on the training data being pop culture, we may even get a good Han Solo movie from tools like this. Starring young Ford.

Re: Generate videos in Gemini and Whisk with Veo 2

#45

Whisk itself ( https://labs.google/fx/tools/whisk ) was released a few months ago under the radar as a demo for Imagen 3 and it's actually fun to play with and surprisingly robust given its particular implementation. It uses a prompt transmutation trick (convert the uploaded images into a textual description; can verify by viewing the description of the uploaded image) and the strength of Imagen 3's actually modern t…

Why text? why not encode the image into some latent space representation, so that it can survive a round-trip more or less faithfully?

Because Imagen 3 is a text-to-image model, not an image-to-image model, so the inputs have to be some form of text. Multimodal models such as 4o image generation or Gemini 2.0 which can take in both text and image inputs do encode image inputs to a latent space through a Vision Transformer, but not reverseable or losslessly.

Re: Generate videos in Gemini and Whisk with Veo 2

#46

Whisk itself ( https://labs.google/fx/tools/whisk ) was released a few months ago under the radar as a demo for Imagen 3 and it's actually fun to play with and surprisingly robust given its particular implementation. It uses a prompt transmutation trick (convert the uploaded images into a textual description; can verify by viewing the description of the uploaded image) and the strength of Imagen 3's actually modern t…

> This tool isn’t available in your country yet

> Enter your email to be notified when it becomes available

(Submit)

> We can't collect your emails at the moment

Re: Generate videos in Gemini and Whisk with Veo 2

#47
post #8
post #7

Earlier quoted context omitted.

We've got a pretty good datapoint along that trajectory with Flow. Almost entirely one person and has grossed $36 million. https://en.wikipedia.org/wiki/Flow_(2024_film)

It was a small team for sure but not a one man show, there are 22 credits for the animation work alone, plus 13 more for sound and music, not counting the director.

> Almost entirely one person

It is closer to one than number the staff of other animated films. It's a good data point to keep in mind as AI tools enable even smaller teams to do more.

Re: Generate videos in Gemini and Whisk with Veo 2

#48
post #4

I think I would buy "yes" shares in a Polymarket event that predicts a motion picture created by a single person grossing more than $100M by 2027.

Well text generation is way ahead of video generation. Have we seen anyone create something like a best selling or high grossing novel with an LLM yet?

Re: Generate videos in Gemini and Whisk with Veo 2

#49

I am not really technical in this domain, but why is everything text-to-X? Wouldn't it be possible to draw a rough sketch of a terrain, drop a picture of the character, draw a 3D spline for the walk path, while having a traditional keyframe style editor, and give certain points some keyframe actions (like character A turns on his flashlight at frame 60) - in short, something that allows minute creative control just l…

Everything is text-to-X because it's less friction and therefore more fun. It's more a marketing thing.

There are many workflows for using generative AI to adhere to specific functional requirements (the entire ComfyUI ecosystem, which includes tools such as LoRAs/ControlNet/InstantID for persistence) and there are many startups which abstract out generative AI pipelines for specific use cases. Those aren't fun, though.

Re: Generate videos in Gemini and Whisk with Veo 2

#50

Earlier quoted context omitted.

Everyone keeps ignoring supply and demand when talking about the impacts of AI. Let's just assume it really gets so good you can do this and it doesn't suck. Yes the costs will get so low that there will be almost no barrier to making content but if there is no barrier to making content, the ROI will be massive, and so everyone will be doing it, you can more or less have the exact movie you want in your head on deman…

Totally agree. This is what Instagram and YouTube did and we got MrBeast and Kylie Jenner making billions of dollars. The cost of creating content is tapping record on your phone and the traditional "quality" as defined by visuals doesn't matter (see Quibi). Viral videos are selfies recorded in the bedroom. When you lower the barrier to entry things get more heterogeneous, not less. So you have bigger outcomes, not s…

Will the AI itself never be a good content creator/director/artist?

People are always out there tying to convince others that AI is better than humans at X. How close is it to being better than humans at being a content creator itself? Or how long before that threshold is crossed?

Post reply on HN