Google is the new Microsoft in the sense that they can Embrace, extend, and extinguish their competition. No matter what xAI or OpenAI or "anything"AI tries to build Google will eventually copy and destroy them at scale. AI (or A1 as our Secretary of Education calls it) is interesting because it is more difficult to protect the IP other than as trade secrets.
Generate videos in Gemini and Whisk with Veo 2
51–60 of 142 posts
Re: Generate videos in Gemini and Whisk with Veo 2
#52Google is the new Microsoft in the sense that they can Embrace, extend, and extinguish their competition. No matter what xAI or OpenAI or "anything"AI tries to build Google will eventually copy and destroy them at scale. AI (or A1 as our Secretary of Education calls it) is interesting because it is more difficult to protect the IP other than as trade secrets.
> Google will eventually copy… Weird take given Google basically invented and released through well written papers and open-source software the modern deep learning stack which all others build on. Google was being disses because they failed to make any product and were increasingly looking like Kodak/Xerox one trick pony. It seems they have woken up from whatever slumber they were in
Re: Generate videos in Gemini and Whisk with Veo 2
#53Earlier quoted context omitted.
Everyone keeps ignoring supply and demand when talking about the impacts of AI. Let's just assume it really gets so good you can do this and it doesn't suck. Yes the costs will get so low that there will be almost no barrier to making content but if there is no barrier to making content, the ROI will be massive, and so everyone will be doing it, you can more or less have the exact movie you want in your head on deman…
Totally agree. This is what Instagram and YouTube did and we got MrBeast and Kylie Jenner making billions of dollars. The cost of creating content is tapping record on your phone and the traditional "quality" as defined by visuals doesn't matter (see Quibi). Viral videos are selfies recorded in the bedroom. When you lower the barrier to entry things get more heterogeneous, not less. So you have bigger outcomes, not s…
Re: Generate videos in Gemini and Whisk with Veo 2
#54Whisk itself ( https://labs.google/fx/tools/whisk ) was released a few months ago under the radar as a demo for Imagen 3 and it's actually fun to play with and surprisingly robust given its particular implementation. It uses a prompt transmutation trick (convert the uploaded images into a textual description; can verify by viewing the description of the uploaded image) and the strength of Imagen 3's actually modern t…
Why text? why not encode the image into some latent space representation, so that it can survive a round-trip more or less faithfully?
Re: Generate videos in Gemini and Whisk with Veo 2
#55Earlier quoted context omitted.
Purely AI-generated content -- with no human authorship -- is not eligible for US copyright protection. However if a human contributes meaningfully to the final output (editing, selection, arrangement, etc.) it becomes eligible. See Thaler v. Perlmutter (2023).
>is not eligible for US copyright protection Once industry adopts AI generation, which it will, a new law will be quickly signed. In a way, not allowing copyright of AI material really only serves a tiny group of people. "We want to empower everyone to bring their ideas to market, not just those with the ability to draw them" is not a particularly evil or amoral sentiment.
Re: Generate videos in Gemini and Whisk with Veo 2
#56I am not really technical in this domain, but why is everything text-to-X? Wouldn't it be possible to draw a rough sketch of a terrain, drop a picture of the character, draw a 3D spline for the walk path, while having a traditional keyframe style editor, and give certain points some keyframe actions (like character A turns on his flashlight at frame 60) - in short, something that allows minute creative control just l…
To train these models you need inputs and expected output. For text-image pairs there exists vast amounts of data (in the billions). The models are trained on text + noise to output a denoised image.
The dataset of sketch-image pairs are significantly smaller, but you can finetune an already trained text->image model using the smaller dataset by replacing the noise with a sketch, or anything else really, but the quality of the output of the finetuned model will highly depend on the base text->image model. You only need several thousand samples to create a decent (but not excellent) finetune.
You can even do it without finetuning the base model and training a separate network that applies on top of base text->image model weights, this allows you to have a model that essentially can wear many hats and do all kinds of image transformations without affecting the performance of the base model. These are called controlnets and are popular with the stable diffusion family of models, but the general technique can be applied to almost any model.
Re: Generate videos in Gemini and Whisk with Veo 2
#57I think I would buy "yes" shares in a Polymarket event that predicts a motion picture created by a single person grossing more than $100M by 2027.
Well text generation is way ahead of video generation. Have we seen anyone create something like a best selling or high grossing novel with an LLM yet?
Re: Generate videos in Gemini and Whisk with Veo 2
#58Whisk redraws the entire thing and it barely resembles source picture.
Re: Generate videos in Gemini and Whisk with Veo 2
#59Earlier quoted context omitted.
Totally agree. This is what Instagram and YouTube did and we got MrBeast and Kylie Jenner making billions of dollars. The cost of creating content is tapping record on your phone and the traditional "quality" as defined by visuals doesn't matter (see Quibi). Viral videos are selfies recorded in the bedroom. When you lower the barrier to entry things get more heterogeneous, not less. So you have bigger outcomes, not s…
Will the AI itself never be a good content creator/director/artist? People are always out there tying to convince others that AI is better than humans at X. How close is it to being better than humans at being a content creator itself? Or how long before that threshold is crossed?
Even when AI is objectively better and dominates in blind ratings tests, there will still be a strong market for "authentic" media.
For instance we already have factories that churn out wares that are cheaper, stronger, better looking, and longer lasting than "hand made", yet people still seek out malformed $60 coffee mugs from the local artistan section in country shops.
Re: Generate videos in Gemini and Whisk with Veo 2
#60I think I would buy "yes" shares in a Polymarket event that predicts a motion picture created by a single person grossing more than $100M by 2027.
Everyone keeps ignoring supply and demand when talking about the impacts of AI. Let's just assume it really gets so good you can do this and it doesn't suck. Yes the costs will get so low that there will be almost no barrier to making content but if there is no barrier to making content, the ROI will be massive, and so everyone will be doing it, you can more or less have the exact movie you want in your head on deman…