Live data from Hacker News

Veo

deepmind.google

331–340 of 539 posts

Re: Veo

#331
post #291

Earlier quoted context omitted.

Well, as a counterpoint, Apple did become a $2 trillion dollar company... Distortion is easiest when the products really work. :)

Apple got up to $3 trillion back in 2023.

Indeed, and they’re at 2.87T today… Built largely on differentiated high-margin products, which is not how I would describe OpenAI. I should clarify that I’m a fan of both companies, but the reality is that OpenAI’s business model depends on how well it can commoditize itself.

Re: Veo

#332

Made an album in 10 mins. Typically as a techno DJ I'd mix them together so they sound kinda bare right now. Here's my 10 minutes to 12:09 album debut: https://on.soundcloud.com/FAXkJrLrC2JjoAyu7

Even as a very inexperienced musician I think I can say these are not very compelling examples? They sound like unfinished sketches that took a few minutes to make each, but with no overarching theme and weirdly low fidelity. An absolute beginner could make better things just by messing around with a groovebox.

Re: Veo

#333

Earlier quoted context omitted.

We have to take account that this community (good chunk have stakes in YC and a lot to gain from secondary shares in OpenAI) and platform is going to favor its own and be aware that Sam Altman is the golden boy of YC's founder after all. So of course you are going to see snarky comments and straight up denial in the competition. We saw that yesterday in the comments with the release of GPT4o in anticipation of Gemini…

The amount of copium in this response is astounding. Yes, there is a noticeable negative response from HN towards Google, and there has always been especially when speaking about their weird product management practices and incentives. Google hasn't launched any notable (and still surviving, Stadia being a sad example of this) consumer product or service in the last 10 years. But to suggest there is a Sam Altman / Op…

My argument wasn't that there was a cabal of shady investors trying to influence perception here. your observation is certainly valid there is general disdain for Google but specifically I'm calling out people that were blatantly telling lies and making outlandish claims and attacking others who were simply pointing out that some of those people have financial motives (either being backed by YC or seek to benefit from the work of others).

None of this is surprising to me and shouldn't shock you. You are literally on a site called Ycombinator. Had this been another platform without ties to investments or drawing from crowd that actively seeks to enrich themselves through participation in a narrative, this wouldn't even be a thing.

Large number of people who read my comment seems to agree and this whole worldcoin thing seems to me just another distraction (We've already been through why that was shady but we are talking about something different here).

Re: Veo

#334

Earlier quoted context omitted.

There is now on that second link: >The videos below were edited by the artists, who creatively integrated Sora into their work, and had the freedom to modify the content Sora generated.

Ha, here's an archive from yesterday for posterity. https://web.archive.org/web/20240513050023/https://openai.co... They also just added a link to the making-of video.

That's hilarious. Your comment clearly got seen by someone.

Re: Veo

#335
post #213

An interesting thing that Google does is to watermark the AI generated videos using the [SynthID technology]( https://deepmind.google/technologies/synthid/ ). It seems that the SynthID is not only for AI generated video but for image, text and audio.

I would like a bit more convincing that the text watermark will not be noticeable. AI text already has issues with using certain words to frequently. Messing with the weights seems like it might make the issue worse

Re: Veo

#336
post #46

Not nearly as impressive as Sora. Sora was impressive because the clips were long and had lots of rapid movement since video models tend to fall apart when the movement isn't easy to predict. By comparison, the shots here are only a few seconds long and almost all look like slow motion or slow panning shots cherrypicked because they don't have that much movement. Compare that to Sora's videos of people walking in rea…

Not just that, but anything with a subject in it felt uncanny valleyish... like that cowboy clip, the gate of the horse stood out as odd and then I gave it some attention . It seems like a camel's gate. And whole thing seems to be hovering, gliding rather than walking. Sora indeed seems to have an advantage

I thought a camel's gait is much closer to two legs moving almost at the same time. Granted, I don't see camels often. Out of curiosity can you explain that more?

Re: Veo

#337

Made an album in 10 mins. Typically as a techno DJ I'd mix them together so they sound kinda bare right now. Here's my 10 minutes to 12:09 album debut: https://on.soundcloud.com/FAXkJrLrC2JjoAyu7

[deleted]

Re: Veo

#338
post #219

From a filmmaking standpoint I still don't think this is impactful. For that it needs a "director" to say: "turn the horse's head 90˚ the other way, trot 20 feet, and dismount the rider" and "give me additional camera angles" of the same scene. Otherwise this is mostly b-roll content. I'm sure this is coming.

There's also the whole "oh you have no actual model/rigging/lighting/set to manipulate" for detail work issue.

That said, I personally think the solution will not be coming that soon, but at the same time, we'll be seeing a LOT more content that can be done using current tools, even if that means a dip in quality (severely) due to the cost it might save.

Re: Veo

#339
post #335
post #213

An interesting thing that Google does is to watermark the AI generated videos using the [SynthID technology]( https://deepmind.google/technologies/synthid/ ). It seems that the SynthID is not only for AI generated video but for image, text and audio.

I would like a bit more convincing that the text watermark will not be noticeable. AI text already has issues with using certain words to frequently. Messing with the weights seems like it might make the issue worse

Not to mention, when does he get applied? If I am asking an llm to transform some data from one format to another, I don't expect any changes other than the format.

Re: Veo

#340

> Veo's cutting-edge latent diffusion transformers reduce the appearance of these inconsistencies, keeping characters, objects and styles in place, as they would in real life. How is this achieved? Is there temporal memory between frames?

Probably similar to Sora, a patchified vision transformer, you sample a 3d patch (third dimension is time) instead of a 2d patch
Post reply on HN