Live data from Hacker News

Veo

deepmind.google

471–480 of 539 posts

Re: Veo

#471
post #324

Earlier quoted context omitted.

I know you weren't implying this, but not every storyboard is for sharing with (or seeking approval from) decision makers. I could see this being really useful for exploring tone, movement, shot sequences or cut timing, etc.. Right now you scrape together "kinda close enough" stock footage for this kind of exploration, and this could get you "much closer enough" footage..

I think of it in terms of the anchoring bias. Imagine that your most important decisions are anchored for you by what a 10 year old kid heard and understood. Your ideas don’t come to life without first being rendered as a terrible approximation that is convincing to others but deeply wrong to you, and now you get to react to that instead of going through your own method. So if it’s an optional tool, great, but some p…

Absolutely. Everyone's creative process is different (and valid).

Re: Veo

#472
post #460

I think the thing that most perturbs me about AI is that it takes jobs that involve manipulating colours, light, shade and space directly and turns them into essay writing exercises. As a dyslexic I fucking hate writing essays. 40% of architects are dyslexic. I wouldn't be surprised if that was similar or higher in other creative industries such as filmmaking and illustration. Coincidentally 40% of the prison populat…

I would imagine and hope for interfaces to exist where the natural language prompt is the initial seed and then you'd still be able to manipulate visual elements through other ways.

This is the case today. You won't get a "perfect" image without heavy post-processing, even if that post-processing is AI enhanced. ComfyUI is the new PhotoShop and although its not an easy app to understand, once it "clicks" its the most amazing piece of software to come out of the opensource oven in a long time.

Re: Veo

#473

I think the thing that most perturbs me about AI is that it takes jobs that involve manipulating colours, light, shade and space directly and turns them into essay writing exercises. As a dyslexic I fucking hate writing essays. 40% of architects are dyslexic. I wouldn't be surprised if that was similar or higher in other creative industries such as filmmaking and illustration. Coincidentally 40% of the prison populat…

It’s seems exceedly clear to me that the primary interface for LLMs will voice.

Re: Veo

#474
post #21

The videos in this demo are pretty neat. If this had been announced just four months ago we'd all be very impressed by the capabilities. The problem is that these video clips are very unimpressive compared to the Sora demonstration which came out three months ago. If this demo was announced by some scrappy startup it would be worth taking note. Coming from Google, the inventor of the Transformer and owner of the larg…

>these sample videos are underwhelming wow the speed at which we can be blasé is terrifying. 6 months ago this was not possible, and felt this was years away! They're not underwhelming to me, they're beyond anything I thought would ever be possible. are you genuinely unimpressed? or maybe trying to play it cool?

On some level, it's healthy to retain a sense of humility at the technological marvels around us. Everything about our daily lives is impressive.

Just a few years ago, I would have been absolutely blown away by these demo videos. Six months ago, I would have been very impressed. Today, Google is rolling a product that seems second best. They're playing catch-up in a game where they should be leading.

I will still be very impressed to see videos of that quality generated on consumer grade hardware. I'll also be extremely impressed if Google manages to roll out public access to this capability without major gaffes or embarrassments.

This is very cool tech, and the developers and engineers that produced it should be proud of what they've achieved. But Google's management needs to be asking itself how they've allowed themselves to be surpassed.

Re: Veo

#475

Earlier quoted context omitted.

Human eyes are basically black and white in low light since rod cells can't detect color. But when the northern lights are bright enough you can definitely see the colors. The fact that some things are too dark to be seen by humans but can be captured accurately with cameras doesn't mean that the camera, or the AI, is "making things up" or whatever. Finally, nobody wants to see a video or a photo of a dark, gray, and…

> nobody wants to see a video or a photo of a dark, gray, and barely visible aurora Except those who want to see an accurate representation of what it looks like to the naked eye.

Exactly. I went through major gas lighting trying to see the Aurora. I just wasn't sure whether I was actually seeing it, because it always looked so different from the photos. It is absolutely maddening trying to find a realistic photo of what it looks like to the naked eye, so that you can know if what you are seeing is actually the Aurora and not just clouds

Re: Veo

#476

Earlier quoted context omitted.

There's a middle ground here. I saw the northern lights with my own eyes just days ago and it was mostly grey. I saw some color. But when I took a photo with a phone camera, the color absolutely popped . So it may be that you've seen more color than any photo, but the average viewer in Seattle this past weekend saw grey-er with their eyes and huge color in their phone photos. (Edit: it was still super-cool even if gr…

The hubris of suggesting that your single experience of vaguely seeing the northern lights one time in Seattle has now led to a deep understanding of their true "color" and that the other person (perhaps all other people?) must be fooling themselves is... part of what makes HN so delightful to read. I've also seen the northern lights with my own eyes. Way up in the arctic circle in Sweden. Their color changes along w…

The person they were responding to was saying that the people reporting grays were wrong, and that they had seen it and it was colorful. If anything, you should be accusing that person of hubris, not GP. All GPS point was, is that it can differ in different situations. They used the example of Seattle to show that the person they were responding to is not correct that it is never gray and dull.

Re: Veo

#477
post #374

OpenAI has the model advantage. Google and Apple have the ecosystem advantage. Apple in particular has the deeper stack integration advantage. Both Apple and Google have a somewhat poor software innovation reputation. How does it all net out? I suspect ecosystem play wins in this case because they can personalize more deeply.

Google has a deep addiction to AdWords revenue which makes for a significant disadvantage. Nomatter how good their technology, they will struggle internally with deploying it at scale because that would risk their cash cow. Innovator’s dilemma.

As of 2020, AdWords represented over 80% of all Google revenue [1] while in 2021 7% of Google’s revenue came from cloud [2].

[1] https://www.cnbc.com/2021/05/18/how-does-google-make-money-a...?

[2] https://aag-it.com/the-latest-cloud-computing-statistics/?t

Re: Veo

#478
post #353

The first thing I will do when I get access to this is ask it to generate a realistic chess board. I have never gotten a decent looking chessboard with any image generator that doesn't have deformed pieces, the correct number of squares, squares properly in a checkerboard pattern, pieces placed in the correct position, board oriented properly (white on the right!) and not an otherwise illegal position. It seems to be…

Most generative AI will struggle when given a task that requires something more less exact. They're probably pretty good at making something "chessish".

Re: Veo

#479

I think the thing that most perturbs me about AI is that it takes jobs that involve manipulating colours, light, shade and space directly and turns them into essay writing exercises. As a dyslexic I fucking hate writing essays. 40% of architects are dyslexic. I wouldn't be surprised if that was similar or higher in other creative industries such as filmmaking and illustration. Coincidentally 40% of the prison populat…

Terence McKenna predicted this:

“The engineers of the future will be poets.”

Re: Veo

#480
post #468

Earlier quoted context omitted.

Speaking has sound but that is still just words with the same logic structure. "Colours, light, shade and space" have entirely different logic.

Very interesting. Thank you for the perspective, it is extremely illuminating. What is a user interface which can move from color, light, shade, and space to images or text? Could there be an architecture that takes blueprints and produces text or images?

[deleted]
Post reply on HN