Live data from Hacker News

Veo

deepmind.google

531–539 of 539 posts

Re: Veo

#531
post #353

The first thing I will do when I get access to this is ask it to generate a realistic chess board. I have never gotten a decent looking chessboard with any image generator that doesn't have deformed pieces, the correct number of squares, squares properly in a checkerboard pattern, pieces placed in the correct position, board oriented properly (white on the right!) and not an otherwise illegal position. It seems to be…

Mine is generation of any actual IBM PC/XT computer. All of the training sets either didn't include actual IBM PCs in them, or they labeled all PC compatibles "IBM PC". Whatever the reason, no generative AI today, whether commercial or open-source, can generate any picture of an IBM PC 5150. Once that situation improves, I'll start taking notice.

Re: Veo

#532

Earlier quoted context omitted.

I guess in the near future prompts can be replaced by a live editing conversation with the AI, like talking to a phantom draughtsman or a camera operator / movie team. The AI will adjust while you talk to it and can also ask questions. ChatGPT already allows this workflow to some extent. You should try it out. I just talked to ChatGPT on my phone to test it. I think I will not go back to text for these purposes. It's…

I would need to be able to talk and draw at the same time, which is how I interact with co-workers and clients.

This would be feasible. Even right now, but I am not sure how much delay is tolerable.

If you use tablets or screens, I would imagine a two screen/tablet setup, where on one screen there is a variant gallery with AI output and on the other screen there is the drawing area. The drawing constantly refreshes the gallery.

One can click on images in the gallery to move the whole image or parts of it into the drawing area. Additionally voice input leads to a conversation in the background that affects the variants as well. The process would be a mix of sketching, overpainting and voice-controlled image manipulation.

Automatic image segmentation that is automatically applied to all variants would make it easy to move objects/parts from the variants easily. The pulled parts would be stitched automatically into the drawing area, as some kind of super charged collage technique.

Maybe the variant gallery would be more like an idea board. You would say things like: "Can you make a variant with clinkers", "Please add garden furniture near the pond." etc. In the gallery these images would pop up and you can pick what you like from it.

Re: Veo

#533
post #353

The first thing I will do when I get access to this is ask it to generate a realistic chess board. I have never gotten a decent looking chessboard with any image generator that doesn't have deformed pieces, the correct number of squares, squares properly in a checkerboard pattern, pieces placed in the correct position, board oriented properly (white on the right!) and not an otherwise illegal position. It seems to be…

> It seems to be an "AI complete" problem.

Conventionally this term means the opposite -- problems that AI unlocks that conventional computing could not do. Conventional computing can render a very wide range of different stylized chess boards, but when an ML technique like diffusion is applied to this mundane problem, it falls apart.

Re: Veo

#534

Earlier quoted context omitted.

Quote the entire sentence, not just a portion of it.

I don't see how that's relevant, unless you're able to possess people looking at their phones to experience what they're experiencing.

To add a bit of color (ha) I was with my color-sighted spouse at a spot well known for panoramic views. 50ish people there. Many conversations happening around me.

“I can’t see anything” “Maybe that’s something over there?” “What’s everyone looking at?”

Someone shows their phone.

“Ooh!” “How do you turn on night mode?” “Wow it’s so much clearer on the phone!”

So I can’t know what their eyes see or what they really think, I could hear what came out of their mouths.

I don’t think this is an instance that warrants deep philosophical skepticism about the nature of truth or the impossibility of knowledge.

Re: Veo

#535

Earlier quoted context omitted.

> Google’s problem has always been in product follow through. Google is large enough to not care about small opportunities. It ends up focusing on bigger opportunities that only it can execute well. Google's ability to shut down products that dont work is an insult to user but a very good corporate strategy and they deserve kudos for that. Now, coming back to the "follow through". Google Search, Gmail, Chrome, Androi…

> coming back to the "follow through". Google Search, Gmail, Chrome, Android, Photos, Drive, Cloud etc. all are excellent examples of Google's long term commitment Do you have any examples of something they launched in the last decade?

Pixel smartphones: Launched in 2016 Google Home smart speaker: Launched in 2016 Google Wifi mesh Wi-Fi system: Launched in 2016 Google Nest smart display: Launched in 2018 Google Nest Wifi mesh Wi-Fi system: Launched in 2019 Stadia Cloud gaming platform*: Launched in 2019 Google Pay (formerly known as Tez): 2028

Re: Veo

#536

I think we should all take a pause and just appreciate the amazing work Google, OpenAI, MS and many others including those in academia have done. We do not know if Google or OpenAI or someone else is going to win the race but unlike many other races, this one makes the entire humanity move faster. Keep the negativity aside and appreciate the sweat and nights people have poured into making such things happen. Majority…

Majority of the people building the ai are artists having their work stolen or workers earning extremely low wages to label gory and csam data to a point where it hurts their mental health.

> where it hurts their mental health.

Why are they working there then ?

Re: Veo

#537
post #423

Earlier quoted context omitted.

The cost to switch to new models is negligible. People will switch to Sora if its better instantly I’ve switched to Opus from GPT-4 for coding and it was non-trivially easy

I think you used non-trivially wrong there, bud.

hah, I did :)

Re: Veo

#539
post #482

Earlier quoted context omitted.

So, the fall of the skilled professional and the rise of anybody who knows how to write prompts?

The AI we have today has very little to do with writing prompts, you still need to understand, correct, glue and edit the results and that is most of the work so you still need skilled professionals.

Pretty much everythnig I see about using AI is based around the construction of proper prompts to achieve the type of output you require. Could you explain how prompts are not a big part of interrfacing with AI?
Post reply on HN