Live data from Hacker News

Genie 2: A large-scale foundation world model

deepmind.google

421–430 of 436 posts

Re: Genie 2: A large-scale foundation world model

#421
post #146

For all that this is lauded as a "prototyping tool", it's frustrating to see Genie2 discarding entire portions of the concept art demo. The original images drawn by Max Cant have these beautiful alien creatures. Large ones floating, and small ones being herded(?). Genie2 just ignores these beautiful details entirely: > That large alien? That's a tree. > That other large alien? It's a bush. > That herd of small creatu…

That's an odd thing to complain about. Focusing on such a minor issue feels overly critical at this stage, like anything less than a pixel perfect 3D world representation of the source image is unacceptable. Insulting? Come on... Max Cant works at DeepMind so I'm sure he's fine.

> That's an odd thing to complain about. Focusing on such a minor issue feels overly critical at this stage

Welcome to HackerNews.

Re: Genie 2: A large-scale foundation world model

#422
post #104

Earlier quoted context omitted.

Do you want household androids? Because this kind of stuff is on the level of research a bery large step towards that. Think as it as ab example where we can make a model understand a lot of physical common sense stuff, which is the goal for robotics right now.

This is really not the avenue for house-hold robots. Interacting with the actual physical world is very different from creating a video game.

No, they actually did use a genie-like model to train robots on household chores.

Page 8 of the Genie 1 paper: https://arxiv.org/abs/2402.15391

Re: Genie 2: A large-scale foundation world model

#423

Earlier quoted context omitted.

> They only spit out derivative works of what they’ve been trained on How is that different from humans? Do we get magic inspiration totally separate from anything we’ve learned? Show me any great book, song, movie, building, sculpture, painting. I will tell you the influences the artist trained on.

Humans are obviously influenced by others but we can also invent novel things that didn't exist before. LLMs trained on the outputs of LLMs collapse into gobbledygook whereas humans trained on humans build civilisation.

Humans trained on human output also build death cults and other harms. And humans believe that nonsense.

I’m not sure “can produce good outputs, can produce terrible outputs” is a good way to differentiate humans and LLMs.

Re: Genie 2: A large-scale foundation world model

#424

Earlier quoted context omitted.

> They only spit out derivative works of what they’ve been trained on How is that different from humans? Do we get magic inspiration totally separate from anything we’ve learned? Show me any great book, song, movie, building, sculpture, painting. I will tell you the influences the artist trained on.

They're different because they're trying to find the most likely output, and humans usually. You can ask and LLM to make weird combinations and use unusual framings, but it's only going to do so once you've already come up with that.

I asked ChatGPT “ Write a one paragraph pitch for a novel that combines genres and concepts in a way that’s never been done before.”

I’m not going to claim this is Pulitzer-worthy, but it seems fairly novel:

> In Spiritfall: A Symphony of Rust and Rose Petals, readers traverse the borders of time, taste, and consciousness in a genre-bending epic that effortlessly fuses neo-noir detective intrigue, culinary magic realism, and post-biotechnological body horror under the simmering threat of a cosmic opera. Set in a floating, living city grown from engineered coral-harps, the story follows a taste-shaper detective tasked with unraveling the murder of an exiled goddess whose voice once controlled the city’s very tides. As he navigates sentient cooking knives, ink-washed memory fractals, and teahouses that serve liquid soul fragments, he uncovers conspiracies binding interdimensional dream-chefs to cybernetic shamans, and finds forbidden love in a quantum greenhouse of sentient spices. Every chapter refracts expectations, weaving together genres never before dared, leaving readers both spellbound and strangely hungry for more.

Re: Genie 2: A large-scale foundation world model

#425
post #343

Earlier quoted context omitted.

It would not surprise me if most people could not tell whether some story about the human condition is human or AI generated. Excluding actual visual artists that have specific context of the craft, most people already can't tell AI art from human art when put to a blind test.

As far as I know know AI art can't really follow instructions so it's actually very, very easy to tell the difference if you aren't biasing the test by allowing vague instructions permitting random results to be considered acceptable. "Here's a photo of me and my wife, draw me and my wife as a cowboy in the style of a Dilbert cartoon shooting a gun in the air" can't be done by AI as far as I know, which is why artist…

Last time I checked GenAI it wasn't able to handle multiple people, but giving Midjourney a picture of yourself, and asking it to "draw me as a cowboy in the style of a Dilbert cartoon shooting a gun in the air" is totally a thing it will do. Without a picture of you to test on, we can't debate how well the image looks like you, but here's one of Jackie Chan: https://imgur.com/a/6cBrHWd

Re: Genie 2: A large-scale foundation world model

#426

Earlier quoted context omitted.

They're different because they're trying to find the most likely output, and humans usually. You can ask and LLM to make weird combinations and use unusual framings, but it's only going to do so once you've already come up with that.

I asked ChatGPT “ Write a one paragraph pitch for a novel that combines genres and concepts in a way that’s never been done before.” I’m not going to claim this is Pulitzer-worthy, but it seems fairly novel: > In Spiritfall: A Symphony of Rust and Rose Petals, readers traverse the borders of time, taste, and consciousness in a genre-bending epic that effortlessly fuses neo-noir detective intrigue, culinary magic real…

...that pitch is a mess. The majority of it is nonsense and it doesn't sound like a good story to me (I think. I can hardly parse it.)

Re: Genie 2: A large-scale foundation world model

#427

Earlier quoted context omitted.

As far as I know know AI art can't really follow instructions so it's actually very, very easy to tell the difference if you aren't biasing the test by allowing vague instructions permitting random results to be considered acceptable. "Here's a photo of me and my wife, draw me and my wife as a cowboy in the style of a Dilbert cartoon shooting a gun in the air" can't be done by AI as far as I know, which is why artist…

Last time I checked GenAI it wasn't able to handle multiple people, but giving Midjourney a picture of yourself, and asking it to "draw me as a cowboy in the style of a Dilbert cartoon shooting a gun in the air" is totally a thing it will do. Without a picture of you to test on, we can't debate how well the image looks like you, but here's one of Jackie Chan: https://imgur.com/a/6cBrHWd

Are you saying you can upload a picture to mid journey that it will use as a reference?

Jackie Chan is not a good example because he's a famous person it may have been trained on. I used myself as an example because it would be something that is novel to the AI, it would not be able to rely on it's training to draw me, as I am not famous.

Re: Genie 2: A large-scale foundation world model

#428

Earlier quoted context omitted.

Last time I checked GenAI it wasn't able to handle multiple people, but giving Midjourney a picture of yourself, and asking it to "draw me as a cowboy in the style of a Dilbert cartoon shooting a gun in the air" is totally a thing it will do. Without a picture of you to test on, we can't debate how well the image looks like you, but here's one of Jackie Chan: https://imgur.com/a/6cBrHWd

Are you saying you can upload a picture to mid journey that it will use as a reference? Jackie Chan is not a good example because he's a famous person it may have been trained on. I used myself as an example because it would be something that is novel to the AI, it would not be able to rely on it's training to draw me, as I am not famous.

yes. here is a video tutorial where a cat is being used as a reference image

https://youtu.be/9dOECM76l_c?t=45

Re: Genie 2: A large-scale foundation world model

#429

Earlier quoted context omitted.

Tech demo, doesn't generalise.

Well, Waymo.

"Well Waymo" is not DeepMind.

Look. The other poster also said "Waymo" but I'm talking about DeepMind. It's DeepMind that promises to conquer the world with Deep Reinforcement Learning, and it's DeepMind that keeps showing us how great their DRL agents work in virtual worlds, like minecraft or starcraft, or how well they work on Chess and Go, but still haven't been able to demonstrate the application of those powerful learning approaches to real-world environments, except for very strictly controlled ones. Waymo's stuff works in the real world (although they do have remote safety drivers much as they try to downplay the fact) but they're also not pretending that they'll do it all with one big DRL "generalist" agent. That's DeepMind's schtick.

For example, it was, I believe, DeepMind that recently publicised some results about legged robot football, where the robots were controlled by agents trained with DRL in a simulation. That's robot football: two robots (yeah, no teams) kicking a ball in the safest of safe environments: a (reduced-size) football field with artificial grass, probably padded underneath (because robots) and no other objects in the play area (except anxious researchers who have to pull the robots on their feet once in a while). Running in the physical world in principle, but in practice nothing but a tech demo.

Or take the other Big Idea, where they had a few dozen robot arms reaching for various little plastic bits in a (specially-made) box to try and learn object manipulation by real-world DRL. I can find a link to those things if you want, but that robot arm project was a few years ago and you haven't heard anything from them since because it was a whole load of overpromising and it failed.

That kind of thing just doesn't generalise. More than that: it's a total waste of time and money. And yet DeepMind keeps banging the drum. They keep trying to convince everyone and themselves that training DRL agents in virtual environments has anything to do with the real world, and that it's somehow the road to AGI. "Reward is all you need". Yeah, OK.

Btw, Waymo is not using DRL, at least not exclusively. They use all sorts of techniques but from what I understand they do a hell of a lot of good, old-fashioned, manual programming to deal with all the stuff that magickal deep learning in the sky can't deal with.

Re: Genie 2: A large-scale foundation world model

#430

Earlier quoted context omitted.

I'll take that bet

Email on profile! I'll match any 5-figure amount you propose. I also know an escrow service we can trust.

Day two of the two year bet:

https://x.com/mrjonfinger/status/1865161230706520472

Let's do this, Shasseem.

Post reply on HN