Live data from Hacker News

Project Genie: Experimenting with infinite, interactive worlds

blog.google

161–170 of 347 posts

Re: Project Genie: Experimenting with infinite, interactive worlds

#161

If only Google had the technology for game streaming... Oh wait RIP Stadia.

Stadia was lightyears ahead, but pro-Microsoft media assassinated it with FUD

While "journalists" were busy bootlicking a laggy 720p Android only xCloud beta, Stadia was already delivering flawless 4K@60FPS in a web browser

They killed the only platform that actually worked just to protect Microsoft

This will be a textbook case study in how a legacy monopoly kills innovation to protect its own mediocrity

Microsoft won't survive the century, they are a dinosaur on borrowed time that has already lost the war in mobile, AI, and robotics

They don't create,, they just buy marrket share to suffocate the competition and ruin every product they touch

Even their cloud dominance is about to end, as they are already losing their grip on the European market to antitrust and sovereign alternatives

Re: Project Genie: Experimenting with infinite, interactive worlds

#162
post #63

Everyone here seems too caught up in the idea that Genie is the product, and that its purpose is to be a video game, movie, or VR environment. That is not the goal. The purpose of world models like Genie is to be the "imagination" of next-generation AI and robotics systems: a way for them to simulate the outcomes of potential actions in order to inform decisions.

Soft disagree; if you wanted imagination you don't need to make a video model. You probably don't need to decode the latents at all. That seems pretty far from information-theoretic optimality, the kind that you want in a good+fast AI model making decisions. The whole reason for LLMs inferencing human-processable text, and "world models" inferencing human-interactive video, is precisely so that humans can connect in…

> you don't need to make a video model. You probably don't need to decode the latents at all.

If you don't decode, how do you judge quality in a world where generative metrics are famously very hard and imprecise? How do you go about integrating RLHF/RLAF in your pipeline if you don't decode, which is not something you can skip anymore to get SotA?

Just look at the companies that are explicitly aiming for robotics/simulation, they *are* doing video models.

Re: Project Genie: Experimenting with infinite, interactive worlds

#163

Everyone here seems too caught up in the idea that Genie is the product, and that its purpose is to be a video game, movie, or VR environment. That is not the goal. The purpose of world models like Genie is to be the "imagination" of next-generation AI and robotics systems: a way for them to simulate the outcomes of potential actions in order to inform decisions.

Yeah and the goal of Instagram was to share quirky pictures you took with your friends. Now it’s a platform for influencers and brainrot; arguably it has done more damage than drugs to younger generations.

As soon as this thing is hooked up to VR and reaches a tipping point with the general public we all know exactly what is going to happen. The creation of the most profitable, addictive and ultimately dystopian technology Big Tech has ever come up with.

Re: Project Genie: Experimenting with infinite, interactive worlds

#164
post #79

Earlier quoted context omitted.

You already can, check out Marble/World Labs, Meshy, and others. It's not really as much of a boon as you'd think though, since throwing together a 3D model is not the bottleneck to making a sellable video game. You've had model marketplaces for a long time now.

> It's not really as much of a boon as you'd think though It is for filmmaking! They're perfect for constructing consistent sets and blocking out how your actors and props are positioned. You can freely position the camera, control the depth of field, and then storyboard your entire scene I2V. Example of doing this with Marble: https://www.youtube.com/watch?v=wJCJYdGdpHg

This I definitely agree with, before you had to massage the I2I and now you can just drag the camera.

Marble definitely changes the game if the game is "move the camera", just most people would not consider that a game (but hey there's probably a good game idea in there!)

Re: Project Genie: Experimenting with infinite, interactive worlds

#165

The actual breakthrough with Genie is being able to turn around and look back, and seeing the same scene that was there before. A few other labs have similar world simulators, but they all struggle badly with keeping coherence of things not in view. Hence why they always walk forwards and never look around.

What about Fei Fei Li's lab? I think they are generating true 3D worlds rather than frames of a video?

Although that probably precludes her from having animations in those worlds...

Re: Project Genie: Experimenting with infinite, interactive worlds

#167
Now I can't stop thinking about _The Experience Machine_ by Andy Clark. It theorizes that this is how humans navigate and experience the real world: Our brains generate what we think the world around is like and our senses don't so much directly process visual information but instead act like a kind of loss function for our internal simulations. Then we use that error to update our internal model of the world.

In this view, we are essentially living inside a high-fidelity generative model. Our brains are constantly 'hallucinating' a predicted reality based on past experience and current goals. The data from our senses isn't the source of the image; it's the error signal used to calibrate that internal model. Much like Genie 3 uses latent actions and frames to predict the next state of a world, our brains use 'Active Inference' to minimize the gap between what we expect and what we experience.

It suggests that our sense of 'reality' isn't a direct recording of the world, but a highly optimized, interactive simulation that is continuously 'regularized' by the photons hitting our retinas.

Re: Project Genie: Experimenting with infinite, interactive worlds

#168
post #28

Really great to see this released! Some interesting videos from early-access users: - https://youtu.be/15KtGNgpVnE?si=rgQ0PSRniRGcvN31&t=197 walking through various cities - https://x.com/fofrAI/status/2016936855607136506 helicopter / flight sim - https://x.com/venturetwins/status/2016919922727850333 space station, https://x.com/venturetwins/status/2016920340602278368 Dunkin' Donuts - https://youtu.be/lALGud1Ynhc?si=…

I liked that first one and I hope someone creates one of going back to dinosaur age, i want to see that.

One step closer to the science-based dinosaur MMO we were promised.

Re: Project Genie: Experimenting with infinite, interactive worlds

#169
I am stumped. Am I misreading, or are the folks at Google deliberately confounding two interpretations of "world model"? Dont get me wrong, this is really cool, and it will undoubtedly have its use. But what I am seeing is an LLM that can generate textures to be fed into a human-coded 3d engine (the "world model" that is demonstrated), and I fail to see how that brings us closer to AGI. For AGI we need "world models" as in "belief systems". The AI model must be able to reason about (learned) dynamics, which I dont see reflected in the text or video.

Re: Project Genie: Experimenting with infinite, interactive worlds

#170

Everyone here seems too caught up in the idea that Genie is the product, and that its purpose is to be a video game, movie, or VR environment. That is not the goal. The purpose of world models like Genie is to be the "imagination" of next-generation AI and robotics systems: a way for them to simulate the outcomes of potential actions in order to inform decisions.

This is a video model, not a world model. Start learning on this, and cascading errors will inevitably creep into all downstream products. You cannot invent data.

Related: https://arxiv.org/abs/2601.03220

This is a paper that recently got popular ish and discusses the counter to your viewpoint.

> Paradox 1: Information cannot be increased by deterministic processes. For both Shannon entropy and Kolmogorov complexity, deterministic transformations cannot meaningfully increase the information content of an object. And yet, we use pseudorandom number generators to produce randomness, synthetic data improves model capabilities, mathematicians can derive new knowledge by reasoning from axioms without external information, dynamical systems produce emergent phenomena, and self-play loops like AlphaZero learn sophisticated strategies from games

In theory yes, something like the rules of chess should be enough for these mythical perfect reasoners that show up in math riddles to deduce everything that *can* be known about the game. And similarly a math textbook is no more interesting than a book with the words true and false and a bunch of true => true statements in it.

But I don't think this is the case in practice. There is something about rolling things out and leveraging the results you see that seems to have useful information in it even if the roll out is fully characterizable.

Post reply on HN