Live data from Hacker News

Genie 2: A large-scale foundation world model

deepmind.google

181–190 of 436 posts

Re: Genie 2: A large-scale foundation world model

#181
post #134

Earlier quoted context omitted.

No Man's Sky is kind of what you're looking for, except you may notice its quests (and worlds) become redundant quickly...I say quickly, but that became the case for me after like 30 hours of game play.

That's the kicker, LLM driven stories are likely to fall into the same trap that "infinite" procedurally generated games usually do - technically having infinite content to explore doesn't necessarily mean that content is infinitely engaging. You will get bored when you start to notice the same patterns coming up over and over again. Procgen games mainly work when the procedural parts are just a foundation for hand-c…

Yeah, generative AI can create cool looking pictures and video but so far it hasn't managed to create infinitely engaging stories. The models aren't there yet.

Re: Genie 2: A large-scale foundation world model

#182

Forget video games. This is a huge step forward for AGI and Robotics. There's a lot of evidence from Neurobiology that we must be running something like this in our brains--things like optical illusions, the editing out of our visual blind spot, the relatively low bandwidth measured in neural signals from our senses to our brain, hallucinations, our ability to visualize 3d shapes, to dream. This is the start of addin…

This is akin to navigating a lucid dream, nothing more. Conscious inputs to a visual stream synthesized from long term memory.

Re: Genie 2: A large-scale foundation world model

#183

Forget video games. This is a huge step forward for AGI and Robotics. There's a lot of evidence from Neurobiology that we must be running something like this in our brains--things like optical illusions, the editing out of our visual blind spot, the relatively low bandwidth measured in neural signals from our senses to our brain, hallucinations, our ability to visualize 3d shapes, to dream. This is the start of addin…

This looks like my dream worlds already but more colorful and a bit more detailed. But the way it hallucinates and becomes inconsistent going back and forth the same place is same as dreams.

Re: Genie 2: A large-scale foundation world model

#184
post #104
post #61

What is actually of value here? There's no actual game, it's incredibly expensive to compute, the behavior is erratic.. It's cool because it's new - but that will quickly wear off, and once that's gone, what's left? There's insane amounts of money being spent on this, and for what?

Do you want household androids? Because this kind of stuff is on the level of research a bery large step towards that. Think as it as ab example where we can make a model understand a lot of physical common sense stuff, which is the goal for robotics right now.

I don't understand how that is relevant. I certainly would not want household androids unless I'm completely disabled.

Re: Genie 2: A large-scale foundation world model

#185

Earlier quoted context omitted.

> I want a Minecraft that tells intriguing stories with infinite quest generation. Procedural infinite world gen recharged gaming, where is the procedural infinite story generation? You're not gonna get new intriguing stories from AI which only regurgitates what it's stolen. You're going to get a themeless morass without intention. I also find it amusing how your example to Siri uses one of the oldest pieces of liter…

Actually, all you need to do is to apply structured randomness to get diversity from a LLM. For example in TinyStories paper, a precursor of the Phi models: > We collected a vocabulary consisting of about 1500 basic words, which try to mimic the vocabulary of a typical 3-4 year-old child, separated into nouns, verbs, and adjectives. In each generation, 3 words are chosen randomly (one verb, one noun, and one adjectiv…

A story is not just words crammed together that sound plausible. Is the AI going to know about pacing? About character motivations? About interconnecting disparate plots? That paper sounds like it has a scientist’s conception that a story is just words, and not complex trade offs between the start of a story and its end and middle, complexity and planning that won’t come from any sort of next-token generation.

These are “stories” in the most vacuous definition possible, one that is just “and then this happened” like a child’s conception of plot

Re: Genie 2: A large-scale foundation world model

#186

Earlier quoted context omitted.

I benchmark these for my job. Just did one a couple days ago, fortitously. Gemini Advanced at $20/month is the worst of any commercial model. One constant over the last 6 months is it is indistinguishable from Llama 3.1 8B with search snippets.

I'm very curious about this. How do you benchmark them?

Good Q: this is my technically-unlaunched app site, full deets are here. https://telosnex.com/compare/ (excuse the marketing, scroll to technical details)

Context / tl;dr:

- I'm making a xplatform app, easiest way to think about it is "what if Perplexity had scripts and search was just a `script` that could be customized", and the AI provider is an abstraction that you can pick, either the bigs via API, or run locally via llama.cpp integration.

- I left my FAANG job where my last project was search x LLM x UI. I really, really want to avoid wasting a couple years building a shadow of what the bigs are. I don't want to be delusional, I want to make sure I'm building something that's at least good, even if it never succeeds in the market.

- I could test providers via API with standard benchmark Qs, but that leaves out my biggest competitors, Perplexity and SearchGPT. Also, Claude's hidden prompt has gotten long enough (6K+ tokens), that I think Claude.ai is a distinct provider.

- So, I hunt down the best two QA sets I can find for legal and medical stuff. Calculate the sample size that gives me a 95% confidence interval that scores are meaningfully different.

- Tediously copy and paste all ~180 questions into Gemini, Claude, Perplexity, Perplexity Pro with GPT-4o and SearchGPT.

There's some things that aren't well understood, and are constants for 6 months now:

- Llama 3.1 8B x Search is indistinguishable from Gemini Advanced (Google's $20/month Gemini frontend)

- Perplexity baseline is absolutely horrid, Llama 3.1 8B x search kicks its ass. Perplexity Pro isn't very good. If you switch Perplexity Pro to use gpt-4o, it's slightly worse than SearchGPT.

- Regular RAG kicks everythings ass. That's the only explanation I can come up with for why Telosnex x GPT-4o beats SearchGPT and Perplexity Pro using 4o. All I'm doing is bog-standard RAG with a nice long prompt with instructions. Search results from API => render in webview => get HTML => embeddings => pick top N tokens => attach instructions and inference. I get the vibe Perplexity has especially crappy instructions and input formatting, and both are too optimized for latency over "reading" the web sites, SearchGPT more so.

Re: Genie 2: A large-scale foundation world model

#187
post #181

Earlier quoted context omitted.

That's the kicker, LLM driven stories are likely to fall into the same trap that "infinite" procedurally generated games usually do - technically having infinite content to explore doesn't necessarily mean that content is infinitely engaging. You will get bored when you start to notice the same patterns coming up over and over again. Procgen games mainly work when the procedural parts are just a foundation for hand-c…

Yeah, generative AI can create cool looking pictures and video but so far it hasn't managed to create infinitely engaging stories. The models aren't there yet.

I'd argue that the same principle applies to pictures, there are many genres of AI image that are cool the first time you see them, but after you've seen the exactly the same idea rehashed dozens of times with no substantial variety it starts wearing really thin. AI imagery is often recognizable as AI not just because of charactistic flaws like garbled text but because it's so hyper-clichéd.

Re: Genie 2: A large-scale foundation world model

#188

It’s interesting to me that we continue to see such pressure on video and world generation, despite the fact that for years now we’ve gotten games and movies that have beautiful worlds filled with lousy, limited, poorly written stories. Star Wars movies have looked phenomenal for a decade, full of bland stories we’ve all heard a thousand times. Are there any game developers working on infinite story games? I don’t ca…

If stories (and AAA games in general) are bland in games is due in large part to how expensive are to produce. Risk tolerance is low.

If game assets are cheap to generate you’ll see small teams or even solo developers willing to take more creative risks

Re: Genie 2: A large-scale foundation world model

#189

Its so much like my lucid dreams where world sometimes stays consistent for a while when I take its control. It's a strange feeling seeing computer hallucinating a world just like I hallucinate a world in dreams. This also means that my dreams will keep looking like this iteration of Genie 2, but computer will scale up and the worlds won't look anything like my dreams anymore in next versions (its already more colorf…

Soon enough I imagine we'll have dream state to cohesive reality models. Our desires and world events can be dissected and analyzed by fine grain and hint authorities to your intent before you know what they mean to you /s.

Re: Genie 2: A large-scale foundation world model

#190
post #4

This is.. super impressive. I'd like to know how large this model is. I note that the first thing they have it do is talk to agents who can control the world gen; geez - even robots get to play video games while we work. That said; I cannot find any: - architecture explanation - code - technical details - API access information Feels very DeepMind / 2015, and that's a bummer. I think the point of the "we have no moat…

Any estimates of how much one of these cost to generate and keep a minute of context?

Secondly, any estimate of how much the price could fall in 5-10 years?

Post reply on HN