Live data from Hacker News

Marble: A Multimodal World Model

worldlabs.ai

61–70 of 87 posts

Re: Marble: A Multimodal World Model

#61
post #41

Earlier quoted context omitted.

histrionic and meretricious

those two words only describe AI models, as they are models. A "world model" is worse than those two words as it is oxymoronic. The idea that words and space are being conflated as a formula for spatial intelligence is fundamentally absurd as our relationships to space have no resolution, both within any one language and worse, between them, as language is arbitrary. Language and thought are entirely separate forms.…

Is that a more convoluted way to say that a next thing predictor can't exhibit complex behavior? Aka the stochastic parrot argument. Or that one modality can't be a good enough proxy for the other. If so, you probably have to pay more attention to the interpretability research.

But actually most people should start with strong definitions. Consciousness, intelligence, and other adjacent terms have never been defined rigorously enough, even if a ton of philosophers think otherwise. These discussions always dance around ill-defined terms.

Re: Marble: A Multimodal World Model

#62

Earlier quoted context omitted.

those two words only describe AI models, as they are models. A "world model" is worse than those two words as it is oxymoronic. The idea that words and space are being conflated as a formula for spatial intelligence is fundamentally absurd as our relationships to space have no resolution, both within any one language and worse, between them, as language is arbitrary. Language and thought are entirely separate forms.…

Is that a more convoluted way to say that a next thing predictor can't exhibit complex behavior? Aka the stochastic parrot argument. Or that one modality can't be a good enough proxy for the other. If so, you probably have to pay more attention to the interpretability research. But actually most people should start with strong definitions. Consciousness, intelligence, and other adjacent terms have never been defined…

Neurobio is built from the base units of consciousness outwards, not intuited interpretation. Eg prediction has nothing inherent to do with consciousness directly. That’s a process imposed on brains post hoc.

https://pubmed.ncbi.nlm.nih.gov/38579270/

And

https://mitpress.mit.edu/9780262552820/the-spontaneous-brain...

Easily refute prediction or error prediction as fundamental.

The path to intelligence or consciousness isn’t mimicry of interpretation.

In terms of strong definitions, start at the base, coders: oscillation, dynamics, Topologies, sharp wave ripples, and I would say roughly 60 more strongly defined material units and processes. This reverse intuition is going nowhere and it’s pseudoscientific nonsense for social media timeline filling.

Re: Marble: A Multimodal World Model

#63
post #14
post #7

I understand that DeepMind is working on this too: https://deepmind.google/blog/genie-3-a-new-frontier-for-worl... I wonder how their approaches and results compare?

Genie delivers on-the-fly generated video that responds to user inputs in real time. Marble renders a static Gaussian Splat asset (like a 3D game engine asset) that you then render in a game engine. Marble seems useful for lots of use cases - 3D design, online games, etc. You pay the GPU cost to render once, then you can reuse it. Genie seems revolutionary but expensive af to render and deliver to end users. You neve…

Graphics have long reached diminishing returns in gameplay, people aren't going to playing VRChat tomorrow for the same reasons today.

AI can speed up asset development, but that is not a major bottleneck for video games, what matters is the creative game design and backend systems, which existing on the interaction between players and systems is just about as hard as any management role, if not harder.

Re: Marble: A Multimodal World Model

#64
post #4

I'm floored. Incredible work. also check out their interactive examples on the webapp. It's a bit more rough around the edges but shows real user input/output. Arguably such examples could be pushed further to better quality output. e.g. https://marble.worldlabs.ai/world/b75af78a-b040-4415-9f42-6d... e.g. https://marble.worldlabs.ai/world/cbd8d6fb-4511-4d2c-a941-f4...

Unsurprisingly the results are by far the best in the area shown in the image in the prompt, and quickly deteriorate beyond it, or more than a couple meters behind the camera.

It's worlds better than just doing gaussian splats from images, but given how much the quality is influenced by images the limit to four images with text prompt or eight images without prompt is quite limiting. That's plenty to describe a chair, but almost nothing to describe a home or a space station. I hope they can extend those limits in future updates

Re: Marble: A Multimodal World Model

#65
post #8

As someone with barebones understanding of "world models," how does this differ from sophisticated game engines that generate three-dimensional worlds? Is it simply the adaptation of transformer architecture in generating the 3-D world v/s using a static/predictable script as in game engines (learned dynamics vs deterministic simulation mimicking 'generation')? Would love an explanation from SMEs.

Games are still mostly polygon based due to tooling (Even Unreal Nanite is a special variation of handling polygons), some engines have tried voxels (Teardown, Minecraft genererates polygons and would fall in the previous category as far as rendering goes) or even implict surface modes by composing SDF'y primitives (Dreams on Playstation and more recently unbound.io). All of these have fairly "exact" representations,…

What does the Gaussian approach do that resolves the issue with voxel engines? I recall if you wanted to start doing animation it becomes a mess of computational complexity.

Re: Marble: A Multimodal World Model

#66

Is Marble's definition of a "world model" the same as Yann LeCun's definition of a world model? And is that the same as Genie's definition of a world model?

Very different, it would seem. Then again, it’s never been clear to me why LeCun believes that LLM architectures don’t inherently produce world models in the course of training.

Re: Marble: A Multimodal World Model

#67

This is going to take movie making to another level because now we can: 1. Generate a full scene 2. Generate a character with specific movements. Combine these 2, and we can have moving cameras as needed (long takes). This is going to make storytelling very expressive. Incredible times! Here's a bet: We'll have a AI superstar (A-list level) in the next 12 months.

Hard disagree. CG in films is awful when done cheaply, and this all looks like really cheap CG.

Re: Marble: A Multimodal World Model

#68

Earlier quoted context omitted.

Is that a more convoluted way to say that a next thing predictor can't exhibit complex behavior? Aka the stochastic parrot argument. Or that one modality can't be a good enough proxy for the other. If so, you probably have to pay more attention to the interpretability research. But actually most people should start with strong definitions. Consciousness, intelligence, and other adjacent terms have never been defined…

Neurobio is built from the base units of consciousness outwards, not intuited interpretation. Eg prediction has nothing inherent to do with consciousness directly. That’s a process imposed on brains post hoc. https://pubmed.ncbi.nlm.nih.gov/38579270/ And https://mitpress.mit.edu/9780262552820/the-spontaneous-brain... Easily refute prediction or error prediction as fundamental. The path to intelligence or consciousnes…

I started writing the counterargument, but somehow I think you have a weird idea of what both interpretability in ML and neurobiology are, especially seeing how you're dealing with things nobody has a full idea about in such absolutes

Re: Marble: A Multimodal World Model

#69

Earlier quoted context omitted.

Fair warning, when I last put up a bet in AI video arena, I won! https://www.linkedin.com/posts/anilgulecha_kitsune-activity-... Same terms - gentlemen's agreement. The loser owes the winner a meal whenever they meet :). For a HN visitor to blore, I'll happy to host a meal anyway :)

What was the 90 minute movie?

The LinkedIn thread also seems as AI generated there

Re: Marble: A Multimodal World Model

#70

Incredibly disappointing release, especially for a company with so much talent and capital. Looking at the worlds generated here https://marble.worldlabs.ai/ it looks a lot more like they are just doing image generation for a multiview stereo 360 panoramas and then reprojecting that into space. The generations exhibit all the same image artifacts that come from this type of scanning/reconstruction work, all the same…

Yeah, I'm likewise a bit underwhelmed by the results.

If you go in with the expectation that you give it a single image and it's doing gaussian splatting from a single image and a prompt it's phenomenal. If you deviate too far from the image viewpoint it breaks down, but it looks decent long enough to be very usable. But if you go in with the expectation that it's generating "worlds" it's not very good. This only passes as a world in a 20 second tech demo where the user isn't given camera controls

My best guess is that they are forced (by investors, lack of investors, fear of the AI bubble, or whatever) to release something, and this was something they could polish up to production quality and host with reasonable GPU resources

Post reply on HN