Earlier quoted context omitted.
Bitter lesson strikes again!
_Especially_ given the goal of a world model using a rasters-only frame-by-frame approach. Holy shit.
It makes me think that Stargate might actually lead to AGI/ASI
491–500 of 512 posts
Earlier quoted context omitted.
Bitter lesson strikes again!
_Especially_ given the goal of a world model using a rasters-only frame-by-frame approach. Holy shit.
It makes me think that Stargate might actually lead to AGI/ASI
Earlier quoted context omitted.
Is it actually unbelievable? It's basically what every major AI lab head is saying from the start. It's the peanut gallery that keeps saying they are lying to get funding.
Even as a layman and AI skeptic, to me this entirely matches my expectations, and something like this seemed like it was basically inevitable as of the first demos of video rendering responding to user input (a year ago? maybe?). Not to detract from what has been done here in any way, but it all seems entirely consistent with the types of progress we have seen. It's also no surprise to me that it's from Google, who I…
Hard to fault them as the process towards ASI now appears to be runaway and uncontrollable.
Earlier quoted context omitted.
Probably depends on how you engage with GTA. “Drive on the street simulator” along with arrays of weapons and explosions is the majority of my hours in GTA. I despise the creative and artistic vision of GTA online, but I’m clearly in a minority there gauging by how much money they’ve made off it.
I took the "creative and artistic vision" line to refer to the story mode.
I didn't think the story was earth-shattering; it was fine, but no Baldur's Gate.
Edit: In retrospect, the characters were fairly iconic. I still distinctly remember Trevor.
I wonder how hard it would be to get VR output? That's an insane product right there just waiting to happen. Too bad Google sleeps so hard on the tech they create.
Consistent output and spatial coherence across each eye, maybe a couple years? But meeting head tracking accuracy and latency requirements, I’d bet decades. There’s no way any of this tech reduces end to end latency to acceptable levels, without a massive change in hardware. We’ll probably see someone use reprojection techniques in a year or so and claim they’ve done it. But true generated pixels straight to the head…
This model already runs at 24fps, and I bet could be made to run at >75fps by scaling hardware and distilling/quantizing the model to only work on certain environments.
The two eye problem seems pretty trivial to me: add another image decoding head with the sole task of decoding the other eye. Training data for this can be plentifully gathered through simulated 3D data, or running existing 2D data (e.g. youtube videos) through slow mono to stereo models. This should add minimal latency as it's another head vs. subsequent layers.
If you can train the model to allow WASD movement + mouse, head tracking is not very different. I think with enough effort we could probably build a VR experience using this today. Getting it onto affordable hardware could be a totally different story, but certainly not decades.
Maybe I'm missing something though!
Would progress in these be faster if they created 3d meshes and animations instead of full frame videos?
Earlier quoted context omitted.
Your brain does not need to render any environments, just the experience of being in them.
What do you think the difference is?