Finally, my interest in LLMs is piqued! Seems like everyone has been getting excited around the search or code-generation use cases ... or simply trying to make it say naughty things (boring, not interested, wake up in a few more years), but this is eye opening. The idea of this as a "universal coupler" is fascinating, and I think I agree with the author that we are probably standing at an early-90s-web moment with L…
We are months away from being able to do this with images too. All the pieces are there, and multi-modal (smallish) large language + image models are already being used in research labs; eg MS Kosmos-1[1]. Check out the visual IQ test results in the paper. Kosmos-1 is only 1.6B parameters. When that or similar models scale out to 50B_ params they will be pretty amazing. [1] https://arxiv.org/abs/2302.14045
I am currently creating reference images for a game which I expect within 6-12 months I'll be able to feed into a multimodal ChatGPT to create 3D assets out of the 2D pics.
We'll be able to conjure up worlds at a whim - so start imagining them already!