Earlier quoted context omitted.
Yesterday my dad, in his late 70's, used Gemini with a video stream to program the thermostat. He then called me to tell me this, rather then call me to come stop by and program the thermostat. You can call this hype, maybe it is all hype until LLMs can work on 10M LOC codebases, but recognize that LLMs are a shift that is totally incomparable to any previous AI advancement.
> He then called me to tell me this, rather then call me to come stop by and program the thermostat. Sounds like AI robbed you of an opportunity to spend some time with your Dad, to me
John Carmack talk at Upper Bound 2025
311–320 of 387 posts
Re: John Carmack talk at Upper Bound 2025
#312Earlier quoted context omitted.
Can you give an example how it would be transformative compared to specialized AI?
Because it could very well exceed our capabilities beyond our wildest imaginations. Because we evolved to get where we are, humans have all sorts of messy behaviours that aren't really compatible with a utopian society. Theft, violence, crime, greed - it's all completely unnecessary and yet most of us can't bring ourselves to solve these problems. And plenty are happy to live apathetically while billionaires become t…
Re: John Carmack talk at Upper Bound 2025
#313What Carmack is doing is right. More people need to get away from training their models just with words. AI need the physicality.
> More people need to get away from training their models just with words. They started doing that a couple of years ago. The frontier "language" models are natively multimodal, trained on audio, text, video, images. That is all in the same model, not separate models stitched together. The inputs are tokenized and mapped into a shared embedding space. Gemini, GPT-4o, Grok 3, Claude 3, Llama 4. These are all multimoda…
Are the audio/video/images tokenized the same way as text and then fed in as a stream? Or is the training objective different than "predict next token"?
If the former, do you think there are limitations to "stream of tokens"? Or is that essentially how humans work? (Like I think of our input as many-dimensional. But maybe it is compressed to a stream of tokens in part of our perception layer.)
Re: John Carmack talk at Upper Bound 2025
#314Interesting reply from an openai insider: https://x.com/unixpickle/status/1925795730150527191
I think some replies here are reading the full twitter thread, while others (not logged in?) see only the first tweet. The first tweet alone does come off as a dismissal with no insight.
Re: John Carmack talk at Upper Bound 2025
#315Earlier quoted context omitted.
My bet is on Carmack.
Appeal to authority is a logical fallacy. People often fall into the trap of thinking that because they are highly intelligent and an expert in one domain that this makes them an expert in one or more other domains. You see this all the time.
While this is certainly true, I'm not aware of any evidence that Carmack thinks this way about himself. I think he's been successful enough that's he's personally 'post-economic' and is choosing to spend his time working on unsolved hard problems he thinks are extremely interesting and potentially tractable. In fact, he's actively sought out domain experts to work with him and accelerate his learning.
Re: John Carmack talk at Upper Bound 2025
#316Earlier quoted context omitted.
> More people need to get away from training their models just with words. They started doing that a couple of years ago. The frontier "language" models are natively multimodal, trained on audio, text, video, images. That is all in the same model, not separate models stitched together. The inputs are tokenized and mapped into a shared embedding space. Gemini, GPT-4o, Grok 3, Claude 3, Llama 4. These are all multimoda…
(If you know) how does that work? Are the audio/video/images tokenized the same way as text and then fed in as a stream? Or is the training objective different than "predict next token"? If the former, do you think there are limitations to "stream of tokens"? Or is that essentially how humans work? (Like I think of our input as many-dimensional. But maybe it is compressed to a stream of tokens in part of our percepti…
Re: John Carmack talk at Upper Bound 2025
#317Earlier quoted context omitted.
I'm sure there were offline rendering and 3D graphics workstation people saying the same about the comparatively crude work he was doing in the early 90s... Obviously both Carmack and the rest of the world has changed since then, but it seems to me his main strength has always been in doing more with less (early id/Oculus, AA). When he's working in bigger orgs and/or with more established tech his output seems to suf…
The thing about Carmack in the 90s... There was a lot of research going on around 3d graphics. Companies like SGI and Pixar were building specialized workstations for doing vector operations for 3d rendering. 3d was a thing. Game consoles with specialized 3d hardware would launch in 1994 with the Sega Saturn and the Sony Playstation (in Japan only for one year) What Carmack did was basically get a 3d game running on…
Re: John Carmack talk at Upper Bound 2025
#318Earlier quoted context omitted.
> If DOOM released in 1994 or 1995, would we still remember it in the same way? Maybe. One aspect of Wolfenstein and Doom's popularity is that it was years ahead of everyone else technically on PC hardware. The other aspect is that they were genre defining titles that set the standards for gameplay design. I think Doom Deathmatch would have caught on in 1995, as there really were very few (just Command and Conquer?)…
I guess the thing about rapid change is... it's hard to imagine what kind of games would exist in a DOOMless world in an alternate 1995. The first 3d console games started to come out that year, like Rayman. Star Wars Dark Forces with its own custom 3d engine also came out. Of course Dark Forces was, however, an overt clone of DOOM. It's a bit ironic, but I think the gameplay innovation of DOOM tends to hold up more…
Rayman was a 2D game.
Re: John Carmack talk at Upper Bound 2025
#319Earlier quoted context omitted.
I'm sure there were offline rendering and 3D graphics workstation people saying the same about the comparatively crude work he was doing in the early 90s... Obviously both Carmack and the rest of the world has changed since then, but it seems to me his main strength has always been in doing more with less (early id/Oculus, AA). When he's working in bigger orgs and/or with more established tech his output seems to suf…
> his main strength has always been in doing more with less Carmack builds his kingdom and then runs it well. I makes me wonder how he would fare as an unknown Jr. developer with managers telling him "that's a neat idea, but for now we just need you to implement these Figma designs".
> "makes me wonder how he would fare as an unknown Jr. developer with managers telling him (...)"
he would probably write an open letter and left Meta. /sRe: John Carmack talk at Upper Bound 2025
#320Earlier quoted context omitted.
The thing about Carmack in the 90s... There was a lot of research going on around 3d graphics. Companies like SGI and Pixar were building specialized workstations for doing vector operations for 3d rendering. 3d was a thing. Game consoles with specialized 3d hardware would launch in 1994 with the Sega Saturn and the Sony Playstation (in Japan only for one year) What Carmack did was basically get a 3d game running on…
The world seems to have rewritten history, and forgotten Ultima Underworld, which shipped prior to Doom..