So is the translation endless scaling has stopped being as effective?
It's stopped being cost-effective. Another order of magnitude of data centers? Not happening. The business question is, what if AI works about as well as it does now for the next decade or so? No worse, maybe a little better in spots. What does the industry look like? NVidia and TSMC are telling us that price/performance isn't improving through at least 2030. Hardware is not going to save us in the near term. Major i…
Ilya Sutskever: We're moving from the age of scaling to the age of research
81–90 of 374 posts
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#82Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#83The impactful innovations in AI these days aren't really from scaling models to be larger. It's more concrete to show higher benchmark scores, and this implies higher intelligence, but this higher intelligence doesn't necessarily translate to all users feeling like the model has significantly improved for their use case. Models sometimes still struggle with simple questions like counting letters in a word, and most p…
> this implies higher intelligence Models aren't intelligent, the intelligence is latent in the text (etc) that the model ingests. There is no concrete definition of intelligence, only that humans have it (in varying degrees). The best you can really state is that a model extracts/reveals/harnesses more intelligence from its training data.
Note that if this is true (and it is!) all the other statements about intelligence and where it is and isn’t found in the post (and elsewhere) are meaningless.
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#84> These models somehow just generalize dramatically worse than people. The whole mess surrounding Grok's ridiculous overestimation of Elon's abilities in comparison to other world stars, did not so much show Grok's sycophancy or bias towards Elon, as it showed that Grok fundamentally cannot compare (generalize) or has a deeper understanding of what the generated text is about. Calling for more research and less scali…
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#85He is, of course, incentivised to say that.
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#86Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#87The impactful innovations in AI these days aren't really from scaling models to be larger. It's more concrete to show higher benchmark scores, and this implies higher intelligence, but this higher intelligence doesn't necessarily translate to all users feeling like the model has significantly improved for their use case. Models sometimes still struggle with simple questions like counting letters in a word, and most p…
> this implies higher intelligence Models aren't intelligent, the intelligence is latent in the text (etc) that the model ingests. There is no concrete definition of intelligence, only that humans have it (in varying degrees). The best you can really state is that a model extracts/reveals/harnesses more intelligence from its training data.
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#88Scaling is not over, there's no wall. Oriol Vinyals VP of Gemini research https://x.com/OriolVinyalsML/status/1990854455802343680?t=oC...
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#89> These models somehow just generalize dramatically worse than people. The whole mess surrounding Grok's ridiculous overestimation of Elon's abilities in comparison to other world stars, did not so much show Grok's sycophancy or bias towards Elon, as it showed that Grok fundamentally cannot compare (generalize) or has a deeper understanding of what the generated text is about. Calling for more research and less scali…
Re: Ilya Sutskever: We're moving from the age of scaling to the age of research
#90Earlier quoted context omitted.
Fridman is a morally broken grifter, who just built a persona and a brand on proven lies, claiming an association with MIT that was de facto non-existent. Not wanting to give the guy recognition is not a matter of being liberal or conservative, but just interested in truthfulness.
Patel takes anticommunism to such an extreme that he repeatedly brings up and speculates (despite being met with repudiation by even the staunchest anticommunist of guests) whether naziism is preferable, that Hitler should have the war against Soviets, that the US should have collaborated with Hitler to defeat communism, and that the enduring spread of naziism would have been a good tradeoff to make.
It being the first (and so far only) interview of his I'd seen, between that and the AI boosterism, I was left thinking he was just some overblown hack. Is this a blind spot for him so that he's sometimes worth listening to on other topics? Or is he in fact an overblown hack?