Earlier quoted context omitted.
> current frontier models need laborious oversight and guardrails on even the simplest tasks It is literally denialist about current capabilities
why don't anthropic and openai ship yolo mode by default?
Why I'm still bearish on LLMs after Navier-Stokes
271–280 of 642 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#272"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, " No, they're really not. They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be. And that they will capture most of that ... which they won't. The Frontier Labs ar…
ai has a >10% chance of causing human extinction, according to anthropic big heads. if that's true, you are wrong. if that's false, anthropic is dishonest. why trust a dishonest company to be worth anything?
This logic doesn't follow at all.
If their argument is that there is 10% chance of extinction then they also believe there is a 90% chance it won't.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#273Lots of people make this claim about "specific task[s] enjoying clearly defined levels of task performance" but they forget that generative AI is also extremely good at generating a) art and b) prose in literary style. None of those things has "clearly defined levels of task performance", in fact they are both the complete opposite of well-defined tasks. Who knows what counts for "good" art? [1]
For me the right model for generative AI is "a million monkeys on typewriters" [2]. Holding any other model to heart will at some point fail to predict observations and cause you to be unpleasantly surprised. Not least because AI companies are actively engineering their systems to optimise for this model and they have a lot of people working on that engineering and shedloads of money to throw at it.
Don't underestimate what a million monkeys on typewriters can do. They can do anything and everything, given enough time. Geneartive AI can also do anything and everything given enough resources. The only question is: how much is going to be "enough"?
____________________
[1] Yes yes, AI art tends to be slop. Not denying that. But part of the problem with slop is that it presents as technically very competent except that it lacks a certain je-ne-sais-quoi, which makes it good art; aesthetics. The point is that there is no clear measure of what makes technically competent art, any more than there is for aesthetics.
And yet generative AI is very good at it.
[2] There's even an article on wikipedia except it's about one monkey on one typewriter with infinite time. There's a proof too.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#274Re: Why I'm still bearish on LLMs after Navier-Stokes
#275I think that's incorrect from the investment point of view. They'd still be worth a lot if they can produce a drop-in replacement but it takes five or ten years as long as they dominate that. The danger from an investment point of view is they become AltaVista, replaced by some Google that does the job better.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#276Earlier quoted context omitted.
Thats just absolutly not true. A human being has general intelligence and needs A LOT of training and finetuning to become good in chess. And there is a relevant and significant difference between the expectation of an AGI and an ASI system.
Humans don't need a lot of training and finite tuning to make only legal moves. An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves. An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of c…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#277Earlier quoted context omitted.
The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for…
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…
This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#278Earlier quoted context omitted.
Write the same sentence you just wrote back to me, but in only four words and let’s see if it has the same meaning.
"Any reason why that can't be solved through context management and keep-forward scaffolding?" becomes "load bearing context seam" /s
Dabadooba, ba dabadooba!
Re: Why I'm still bearish on LLMs after Navier-Stokes
#279Earlier quoted context omitted.
> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
I tested both myself and a weak bot against Astra xhigh, https://lichess.org/study/27lCQqDa . It's still pretty bad at chess, though it takes longer to devolve into illegal moves.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#280Earlier quoted context omitted.
that is if it even a human commenter at all
State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behav…
I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to do that in the process of learning to represent the entire text on the web. On the other hand they have gotten say very good at machine translation without being trained exclusively (and I select the preceding word carefully) on machine translation.
Edit: I'm saying this because there is this idea expressed by e.g. Ilya Sutskever, that in order to predict the next token accurately an LLM has to learn something about all of underlying reality. See for example this interview with Dwarkesh:
https://x.com/biobootloader/status/1640512444958396416
Where Sutskever claims that "Predicting the next token well means you understand the underlying reality that led to the creation of that token".
If that were true, we should have seen LLMs play good chess by now. There is a huge amount of data on playing chess floating around on the web in the form of algebraic chess notation and if LLMs were capable of learning the "underlying reality" of chess, they would already have. They haven't. Because they can't. What Sutskever is saying flies in the face of literally hundreds of years of statistical modelling, which is to say, building predictive models that, very explicitly, do not have to understand any "underlying reality" and only have to be good at modelling a dataset.