Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

331–340 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#331

Earlier quoted context omitted.

State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behav…

I'm not sure why anyone is expecting stochastic systems to be deterministic. Chess is a deterministic game won by a combination of known movesets and constrained multi-level forward search. LLMs do neither of these things. They don't reproduce training data exactly, their next response is more 'inspired by' prompts and its own memory than produced deterministically, and they don't have the capability to do general fo…

I'm expecting that they at least don't forget about pieces between turns, we're in AGI era after all, according to the tech overlords.

I, as a human AGI, would jever just forget and remove a piece from the board from one turn to the next.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#332

Earlier quoted context omitted.

If humans were actually intelligent, they wouldn't need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?

OpenAI making the next model good at chess is not analogous to a human training to get good at chess. It is analogous to God creating Human 2.0 which now has increased chess playing ability. If LLMs were intelligent the way humans are, then the models that exist right now would be able to spend time improving themselves at chess and become good at it. They can't do this because they are not, in fact, intelligent.

[dead]

Re: Why I'm still bearish on LLMs after Navier-Stokes

#333

Earlier quoted context omitted.

And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That's just not how the brain works. If LLMs could do that they would truly be superintelligence.

> Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. Maybe not, but you'd be surprised how little it takes. A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I've witnessed significant improv…

It is interesting, but people are drawing the wrong conclusion from it. For one, LLMs don't go through a "chess learning phase". They're not analyzing a board as they're learning the rules or studying games to create a coherent model of chess. They're just imbibing raw relationships as isolated fragments of information. The fact that they can't unify this into a coherent model of chess playing in one shot and execute a competent game says nothing interesting about the limits of their intelligence. If you give frontier models the rules of chess in their context window, could they perform only legal moves? I bet they could, excepting trickier scenarios like pins and failing to respond to a check. But those kinds of scenarios have to be reinforced in any human player as well. Even Super GMs fall for mate-in-1's occasionally which is functionally equivalent to those kinds of failures.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#334

Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT…

For me prototyping whas always about going in trying to surface the unknown unknowns.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#335
post #93

Earlier quoted context omitted.

I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well. People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

Then why respond at all for the sake of responding? We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own. You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess. I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the ot…

> We all know AI can code...

We know no such thing. LLMs are quite bad at generating code, worse than any capable human.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#336
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access. I don't believe this. You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up L…

[dead]

Re: Why I'm still bearish on LLMs after Navier-Stokes

#337

Earlier quoted context omitted.

State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behav…

For me the useful intuition is that LLMs haven't somehow magickally learned to implement any of the algorithms we know that we have used to make strong chess engines: alpha-beta minimax and Monte-Carlo Tree Search on the one hand, and obviously the ability to learn accurate evaluation functions by self-play. I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to…

But it speaks in words, therefore it must be super duper extra smart!!11 /s

Sarcasm aside, I think this is an easy cognitive trap to fall into. It does sometimes feel like the LLM must have some world model because it converses somewhat coherently. Examples like this failure to understand chess, or to count the number of Rs in "strawberry", seem difficult to explain if the models are intelligent. But that doesn't stop people believing they are anyway. I think there must be something about the conversational interface that fools us easily. I wonder if people trained in interrogation techniques are also fooled?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#338

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

"Transatlantic flight will never be commercially viable, we conclude based on careful study of several aircraft designs from the 1920s"

How about supersonic flight?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#339

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

The fact that they can play chess at all despite having no specific training for it blows my mind, and the fact it doesn’t do the same for many others shows just how far they’ve come and how fast.

It doesn't blow anyone's mind because it hasn't been impressive for a computer to play chess for 40 years. "We made something worse than existing solutions by using a new technique" is not an impressive feat.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#340

Earlier quoted context omitted.

So we humans are not a general intelligence then? And the stuff i'm using LLMs daily is just fake? I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

> So we humans are not a general intelligence then? No, because we can , in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us. > And the stuff i'm using LLMs daily is just fake? This misses the point completely.

> generally read the rules of a game and then follow them

How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.

Post reply on HN