Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

291–300 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#291

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for…

> The actual current frontier plays somewhere around GM level.... It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors

Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#292

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

" the lack of sentence capitalization makes it unnecessarily difficult to read." If it had proper caps etc, people here would accuse it of written using LLMs. You just can't win...

> You just can’t win…

That’s life )

Re: Why I'm still bearish on LLMs after Navier-Stokes

#293

Earlier quoted context omitted.

So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.

There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1. The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier. That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.

Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)...

I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.

[0] https://news.ycombinator.com/item?id=49720751

Re: Why I'm still bearish on LLMs after Navier-Stokes

#294
post #286
post #276

Earlier quoted context omitted.

How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?

I don't know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.

I've not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#295

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

"Transatlantic flight will never be commercially viable, we conclude based on careful study of several aircraft designs from the 1920s"

Re: Why I'm still bearish on LLMs after Navier-Stokes

#296
post #78
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…

> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.

This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#297
post #84
post #19

Short and to the point! Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.

Open models wont be open for long. No one is going to release an open model capable of chaining zero-days. Even the Chinese aren't that reckless because it will just be turned around and used against them.

They already aren't really open, try asking an open model on advice for constructing a nuclear bomb. There's no available model that's even remotely near the frontier that doesn't have restrictive safeguards built in.

(Mind you, this may be for the better. I'm just saying that the safeguards driven by cybersecurity concerns aren't some new quality that wasn't there before.)

Re: Why I'm still bearish on LLMs after Navier-Stokes

#298

Earlier quoted context omitted.

> The only way you get better at chess is by playing a lot of games and learning from mistakes How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice. This kickstarting then gradual refinement is how most people learn.…

Reading can kickstart the process, but you can also make random moves guided by some sort of system (such as a computer GUI) or learn by watching other players play. The overall point is that you learn through observation and lots of trial and error (whether you are a human or a computer). And beginners in chess often make illegal moves even after learning the rules, it's fairly common. It feels like you're trying to…

> The overall point is that you learn through observation and lots of trial and error

That’s the most inefficient way and people usually avoid doing that. Instead they find someone that knows how to do the thing and ask him to be a teacher. Or use a proxy like a book or videos.

> It feels like you're trying to say that humans never make illegal moves while learning chess, which doesn't match with my experience. I'm trying to understand your overall point

There’s learning the basic stuff (which is done after a few games) and there’s mastery. The thread started with the observation that even with all that knowledge (through content ingested in training), LLMs still makes illegal moves. Humans can be erratic, but they can constrain themselves to the rules for the task at hand after learning them.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#299

Earlier quoted context omitted.

If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this. So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

So we humans are not a general intelligence then? And the stuff i'm using LLMs daily is just fake? I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

> So we humans are not a general intelligence then?

No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us.

> And the stuff i'm using LLMs daily is just fake?

This misses the point completely.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#300
post #286
post #276

Earlier quoted context omitted.

How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?

I don't know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.

More importantly, beginner human players don't exhibit that tendency. The history of the position doesn't bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.
Post reply on HN