Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

471–480 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#471

I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures. Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon ag…

> This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures. Any reason why that can't be solved through context management and keep-forward scaffolding?

And who’s to manage context? And who’s building the scaffolding? Yes, AI can be used for both, but you do realize that all this being self-contained and regulated internally is what makes biological agents successful agents, right? If you break the process apart and need to dial back in these aspects, and can only do so with human input, or another agent which will need the same handholding the one whose issues you’re solving for, where’s the agency?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#472
post #64

Earlier quoted context omitted.

The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for…

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

I'm just dropping this all over this thread but you're unfortunately mistaken

https://dynomight.net/more-chess/

Re: Why I'm still bearish on LLMs after Navier-Stokes

#473

Earlier quoted context omitted.

My point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code. The illegal move aspect has more to do with a failure of online/in-context learning, which would support y…

chess is not just a rigorous set of rules, it is rules as foundation with layers of strategy on top. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun that comprises human knowledge knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from…

> an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain.

AGI != ASI. You are confusing the two.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#474

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

Yes "oversight and guardrails" are still needed, but even that is becoming easier to build, and imo already no longer in the "laborious" category. Even far from the frontier, you can tune a 0.2B LLM into a decent 2000 Elo player as a weekend project:

https://x.com/maximelabonne/status/2100137121264828901

Re: Why I'm still bearish on LLMs after Navier-Stokes

#475

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

The story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.

there are far more reams of chess moves than there are math papers. Lichess is pretty open...

But hey, they're actually good at chess if you prompt correctly so.... https://dynomight.net/more-chess/

Re: Why I'm still bearish on LLMs after Navier-Stokes

#476

Earlier quoted context omitted.

It's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperform on that test. Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have…

>> There are many claims that current LLMs surpass humans in cognitive abilities Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?

Why the aggressive tone and the strawman rhetoric? I never said there was a consensus. Yes I'm talking more about the "optimistic" commenters and pointing out that this chess thing is a good datum to temper their enthusiasm. What's wrong with that?

Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#477

Earlier quoted context omitted.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…

What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill? Also what harness? If you’re using a general harness of course it’s going to try and give you commentary. I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when…

You're pointing out that the goalposts are not fixed in the problem statement above, and gp's interpretation is not the most generous possible. But as the interpretations get more generous, the claim becomes more and more absurd. Maybe a properly-harnessed model would download the most advanced chess engine and query it to find the best move in each position, but that's not really demonstrating the model's intelligence anymore.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#478

Earlier quoted context omitted.

Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better? I would not be surprised if OpenAI released a model that beats humans at chess this year.

Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the conce…

I find it amusing that you're describing a huge misallocation of capital and a society enabling such, and that is the optimisitic scenario (in my mind anyway).

Re: Why I'm still bearish on LLMs after Navier-Stokes

#479

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look.

I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

Re: Why I'm still bearish on LLMs after Navier-Stokes

#480

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…

Luna is one of the budget lower-end last generation models. It'd be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it's Claude Fable 5.1 with XHigh or Max reasoning.
Post reply on HN