This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…
It all depends on what prompt you use though. You can just tell all current frontier models to write a chess engine first, and then play a game of chess against you using that engine. It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun). On that note, I actually had an overall harness (for experimenting) that was essentially like this:…
Why I'm still bearish on LLMs after Navier-Stokes
441–450 of 646 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#442Earlier quoted context omitted.
Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better? I would not be surprised if OpenAI released a model that beats humans at chess this year.
Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the conce…
The regular model generally does not suffer the same issues he is demonstrating with the real time audio version.
In my view the investment into datacenters is well justified by the current demand, and progress has been very impressive.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#443Earlier quoted context omitted.
Why not ask it to implement a chess engine first, and then use that to play against you? Does the LLM need to learn to play chess if it can build a chess engine to play for it instead?
With that approach, the benchmark falls apart. Of course it can write a chess engine, because it learned on lots of stolen source code of chess engines. This has nothing to do with the LLM's ability to reason. Writing a well understood engine for a super popular problem does not count as reasoning about the problem.
Doesn't writing the engine imply understanding about the problem domain? Tool use is a widely accepted measure of intelligence.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#444This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…
What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those func…
The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#445Earlier quoted context omitted.
It's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperform on that test. Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have…
> The fact that the human would have a much harder time writing a useful program is irrelevant. Why? There's a box. You give it a problem, and it comes up with a solution. Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes? Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it ca…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#446Earlier quoted context omitted.
Because they’ll train it to be good at chess and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even ____ It can’t even count the R’s in strawberry It can’t even add numbers It can’t even solve a millennium puzzle It’s not even a chess GM It’s not even beyond human capability in Go It can’t even drive a car It can’t even self replicate It can’t even build weapons It doesn’t even have…
> and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though. Notably, if you have to retrain the model in order to X then it can't possibly be AGI since if it were _general_ it would be capable o…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#447Earlier quoted context omitted.
Well except AlphaZero played 44 million chess games in that time (and actually played with a 44 core computer). So I'd like to point out that the human is still just a few orders of magnitude more efficient.
Yes, we all know that biological systems are more efficient than machines through billions of years of evolution and natural selection but the overall process is largely the same (interacting with an environment, learning from results, improving underlying architecture, etc); efficiencies will come with more time and improvements.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#448This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…
What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those func…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#449Earlier quoted context omitted.
I would definitely take you up on that.
https://www.chessbench.org/ >GPT-6 Astra xHigh: 0.06% rejected moves
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#450Earlier quoted context omitted.
> The "value" of most knowledge workers -- based on what enterprises currently pay for them -- is $50 - 70 trillion annually. What do you mean? The sum of ALL US salaries is $13.4 Trillion per year. According to google $65T is the sum of ALL salaries Globally (not just knowledge workers). It's not reasonable to assume AI is a drop-in-replacement for any job yet (perhaps bottom tier customer support from oversees?). >…
You are comparing company valuations to annualized revenue (as approximated by some fraction of total knowledge worker compensation). Valuations are (roughly) based on the sum of all discounted future cash flows, not just the current year’s revenue.
There's no strong evidence that openAI or anthropic will have non-linear revenue growth, so I'm not sure what point you're trying to make. Unless you think they'll fire all their engineers and replace them with agents or something crazy?