Earlier quoted context omitted.
I've not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.
I haven't tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: https://news.ycombinator.com/item?id=49720751 Nothing is forcing the LLM to play 'blind'. If it's smart, it should be able to create its own representation of the chess board and update it with every move, just like a human would. Any chess engine that's sensitive to how the moves are formatted is cle…
Why I'm still bearish on LLMs after Navier-Stokes
391–400 of 642 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#392Earlier quoted context omitted.
1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks. 3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and C…
> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly. Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.
If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker?
Probably the only thing saving many jobs from being replaced right now is that it’s hard to have a verification of correctness in the loop, so the agent can’t hill climb very easily.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#393Re: Why I'm still bearish on LLMs after Navier-Stokes
#394Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
didn't realize people struggled to read text that way though, maybe i should change my writing style if this is a common pain point
Re: Why I'm still bearish on LLMs after Navier-Stokes
#395Earlier quoted context omitted.
AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
The fact that LLMs can play chess at any level is a strong indication we are in AGI.
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
Re: Why I'm still bearish on LLMs after Navier-Stokes
#396Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
" the lack of sentence capitalization makes it unnecessarily difficult to read." If it had proper caps etc, people here would accuse it of written using LLMs. You just can't win...
Re: Why I'm still bearish on LLMs after Navier-Stokes
#397The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for th…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#398Earlier quoted context omitted.
Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better? I would not be surprised if OpenAI released a model that beats humans at chess this year.
I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish. Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models ha…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#399The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for th…
> It's reasonable to assume that if AI drop-in-replaced all those knowledge workers, AI companies could credibly charge somewhere in that order of magnitude, because that's what the market is already bearing This assumes you don't change the market, but at the scale of (checks notes...) "all knowledge work", that just doesn't hold. For example if you put 1bn people out of work, you now need some sort of safety net to…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#400Earlier quoted context omitted.
That would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.
My points is more that, while we have a strong intuition about where, as an entity, a human's boundaries are (i.e. where the person begins and ends), philosophically it' not immediately obvious that the analogy applies to the a model in the same way. Why should that be the line drawn that says this is the "thing" and this other stuff is external to the thing? It feels somewhat arbitrary. Of course this is a difficult…