Earlier quoted context omitted.
So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.
I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well. People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
Why I'm still bearish on LLMs after Navier-Stokes
231–240 of 642 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#232the issue most of you seem to not realize is that when you put these models in a loop, you are able to do more and more insane and cool things. have you guys actually designed, built, and deployed agentic workflows? it is actually quite hard, requires tons of time spent on evals and testing to ensure accuracy, but when it starts to work it is mind blowing. there is no going back. listening to people yap about AI when…
Please share some of these insane things that you speak of..
Re: Why I'm still bearish on LLMs after Navier-Stokes
#233Earlier quoted context omitted.
If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn). Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a…
That's quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
Re: Why I'm still bearish on LLMs after Navier-Stokes
#234This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…
An LLM is the wrong approach for playing chess.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#235Earlier quoted context omitted.
I can write a chess bot program that will beat you. Does that mean I’m good at chess? >If they cared to have it perform well in chess games, you'd see a different shape and behavior. So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?
They’ll never be AGI simply because the definition will be constantly updated to be some steps ahead of them.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#236Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT…
Putting that aside, prototype software is recombining existing technologies and concepts in well trodden domains, which is distinct from the genuinely novel scientific work the author was contrasting with. Software prototypes are not in the same league, as much as you may like it to be.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#237This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…
> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#238Earlier quoted context omitted.
The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for…
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…
I don't believe this.
You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.
A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.
I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#239Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT…
I don't understand this comment. We use AI for prototyping all the time. We write ERP software. Client wants to know how process XYZ could be automated? Send them a prototype UI (sans the three A's) so they can play around with it. Takes 30 minutes.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#240Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT…
The aspect of prototype software that the author was calling out was that it is throwaway software. Putting that aside, prototype software is recombining existing technologies and concepts in well trodden domains, which is distinct from the genuinely novel scientific work the author was contrasting with. Software prototypes are not in the same league, as much as you may like it to be.
That being said I didn't compare both, not sure why you brought that up. I specifically discussed about prototyping, quoting a specific sentence, not scientific research.