Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

301–310 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#301
One the main theses is

> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on

Unless frontier labs have surprisingly trained their models in the exact tasks my team works on, this is patently false. We are getting very good results on automation and I'm bullish we will be able to mostly remove humans in the loop for most of our infra tasks by the end of the year.

I have no opinion on the other theses, but given that OP doesn't back up these claims in any way, I have my doubts about the conclusions of this article.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#302
post #159

Earlier quoted context omitted.

the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions

That has not been my experience with new players, they regularly make invalid or incorrect moves even after detailed instructions especially in novel situations.

Maybe it depends on the person? My six year old isn’t great at strategy but they can pretty consistently make valid moves. Sometimes they ask for confirmation on a move which is also not a trait I see in language models (at least unprompted)

Re: Why I'm still bearish on LLMs after Navier-Stokes

#303

Earlier quoted context omitted.

AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

The fact that LLMs can play chess at any level is a strong indication we are in AGI.

This is roughly comparable to observing a cat batting a ball away with its paw and taking this as a "strong indication" that cats can play any sport.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#304

Earlier quoted context omitted.

The fact that LLMs can play chess at any level is a strong indication we are in AGI.

This is roughly comparable to observing a cat batting a ball away with its paw and taking this as a "strong indication" that cats can play any sport.

Yes, a good analogy. Except the cat actually follows the football rules and can beat some humans. And has no physical limitations to play other kinds of sport that you might imply.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#305

Earlier quoted context omitted.

> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

Pretty much this. Feed it a book or two on chess, and you should have a decent (or good) player. That's the generic intelligence people have. The aims is not to be supremely talented at something, but being able to read a manual and figure how to use/play something. Mastery can be gained overtime.

1. The LLMs have surely ingested hundreds if not thousands of books on chess.

2. The study (along with other posters here) show the models can’t even stick to following the rules of the game

Re: Why I'm still bearish on LLMs after Navier-Stokes

#307

Earlier quoted context omitted.

The fact that LLMs can play chess at any level is a strong indication we are in AGI.

No it isn't. Computers could play chess long before LLMs, better than LLMs can in fact. That didn't make them AGI.

You are saying "No it is not" without an argument. The fact that computer systems could play chess yet not being AGI has no relevance to LLMs' ability to play chess being AGI, because the point is about G, not I. There's little doubt about A or I parts.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#308
post #234

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

I suspect (in a probably ignorant fashion) that this is because learning process has been reading a lot of algebraic chess notation (such as "1. e4 e5 2. Nf3 f6 3. Nxf6 gxf6 4. Qh5! +-") then, to play, generating more of it without considering the rules of the game. This is exactly how it's always felt to me when playing chess against LLMs. Sure, "1. e4 e5 2. Nf3 Nc3" looks innocent to somebody simply learning the sy…

Are you saying that modern LLMs cannot play chess now, or that LLMs (GPT architecture) cannot be trained to play chess well?

Or are you saying that neural networks in general cannot (practically) be trained to be an above-average chess player?

Or are you saying that it depends on the input? Would it be better if they were given a picture/drawing/ascii art of the board? If so, surely they can produce it at will?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#309
post #170

Earlier quoted context omitted.

I think that the AI people believe in their own nonsense a bit. Like - the guy on TV talking about 'AI will destroy everything' ... I don't think he's lying. I think they are like we here on HN and Reddit and a bit caught up in our own thoughts. If AI were unleashed, in raw form today, it could cause havoc. Bad. Maybe very bad but I think we'd get over it. It would probably trigger a recession (because we are in a bu…

>If AI were unleashed, in raw form today, it could cause havoc. What is "raw form?"

My exact question and afaict since I’m running models locally and can inspect and retrain etc..I assume I already have the raw form?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#310
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

>> I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.

How many posts if I link that do the same thing will you agree this is the norm here.

Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine.

I don't even know if there is critical thought or we believe what we read/shared/etc

Post reply on HN