Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

381–390 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#381

Earlier quoted context omitted.

AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

>LLM can’t beat an avg chess player. Why should that matter?

It depends on what you are selling it as.

It only matters if you are claiming it to be general purpose.

If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.

The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#382

Earlier quoted context omitted.

How many posts if I link that do the same thing will you agree this is the norm here. Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine. I don't even know if there is critical thought or we believe what we read/shared/etc

No, I don't agree it's the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that's because many are software engineers (or equivalent) and they are used to being expected to "wing it" on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge. And there's plenty of pushback…

It's not been my personal experience on this website in the last 2-3 years atleast, pre-covid perhaps.

But despite that you aren't wrong and the only reason I even visit this website is because people sometimes did/do take time to reflect on things based on their experience and knowledge.

And in hindsight pointing out that hn has issues wasn't even the point but I feel frustrated when everyone is readily agreeing to things on here without reading. When that in this moment feels like the one thing that separates humans from machines that we get to think and learn.

I possibly should just drop reading this place until we have most noisy people go away. I have for one tried to always only comment on things where I could be a value add, this one does feel like I could I have done better.

In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.

Either way I still do think HN as a whole has devolved into mindless herd follower mindset, I can point to more than a few posts that just say adopt the hacker mindset aka move fast don't care about the consequences.

And I for one find this laughable even though that's the reality of my job/work as well.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#383
post #229

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Yep, I had the same thought. As simple heuristic, text written in all lowercase is often just hot takes and so not worth taking the time to read. LC;DR :P

I have a similar heuristic for short HN comments.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#384

Earlier quoted context omitted.

1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks. 3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and C…

Good science, properly digested and presented takes time. The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues

Not sure how that vague truism applies to this paper.

Lots of papers have great results that don’t depend on the latest models.

However in this case it’s problematic:

- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.

- They ask are LLMs capable of X and arrive at a negative result.

If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.

However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#385
post #367

Earlier quoted context omitted.

Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better? I would not be surprised if OpenAI released a model that beats humans at chess this year.

I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish. Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models ha…

I'm wondering if instructing it to track the board state in a file would make a significant difference then.

It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#386

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

What happens when you ask those same frontier models to write a chess-playing program?

I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn't fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.

It's really neat to see what a frontier model can do itself. No doubt.

But "play chess by hand" is a frankly awful metric. It's kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#387

Earlier quoted context omitted.

It's a tech bro thing. Altman does it too, and I've worked with people in the past who do it. I read it as "I'll take literally any conscience for myself no matter how minor, at any cost for you no matter how big".

why take the least charitable possible reading D; i've always read and intended it as inviting informality. i also don't find it harder to read at all (most people don't know some find it harder to read: i didn't)

I didn't take it any particular way besides maybe a hint of informality. Regardless it definitely threw off my reading in an unpleasant way. Anecdata point for the poster maybe. I would read again.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#388

Earlier quoted context omitted.

I'm pretty sure "competent at chess without external aids" has been on the standard AGI checklist since before personal computers were a thing. How can you claim an intelligence is general if it can't make sense of such a highly constrained board game? This is solidly table stakes.

Because they’ll train it to be good at chess and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even ____ It can’t even count the R’s in strawberry It can’t even add numbers It can’t even solve a millennium puzzle It’s not even a chess GM It’s not even beyond human capability in Go It can’t even drive a car It can’t even self replicate It can’t even build weapons It doesn’t even have…

> and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even

One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though.

Notably, if you have to retrain the model in order to X then it can't possibly be AGI since if it were _general_ it would be capable of figuring X out on its own having never seen it before.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#389

The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? IMO the very best case scenario / potential for th…

> The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. That's the thing, it very much does NOT show us that. What happened was mathematicians at openAI learned of an imminent development on this problem, and the insight that it entailed, then they were able to prompt a system in the correct direction and spend 20 million dollars to write down the final…

>What happened was mathematicians at openAI learned of an imminent development on this problem, and the insight that it entailed, then they were able to prompt a system in the correct direction and spend 20 million dollars to write down the final steps.

That's not what happened.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#390

Tired of that trope, I already wrote it before but "Those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc." is not correct. I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do. Rapid prototyping is NOT…

For me prototyping whas always about going in trying to surface the unknown unknowns.

Agreed, that what makes it endlessly exciting.
Post reply on HN