Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

111–120 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#111

> current frontier models need laborious oversight and guardrails on even the simplest task As models advance, we shift the goalpost for what "simplest task" means. Before, "simplest task " meant "write a coherent English sentence." Now, "simplest task" means autonomously fix, review, and merge a bugfix.

Eliza wrote coherent English sentences.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#112
post #107

Earlier quoted context omitted.

If you gave a human a book or two on chess they would not become a decent player (they would be closer to 500-600 than 1100 ELO) and they would only get better after playing hundreds or thousands of games (often making illegal moves and moves that violate the rules of chess as they learn). Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a…

> considering LLMs currently play better than a brand new human player would They’ve ingested all the literature on playing chess, a brand new human player has not.

Yes, but my point is that humans can’t even do the thing that the above comments are claiming humans can do (read a book or two and be decent at chess), and then they complain that LLMs can’t do the same thing (that humans can’t do either).

We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#113
post #97
post #78

Earlier quoted context omitted.

The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…

I can write a chess bot program that will beat you. Does that mean I’m good at chess? >If they cared to have it perform well in chess games, you'd see a different shape and behavior. So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?

They’ll never be AGI simply because the definition will be constantly updated to be some steps ahead of them.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#114
"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, "

No, they're really not.

They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be.

And that they will capture most of that ... which they won't.

The Frontier Labs are a very bad buy at a high price, but that partly has to do with wacky pricing, but actually mostly has to do with their relatively weak place in the value chain.

The money is going to Nvidia, who have the most powerful position.

A bit like how a retailer can take all the margins of some innovative product, if they own the channel.

AI is over-hyped, the Frontier Labs are over priced - but AI is here to stay, and will grow. Not like Skynet, but like a new form of compute. And it will take it's time, and the profits will be reaped by those with the power.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#115
> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers

That's a reason to be bearish about AI companies, not LLMs. But is it even true? OpenAI and Anthropic have each reported ~50 billion in revenue with ~900 billion valuations. That's a high ratio but I'm not sure if follows that the only way it pans out is if we get "fully automated drop-in replacement for most knowledge workers".

It wouldn't shock me to see those revenue numbers scaling up to where they need to be over the next decade ( to, say, ~400 billion) without ever achieving drop-in worker replacements.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#116
post #78

Earlier quoted context omitted.

The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…

It's not quite the same, but the in-flight chess game provided by Delta was known to be absurdly hard: https://news.ycombinator.com/item?id=46593395

I believe I remember reading it was based on Glaurung's code (which eventually evolved into what we now know as the juggernaut Stockfish).

Re: Why I'm still bearish on LLMs after Navier-Stokes

#117

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks. 3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and C…

> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.

A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess."

Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#118

Earlier quoted context omitted.

1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks. 3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and C…

so prove it! get a public repo out there, have it play against some open source engines also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?

I suck at chess. Are you saying I can't be intelligent?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#119

Earlier quoted context omitted.

If humans were actually intelligent, they wouldn't need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?

Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck

If all you do is read chess books, you'll be a shit player. Training and practice is what it takes to be great.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#120
post #81

Earlier quoted context omitted.

So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.

AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

>LLM can’t beat an avg chess player.

Why should that matter?

Post reply on HN