Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

211–220 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#211

The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? IMO the very best case scenario / potential for th…

The whole point of this post was that it's questionable what can be achieved without huge investments into oversight and steering, because navier-stokes was a topic with an unusual level of specification. The problem itself was a specification. Such situations are rare in real-world scenarios.

AI agents are good at solving well-specified tasks, not at solving problems. They do well in fields where the cost/effort of specification is already part of the business.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#212

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

The fact that they can play chess at all despite having no specific training for it blows my mind, and the fact it doesn’t do the same for many others shows just how far they’ve come and how fast.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#213
post #96

Earlier quoted context omitted.

Humans don’t code a $game engine to play $game, they can just play it. It seems like you are the one that has gone insane.

And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That's just not how the brain works. If LLMs could do that they would truly be superintelligence.

> And how many years of direct play and study does it take for a human to get good at chess or any other game?

Time is irrelevant to training; the more relevant comparison is "how many games does a human need to play to get diminishing returns".

Re: Why I'm still bearish on LLMs after Navier-Stokes

#214

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

Llm systems are not really build for adhering to a grammar (other than "a string og tokens"). It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents. Certainly,a harness can easily correct for it.

By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it's a somewhat educated random walk in the space of correct lean theorems.

Harnesses do correct things, sure.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#215

Earlier quoted context omitted.

Llm systems are not really build for adhering to a grammar (other than "a string og tokens"). It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents. Certainly,a harness can easily correct for it.

By the promise of it, llms should be able to both adhere to grammars, or go free form where necessary. I mean, doing math is supposed to be strict but in practice it's a somewhat educated random walk in the space of correct lean theorems. Harnesses do correct things, sure.

You are right. I am imprecise.

Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.

Games and programming languages (including lean) does not allow this flexibility.

A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.

Certainly it must be like that, otherwise reviews in math was rendered moot.

Do we blame research mathematicians for not adhering to the grammar?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#216
post #63

Earlier quoted context omitted.

> People who are good at it rely more on experience and deep domain expertise People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range. A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is q…

1100 at online speed chess or something, could be. I'm not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.

1100 is literally below the ELO you get by default as a beginner.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#217

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

    Array.from(document.body.querySelectorAll('p,li')).filter(e=>e.innerText).map(e=>e.innerText = e.innerText.split('\. ').map(s=>s[0].toUpperCase() + s.slice(1)).join('. ') )
Not perfect but hope it helps.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#218
post #105
post #74

Earlier quoted context omitted.

> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.

This is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.

I feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#219

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

Llm systems are not really build for adhering to a grammar (other than "a string og tokens"). It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents. Certainly,a harness can easily correct for it.

It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#220
post #84
post #19

Short and to the point! Open and cheap models will undercut the big labs continuously. The blast radius won't be pretty once spending commitments knock the door.

Open models wont be open for long. No one is going to release an open model capable of chaining zero-days. Even the Chinese aren't that reckless because it will just be turned around and used against them.

Isn't that assuming that fix won't be implemented?

Zero days are valuable because they can be exploited but if the pace of exploitation is faster (which I'm not sure is the case), then the response WILL be faster, even if it means going offline. Institutions that won't will simply go offline by losing their data or becoming unprofitable due to ransomware.

Now for components that are core to the infrastructure, say OpenSSL, there is already a TON of attention and efforts, including red teaming, so it's not as if it's opening floodgates.

Sure low hanging fruits will get picked either faster or a at a larger scale, say a random outdated IoT device at your local flower shop, but for the rest, I don't think it's realistic to expect no response.

Security, digital or not, has always been an arm race. New threats means new responses specifically by incorporating the threat.

Post reply on HN