Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

551–560 of 646 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#551

Earlier quoted context omitted.

Why do you think judgement is impossible to automate? What aspects of it do you think make it hard?

Judgments are not generally impossible to automate -- judgements are typically binary or quantifiable interpretations, so in some sense are perfect targets for automation, but the sheer volume of judgements needed to build something coherent is overly cumbersome to specify to the point of being intractable. There are also many hidden judgements, ones where the thresholds may not be well understood, and interactions b…

You don't need to describe it; just show samples of the style you want to achieve. Of course it's not perfect, but it's easier than describing it.

I have my own theory about why it's impossible to remove the human from the loop:

1. Any task emerges from a need, from a human context. We need the human to pay and assume the risks and costs of using the model. So intent emerges from context.

2. While the task is being worked on, constant interaction with the context is needed, for action, for feedback, and for steering.

3. At the end of a task, consequences accumulate in the context, they don't fly to the model provider. The cost, risk, liability, gains and losses remain there.

So the LLM is great except for the start, middle and end of a task. Contexts are humans, teams, projects, and they are maximally distributed, you can't copy a context, it is indexical and relational, just as you can't copy my phone number or eat for me.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#552
post #535

Earlier quoted context omitted.

Why do you think it requires general intelligence? The parent and I aren't being obtuse here: the history of artificial intelligence research is littered with examples of humans confidently declaring that task X requires general intelligence, then getting humiliated by a neural network doing task X better than humans a few years later. See Go, driving, art (you can complain about the quality of AI art, but it's winni…

Hence my saying. We will never prove AI is intelligence. We'll only prove humans are not.

I mean I could believe that there are some tasks which can't be done at acceptable performance until AI resembles something like Commander Data, I just don't think being a software PM is one of them.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#553

Earlier quoted context omitted.

It would be more impressive if they could play chess (or do anything they haven't been custom RLVR trained for) by reasoning, rather than just "have a go at it" prediction which is closer to memorization. HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?

They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board? The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.

> They can't possibly remember even a few positions.

Sure they could, but that's irrelevant.

A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.

However, that's not how LLMs work. They don't memorize inputs - they predict them, based on discovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positions). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.

> Don't you know the legend about rice grains on a chess board?

Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.

> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.

Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#554

Earlier quoted context omitted.

With that approach, the benchmark falls apart. Of course it can write a chess engine, because it learned on lots of stolen source code of chess engines. This has nothing to do with the LLM's ability to reason. Writing a well understood engine for a super popular problem does not count as reasoning about the problem.

> Writing a well understood engine for a super popular problem does not count as reasoning about the problem. Doesn't writing the engine imply understanding about the problem domain? Tool use is a widely accepted measure of intelligence.

Definitely not. A human can write a chess engine without being very good at chess.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#555

Earlier quoted context omitted.

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law ( https://en.wikipedia.org/wiki/Parkinson%27s_Law ) is still in effect. You could even expand "Work expands so as to fill the time avail…

I think the fear is that softwaring engineering becomes a low-skill profession. If the models get good enough you won't need years of experience to be a decent programmer.

IMO the fear is that software engineering becomes a _high_-skill profession.

If the models get good enough to solve all of the low-skill problems, then we only need to keep around the people who are highly skilled. That means a fewer number of software engineers, and a difficult path to becoming someone who is highly skilled.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#556
post #479

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look. I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

Someone has already mentioned Matthew Butterick's fonts and his online "book" Practical Typography [1]—absolutely fantastic material there.

Some concrete suggestions from me:

- Fira Sans [2] is a lovely, legible sans-serif font that's free and incredibly easy-to-read. This would make a fine body font.

- Source Sans [3] is another good sans-serif that might be closer to your current font. It's about as readable as Fira, but it's a little less warm imo.

- Make your line length a little narrower and your line spacing (i.e. leading) just a hair wider—that will do wonders for readability.

Definitely check out Practical Typography for more tips from a real professional.

Thanks for the blog post!

[1]: https://practicaltypography.com/

[2]: https://carrois.com/fira/

[3]: https://github.com/adobe-fonts/source-sans

Re: Why I'm still bearish on LLMs after Navier-Stokes

#557
post #522
post #479

Earlier quoted context omitted.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look. I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

> if anyone has font or other recommendations for style tweaks to make I'd love to hear them! I'll play! My recommendation, in short: pick a new font, make the content pane narrower (around 75 characters per line), and increase line height by 15%. The font you use, Montserrat, has nice details, but they're extremely subtle, and our eye doesn't pick them up at text size.The font uses very pure geometry [1], has really…

Ah, fellow Butterick fonts enjoyer—Butterick's fonts have spoiled me: compared to pretty much every other professional font I've seen, they're the cheapest and most conveniently licensed fonts out there.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#558

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

I want you to consider how relevant this is in any practical sense. First -- most _people_ cannot do this, without having a physical board in front of them. Second -- Claude Code is perfectly capable of downloading and running stockfish. People focus too much on LLMs by themselves as the entity of concern instead of the entire harness and all of it's capabilities together.

Because they are obviously testing for general intelligence. If you want a thread about how cool the harness is, that's down the street.

We don't really consider humans downloading stockfish to beat people at chess as noteworthy endeavors.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#559

Earlier quoted context omitted.

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law ( https://en.wikipedia.org/wiki/Parkinson%27s_Law ) is still in effect. You could even expand "Work expands so as to fill the time avail…

I think the fear is that softwaring engineering becomes a low-skill profession. If the models get good enough you won't need years of experience to be a decent programmer.

Expectations around software quality will rise in tandem with gains in productivity so that the same experience curve continues to apply.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#560

Earlier quoted context omitted.

What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those func…

I am bad in chess game by itself like 1200 ELO, but I can write Programm and win player with 2600 ELO. Does it mean I am pro chess gamer?

Do I care if Richard Feynman was only able to do nuclear physics with the help of an abacus?

Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.

Sometimes you care about the Bee, sometimes you care about the results.

Post reply on HN