Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

521–530 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#521

Earlier quoted context omitted.

Because agents lack human judgment. At the very least there's a need for a human-in-the-loop with agentic processes. Otherwise, it's like running a coding harness with --dangerously-skip-permissions all the time.

Why do you think judgement is impossible to automate? What aspects of it do you think make it hard?

Judgments are not generally impossible to automate -- judgements are typically binary or quantifiable interpretations, so in some sense are perfect targets for automation, but the sheer volume of judgements needed to build something coherent is overly cumbersome to specify to the point of being intractable. There are also many hidden judgements, ones where the thresholds may not be well understood, and interactions between them.

But humans still manage to wrangle these, sometimes seemingly effortlessly, through a process which we call by shorthand "taste". This is a largely vibes-based heuristic that combines expertise with life experience and cultural training -- intuition, more or less.

This is likely not possible to automate either -- aspects of it may be automatable for a given expert, in small pieces in narrow subsets of their particular domains of interest, but even those likely will require some manual intervention.

This is in part because it is, to a large degree, a black box, even to the expert deploying it. With some self-awareness and strong language skills we can articulate approximations of the judgements that go into taste. But even those will fall short, as even the most self-aware individual will fail to notice certain judgements and dependencies.

In practice many of these are not even explicitly articulable. Humans are idiosyncratic and messy and dynamic, and the suggestion that we can build a machine that approximates this in a way that pleases our sensibilities and doesn't require supervision is kind of ludicrous, even in light of recent developments.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#522
post #479

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look. I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

> if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

I'll play!

My recommendation, in short: pick a new font, make the content pane narrower (around 75 characters per line), and increase line height by 15%.

The font you use, Montserrat, has nice details, but they're extremely subtle, and our eye doesn't pick them up at text size.The font uses very pure geometry [1], has really big counters [2], is light in weight, and has no visible stroke contrast [3]. The effect when you look at the screen is that you see big blocks of text, but struggle to register individual characters, and it's hard to find the beginning of the next line. Add the lack of capitalization, and it feels like a monologue, like you're playing rubber-duck for someone.

A challenge: everybody uses Google Fonts, and all the good fonts get used so heavily they end up losing their ability to make people feel something.

Pick a font that is good [4] and feels comfortable to you. Ideally, buy one. You get what you pay for. And you go from being one of the millions of people who use a certain font to being one of fifteen.

If you like the feel of Montserrat, Tiny Grotesk [5] and Decimal [6] would both be great choices and are from excellent type designers. Or browse Matthew Butterick's font recommendations [4]. They're good. You're out $50, but it's yours.

If you're absolutely allergic to paying for fonts, DM Sans and Work Sans are Google Fonts, have a similar feel to Montserrat, and solve the above problems. But then you're a robot.

A final note: Montserrat is actually a nice headline font. The details come into focus at larger sizes and weights. So I'm going to open a can of worms: you could pick a _contrasting_ body font. Remember PT Serif from the [3] footnote? Try it on for size.

----

[1] By that, I mean things like: an "o" looks like a perfect circle rather than an oblique oval, and forms are strictly on 90º axes.

[2] Counters are the apertures of letters, the open center of an "o" being one.

[3] This is the contrast in "line" size. Think of it like the width of the line when you write with a chisel-tip marker: some lines end up thinner, some thicker. Look at the "e" on the font PT Serif - the vertical walls of the character are thicker, horizontal thinner. https://fonts.google.com/specimen/PT+Serif?categoryFilters=S...

[4] You can't go wrong with anything Matthew Butterick recommends (https://practicaltypography.com/font-recommendations.html), or from any of the "big" font foundries: Hoefler & Co., Linotype, Monotype, Berthold, URW. If it's a bestseller on myfonts, it's a sure choice: https://www.myfonts.com/collections/best-seller

[5] https://tinytype.co/type/tiny-grotesk

[6] https://www.myfonts.com/collections/decimal-font-hoefler-and...

Re: Why I'm still bearish on LLMs after Navier-Stokes

#523

Earlier quoted context omitted.

Why do you think judgement is impossible to automate? What aspects of it do you think make it hard?

It requires general intelligence and we don't even have a good understanding of how our's works or a particularly good way of quantifying it. The counter argument is of course maybe you don't need to understand our kind of intelligence to create a different kind and that could well be true but then how do you determine if a system is intelligent. Unless the new system is intelligent enough to reason with us on our le…

Why do you think it requires general intelligence? The parent and I aren't being obtuse here: the history of artificial intelligence research is littered with examples of humans confidently declaring that task X requires general intelligence, then getting humiliated by a neural network doing task X better than humans a few years later. See Go, driving, art (you can complain about the quality of AI art, but it's winning competitions with human judges), etc.

A priori I'm not sure why you would think being a PM at a FAANG, deciding what color the login button should be, is any different.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#524

Earlier quoted context omitted.

That would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.

If the goal is merely to "win at chess", then yes, an LLM using stockfish is better than any human alone at performing the task. When you are talking about what AI agents are capable of doing, there is no such thing as "cheating". They are as capable as the tools they can use effectively. The entire history of human civilization was driven by effectively using tools to achieve goals.

But then the human should also get Stockfish, and we're back at a stalemate.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#525

Earlier quoted context omitted.

Threw me off too. Like why???

It's a tech bro thing. Altman does it too, and I've worked with people in the past who do it. I read it as "I'll take literally any conscience for myself no matter how minor, at any cost for you no matter how big".

I actually think of it as a tumblr thing. Like my first thought whenI see it is that the person who wrote this runs in queer or activist circles.

Not my read on OP, but “tech bro” is not my first thought associated with the style

Re: Why I'm still bearish on LLMs after Navier-Stokes

#526

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

> 2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc.

The problem with this angle is that it is still absolutely terrible at doing call center/customer service work, and the profitability story is that the price is going to go up rather than go down.

For repetitive physical labor in a controlled environment I'm slightly more bullish, but if you control the environment, you mostly don't need AI. You just use traditional deterministic methods, and send a person in when things get stuck or things are by nature irregular.

> 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc.

Those who can accept failure cheaply can't necessarily detect failure cheaply. A ton of insane attempts will have to be picked through carefully to find the candidates for success, because the lack of a thought process makes AI bad in random, inhuman ways. This is basically a version of 3) that wishes away tests. It will be (and is) certainly helpful to replace interns and aid in rapid prototyping, but not because failure can be accepted, but because those are things that are tightly supervised. According to the world thus far, that is resulting in anything from -15% to +25% productivity gains. I'm not seeing it as a game changer simply because if it was, I'd expect to have seen a lot more useful, original software products by now and I haven't. I've just seen old ones get buggier or rewritten in Rust.

I'm only buying 3): when you just want a machine to randomly enumerate through a search space looking for things that make the carefully constructed tests pass. That's a very good thing, though. But as you say, it's not a game changer because you still have to write the tests.

but

> 5. navier-stokes and statements in pure mathematics like it are the absolute best case scenario for agentic work against rigorous specification. the theorem statement itself is already a rigorous specification. it has undergone decades of auditing by the mathematical community and its rendering in lean is a straightforward translation defined in terms of battle-tested mathematical objects from mathlib. the verifier, the lean theorem prover, has been extensively audited and specifically designed to avoid the types of unsoundness that would make it vulnerable to reward hacks. even lean and theorem provers like it are not invulnerable: soundness bugs have allowed LLMs to launder bogus proofs through the proof kernel before and it is not improbable that more such bugs exist. this is the rosiest setup; the vast majority of human knowledge work does not look like this. i'll comment below on the few areas of knowledge work that do resemble pure mathematics in this respect.

This is the real deep point, and one I've been repeating since I heard Navier-Stokes was a fraud.

This is exactly where I expected that LLMs would do well, and they are not.

It shows that I have a basic misunderstanding of LLMs, and that misunderstanding is causing me to think that they have more potential than they have actually shown.

Maybe the nature of the architecture, where it picks out features, intrinsically limits its ability to search a solution space?

Maybe the fact that they modally predict what someone might say, and nobody has said a thing as of yet (when many people were knowledgeable enough to have, if it is correct), means that the LLM is not going to say it either?

Maybe the fact that it consumes all information and blends it in a structured way, instead of synthesizing an entire space from a relatively very small amount of input like a human does, means that it won't ever accidentally synthesize something that can't be pieced together from things that have already been said? Is its accuracy its flaw, where a human's "mistaken" synthesis might ultimately correct everyone's understanding?

Really not beating the charge of being a stochastic parrot. It might just be that we were underestimating stochastic parrots; if a million monkeys on a million typewriters were all getting treats when they satisfied a trainer who wanted to see a new work of Shakespeare; they could look at his published work, and they could watch each other type and when each other got treats; whenever they successfully spelled a word or put words into an intelligible phrase, that was made into a keyboard key for a group of sentence monkeys, and the successes of the sentence monkeys were made into keys for the paragraph monkeys, etc... could you get something that passed for mediocre, drunken Shakespeare in a thousand years? Or maybe even 10?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#527
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

I'm just dropping this all over this thread but you're unfortunately mistaken https://dynomight.net/more-chess/

More show and less tell would be appreciated.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#528

Earlier quoted context omitted.

Doesn't really even need to look like it. If you can verify rewards, RLVR will optimize really really well. If you can't... it's a struggle. There are probably fewer fields where you can verify rewards than one might hope.

> There are probably fewer fields where you can verify rewards than one might hope. 2 tasks I've done today that I believe robots are nowhere near being able to do: Cleaning my wardrobe and draining bad fuel out of my generator. As in generic use cases.

Hard to verify that your wardrobe is clean. Also hard to verify that the bad fuel is out without physical sensors. Many, many tasks are quite difficult to verify beyond "you know it when you see it". That doesn't work so well for training a model.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#529

This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for…

What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those func…

I am bad in chess game by itself like 1200 ELO, but I can write Programm and win player with 2600 ELO. Does it mean I am pro chess gamer?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#530

Earlier quoted context omitted.

> an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain. AGI != ASI. You are confusing the two.

I'm not. AGI is almost necessarily closer to ASI than it is to human intelligence by definition. it's become pretty obvious that even what seems like irrelevant domain knowledge has utility applied to other domains - that's why we're pursuing general-use models presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task bec…

> presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion.

I disagree with this definition of AGI, and I disagree that chess skills significantly benefit from generalizing non-chess knowledge, outside of computing moves probabilistically.

AGI has historically been defined as human level or better, with generality to new domains. I think blurring it with ASI makes the terminology confusing to use.

Chess is learned rules and the ability to apply those rules. Strategy as a whole is applying a set of rules to circumstances, that's how it is taught: "here are examples of circumstances and actions, try to pattern match to future circumstance and apply commensurate action."

If you make the point that chess is a large part of the training data, or that LLMs are unable to learn chess well, I'll accept that as refuting that LLMs are AGI, but these other points I disagree with.

Post reply on HN