Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

491–500 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#491

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law ( https://en.wikipedia.org/wiki/Parkinson%27s_Law ) is still in effect. You could even expand "Work expands so as to fill the time avail…

>with "time and tokens available".

That's like 1/4 of Codex users burning through their free resets just to build shinier todo apps.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#492

Earlier quoted context omitted.

Why does a 6 year old not need any of these guardrails? Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about. Maybe not all AGIs have a path to digital singularity. Maybe ou…

Ask a 6 year old to draw a chess board from scratch every turn and they too will make mistakes.

No one has spent the past three years telling me that a 6 year old will take my job!!!

Re: Why I'm still bearish on LLMs after Navier-Stokes

#493

Earlier quoted context omitted.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…

Luna is one of the budget lower-end last generation models. It'd be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it's Claude Fable 5.1 with XHigh or Max reasoning.

out of curiosity, do you think fable would get this right? (I'm not sure myself, and haven't tried yet.)

Re: Why I'm still bearish on LLMs after Navier-Stokes

#494

Earlier quoted context omitted.

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law ( https://en.wikipedia.org/wiki/Parkinson%27s_Law ) is still in effect. You could even expand "Work expands so as to fill the time avail…

Why do you belive that agents wouldn't be able to take over product management, and generate prompts for the "software engineer" agents?

Because agents lack human judgment. At the very least there's a need for a human-in-the-loop with agentic processes. Otherwise, it's like running a coding harness with --dangerously-skip-permissions all the time.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#495

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Threw me off too. Like why???

Reverse signaling

Re: Why I'm still bearish on LLMs after Navier-Stokes

#496

Earlier quoted context omitted.

chess is not just a rigorous set of rules, it is rules as foundation with layers of strategy on top. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun that comprises human knowledge knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from…

> an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain. AGI != ASI. You are confusing the two.

I'm not. AGI is almost necessarily closer to ASI than it is to human intelligence by definition. it's become pretty obvious that even what seems like irrelevant domain knowledge has utility applied to other domains - that's why we're pursuing general-use models

presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion. it's like the parable of Newton and the apple - the domain knowledge that an apple falls according to certain rules observed through historic experience igniting the creative spark that led to universal gravitation

Re: Why I'm still bearish on LLMs after Navier-Stokes

#497

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

As a programming lead on a hobbyist video game project, we're reaping massive rewards being in category #1. I keep the architecture and important details in check while letting frontier models go ham. As alluded, it's a video game, not a life support system, so bugs are low-impact. But even better, the defect rate is actually the lowest it's ever been. Insidious bugs baked in by years of accumulated human error are trivial for Sol or Astra to untangle. This is the best time to be alive so far if you enjoy hobby game dev.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#499

Earlier quoted context omitted.

Why do you belive that agents wouldn't be able to take over product management, and generate prompts for the "software engineer" agents?

Because agents lack human judgment. At the very least there's a need for a human-in-the-loop with agentic processes. Otherwise, it's like running a coding harness with --dangerously-skip-permissions all the time.

Why do you think judgement is impossible to automate? What aspects of it do you think make it hard?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#500

Earlier quoted context omitted.

> a topic with an unusual level of specification Solving cancer also has an unusual level of specification. Many real world problems have that characteristic.

> Solving cancer also has an unusual level of specification. Where did you read that? "Cancer" is not just a single disease, even though we layman use the term that way. Cancer is a family of diseases, each probably having their own specific solution, but even in each of these individual diseases, there is no specification at the level of any open maths problem.

I know what cancer is. Many scientists believe that all human adults above the age of about 30 have "cancer".

I really meant solving a specific cancer disease for a specific individual, which will require individualized medicine, which will require us to leverage AI to make it possible to do for the general population as opposed to doing it just for Lance Armstrong and the like.

I know I'm an optimist, but I really think AI is gonna result in drastically improved healthcare for a much lower price.

Post reply on HN