Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

261–270 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#261
AI is just good at what it's got the most elaborate training data on. And by "good" I mean, is statistically most likely to spit something out that makes some kind of sense.

I don't know how well it is studied, but I suspect it is possible there is language-related complexity constraints to the effectiveness of the LLM algorithms. Like perhaps context-free grammar related problems with adequate training data can be more and more effectively solved, but maybe natural language related problems will not so much be effectively solved.

Would be curious if there is active research here.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#262

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

So give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.

Well I guess this excuse is finally gonna become obsolete soon with all the "pacing" nonsense.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#263

I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures. Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon ag…

My belief is that LLMs will fundamentally change how we approach domain expertise. From what I've seen, SDEs tend to be over-specialized compared to what the company actually needs to implement due to the need to understand enough of the domain to pick a best path. If an LLM can see the domain enough so that someone in an adjacent field can be confident in their approach and quickly change course then you don't need as many niche SDEs

Re: Why I'm still bearish on LLMs after Navier-Stokes

#264

Earlier quoted context omitted.

so prove it! get a public repo out there, have it play against some open source engines also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?

I suck at chess. Are you saying I can't be intelligent?

The issue with the models isn't that they play a bad game, but that they persist in making illegal moves. An average intelligent human can be told the rules of chess and then play chess, badly, within the rules.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#265

Earlier quoted context omitted.

> Why should that matter? Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games. So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.

We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules. Why is it so hard for people to keep track of the thread of discussion?

> We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules.

Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.

The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.

This does not point to generalisable and adaptable intelligence, such as we see in the average human.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#267

Earlier quoted context omitted.

> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

Thats just absolutly not true. A human being has general intelligence and needs A LOT of training and finetuning to become good in chess. And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

Humans don't need a lot of training and finite tuning to make only legal moves.

An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.

An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).

Re: Why I'm still bearish on LLMs after Navier-Stokes

#268
post #73

The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for th…

There's a leverage issue.

In one case (financial services) it's thought that expertise is valuable at V=S^2/b4 where V is value, S is skill and b capacity (the leverage available to the manager/expert. b erodes as it becomes harder to find examples of things that are not done well, so if you manage $1bn you might find lots of miss allocations that you can exploit with just that $1bn really effectively, but if you manage $10bn it's much harder to find good places for the extra $9bn. A low hanging fruit effect.

Anyway, that double hit - raw skill and the amount of times you can supply the skill makes the value of skill (V) convex, and it means that in a perfect market (heh heh heh) someone running $100bn is worth 1000's or maybe 10,000's of an average joe expert.

Now, if AI is trusted to run the top 0.1% of everything and has the skill to do it at human top level expertise, then your calc holds. If it's the case that it isn't then more than half of that value disappears. If it's not even top 1% then chop out another 25%.

That implies that we need a lot of trust and a lot of AI capability before these valuations stack up, and it also implies that all other competitors and incumbants are going away. I do not think that Citidal or Bridgewater are going to let Anthropic or OAI take them without a fight. They might lose - but there is a decent bet that they don't. I don't think that many professions like Lawyers or Doctors are just going to roll over and cede their monopoly rights to OAI or Anthropic either.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#269

The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? IMO the very best case scenario / potential for th…

Wasn't Navier Stokes solved by ripping off a researcher's private chat log?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#270

Earlier quoted context omitted.

I don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are claiming. The l…

Gary Marcus is an especially puzzling addition. If I recall correctly, he has made statements along the lines that superintelligence this century is more likely than not. If you’re AGI-pilled that might read as bearish, but that is still extremely rapid progress in the grand scheme of things.

What even is superintelligence? Is my phone not a superintelligence?
Post reply on HN