Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

621–630 of 646 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#621

Earlier quoted context omitted.

> So we humans are not a general intelligence then? No, because we can , in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us. > And the stuff i'm using LLMs daily is just fake? This misses the point completely.

> generally read the rules of a game and then follow them How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.

Chess.com has to accommodate people who haven't learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they're often outlining sequences of multiple "pre-moves" during the opponent's turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#622

Earlier quoted context omitted.

Luna is one of the budget lower-end last generation models. It'd be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it's Claude Fable 5.1 with XHigh or Max reasoning.

out of curiosity, do you think fable would get this right? (I'm not sure myself, and haven't tried yet.)

Elsewhere in the thread there are reports of Astra on xhigh playing at what I would characterize broadly as a competent casual level, at least given occasional prodding (which a human of that skill level would basically only require when trying to play unreasonably quickly). There seems to be a pattern (even after correcting for relative ELO systems that aren't calibrated) of the LLM bots demonstrating stronger play against traditional bots than against humans.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#623

Earlier quoted context omitted.

No, because there is no coherent, agreed-upon definition. There’s just a million people vibe defining it. Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever. Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not. And for what it’s worth I just watched…

Throughout this exchange you're repeatedly confusing the negative and the positive. I agree with you that there is no rigorous and universally agreed upon criteria for exactly what would constitute AGI (ie the positive). There are some vague shapes that are widely (but not universally) accepted such as largely (vague boundary) being capable of replacing (vague criteria) humans. However there are plenty of disqualifie…

> However there are plenty of disqualifiers that are more or less universally accepted (ie the negative)

Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.

An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#624

Earlier quoted context omitted.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…

What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill? Also what harness? If you’re using a general harness of course it’s going to try and give you commentary. I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when…

> but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.

This isn't just about judging LLM capability. This is about pointing out that these capabilities are not "AGI". If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.

Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it's hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#625
post #608

Earlier quoted context omitted.

Yes, say the sum of all salaries paid be employers is $65T, and if AI accelerates workers by 15.4% -- studies and survey data actually suggest it's closer to 33% already e.g. https://www.stlouisfed.org/on-the-economy/2025/nov/state-gen... -- that is worth 15.4% of 65T which is 10T, which is already "double-digit trillions" as I said. Assuming a 33% boost takes it to ~20T annually, which is technically "10s of trillio…

Did you read your link? It says that 33% of people adopt AI, not that AI makes all workers 33% more effective, that would be absolutely insane. For that to be true we'd expect to see either companies who use it have revenues all suddenly jumping 33% (which we have not seen) or laying off 33% of their staff (which we have not seen and would also be catastrophic in the short term at least). I'm not really going to bela…

Yep, I read the link, their detailed paper (https://s3.amazonaws.com/real.stlouisfed.org/wp/2024/2024-02...) and I also played with their data tracker here: https://www.genaiadoptiontracker.com/#explore-data

Toggle "Work hours using genAI" and "Time savings due to genAI" so see what I mean. You can also toggle between various industries, which is eye-opening because even industries like "Agriculture, Forestry, Fishing, and Hunting" are seeing productivity gains!

And if you do not want to read the detailed paper, they have a follow-up article here: https://www.stlouisfed.org/on-the-economy/2025/feb/impact-ge...

Specific quote:

> Using our data on generative AI use, this estimate implies that, on average, workers are 33% more productive in each hour that they use generative AI. This estimate is in line with the average estimated productivity gain from several randomized experiments on generative AI usage.

What we HAVE seen is that the national labor productivity has gone up by 1.3% since ChatGPT was released, and it lines up very well with all the other data and studies they cover... AND your reference, which predicted a 1.5% growth back in 2023!

Re: Why I'm still bearish on LLMs after Navier-Stokes

#626
post #321

Earlier quoted context omitted.

[flagged]

A Transformer has a massive amount of state - it's entire KV cache, in addition to the user asking it to draw the state after every move, which is really unnecessary. A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded? A Trans…

I just want to make sure it's clear: the reason I was asking it to redraw the board is because last time I tried (which was like a month ago), I didn't ask for that, and basically as soon as the opening was "out of book" it started trying to make illegal moves and made false statements about the position in its running commentary (and after being corrected on these points, started dropping pieces for no reason).

Re: Why I'm still bearish on LLMs after Navier-Stokes

#627
post #170

Earlier quoted context omitted.

>If AI were unleashed, in raw form today, it could cause havoc. What is "raw form?"

My exact question and afaict since I’m running models locally and can inspect and retrain etc..I assume I already have the raw form?

You are not running SOTA models locally, the gap is material. At least for now ... but by next year, the models we can run locally might cross that 'risk' threshold.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#629

Earlier quoted context omitted.

The story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.

> Are there reams of chess moves that the model can train off of? No. This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.

Let me make my statement more clear with a correction:

Was there reams of chess moves that the model trained off of? No.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#630

Earlier quoted context omitted.

The story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.

>Are there reams of chess moves that the model can train off of? No. For real??

I meant if there are reams of chess moves the model was trained off of.
Post reply on HN