Earlier quoted context omitted.
> So we humans are not a general intelligence then? No, because we can , in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us. > And the stuff i'm using LLMs daily is just fake? This misses the point completely.
> generally read the rules of a game and then follow them How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
Why I'm still bearish on LLMs after Navier-Stokes
621–630 of 646 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#622Earlier quoted context omitted.
Luna is one of the budget lower-end last generation models. It'd be useful to at least try to verify the present before being bearish about the future. For OpenAI, the best publicly available model is GPT-6 Astra with XHigh or Max reasoning, and for Anthropic it's Claude Fable 5.1 with XHigh or Max reasoning.
out of curiosity, do you think fable would get this right? (I'm not sure myself, and haven't tried yet.)
Re: Why I'm still bearish on LLMs after Navier-Stokes
#623Earlier quoted context omitted.
No, because there is no coherent, agreed-upon definition. There’s just a million people vibe defining it. Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever. Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not. And for what it’s worth I just watched…
Throughout this exchange you're repeatedly confusing the negative and the positive. I agree with you that there is no rigorous and universally agreed upon criteria for exactly what would constitute AGI (ie the positive). There are some vague shapes that are widely (but not universally) accepted such as largely (vague boundary) being capable of replacing (vague criteria) humans. However there are plenty of disqualifie…
Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.
An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#624Earlier quoted context omitted.
Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…
What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill? Also what harness? If you’re using a general harness of course it’s going to try and give you commentary. I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when…
This isn't just about judging LLM capability. This is about pointing out that these capabilities are not "AGI". If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.
Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it's hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#625Earlier quoted context omitted.
Yes, say the sum of all salaries paid be employers is $65T, and if AI accelerates workers by 15.4% -- studies and survey data actually suggest it's closer to 33% already e.g. https://www.stlouisfed.org/on-the-economy/2025/nov/state-gen... -- that is worth 15.4% of 65T which is 10T, which is already "double-digit trillions" as I said. Assuming a 33% boost takes it to ~20T annually, which is technically "10s of trillio…
Did you read your link? It says that 33% of people adopt AI, not that AI makes all workers 33% more effective, that would be absolutely insane. For that to be true we'd expect to see either companies who use it have revenues all suddenly jumping 33% (which we have not seen) or laying off 33% of their staff (which we have not seen and would also be catastrophic in the short term at least). I'm not really going to bela…
Toggle "Work hours using genAI" and "Time savings due to genAI" so see what I mean. You can also toggle between various industries, which is eye-opening because even industries like "Agriculture, Forestry, Fishing, and Hunting" are seeing productivity gains!
And if you do not want to read the detailed paper, they have a follow-up article here: https://www.stlouisfed.org/on-the-economy/2025/feb/impact-ge...
Specific quote:
> Using our data on generative AI use, this estimate implies that, on average, workers are 33% more productive in each hour that they use generative AI. This estimate is in line with the average estimated productivity gain from several randomized experiments on generative AI usage.
What we HAVE seen is that the national labor productivity has gone up by 1.3% since ChatGPT was released, and it lines up very well with all the other data and studies they cover... AND your reference, which predicted a 1.5% growth back in 2023!
Re: Why I'm still bearish on LLMs after Navier-Stokes
#626Earlier quoted context omitted.
[flagged]
A Transformer has a massive amount of state - it's entire KV cache, in addition to the user asking it to draw the state after every move, which is really unnecessary. A human, at least a trained human (for fairer comparison to an LLM whose training data contained a ton of chess games) can absolutely do this - have you never seen demonstrations of expert players playing a dozen or more games while blindfolded? A Trans…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#627Earlier quoted context omitted.
>If AI were unleashed, in raw form today, it could cause havoc. What is "raw form?"
My exact question and afaict since I’m running models locally and can inspect and retrain etc..I assume I already have the raw form?
Re: Why I'm still bearish on LLMs after Navier-Stokes
#628Re: Why I'm still bearish on LLMs after Navier-Stokes
#629Earlier quoted context omitted.
The story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.
> Are there reams of chess moves that the model can train off of? No. This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.
Was there reams of chess moves that the model trained off of? No.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#630Earlier quoted context omitted.
The story isn't so clear cut. The caveat is: It depends on the task. Are there reams of chess moves that the model can train off of? No. Are there reams of math papers the model can train off of? Yes.
>Are there reams of chess moves that the model can train off of? No. For real??