Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

591–600 of 646 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#591

Earlier quoted context omitted.

Completely arbitrary definition that nobody will agree on, stated as if it’s some self-evident ground truth.

Yes, it is indeed self evident. If it can't figure things out then its intelligence isn't general in which case it can't be AGI by definition .

No, because there is no coherent, agreed-upon definition. There’s just a million people vibe defining it.

Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever.

Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not.

And for what it’s worth I just watched GitHub Copilot figure something out. So your definition is once again lacking.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#592

Earlier quoted context omitted.

>What happens when you ask those same frontier models to write a chess-playing program? they shit out a carbon copy of https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.

I feel like you're saying something as odd as "Transistors still aren't good at playing chess." I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.

I'm not the one posting daily about how "AIs are going to destroy the world because of how smart they are", "humans are finished" and "we've reached super duper mega intelligence". Go see Dario and Sam about that.

>I'm pretty sure Fable could write AlphaZero

If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn't matter if it does: it's writing a solver: it's not good at playing chess. If tomorrow I tell you that I'm so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you're going to laugh me out of the room.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#593
post #367

Earlier quoted context omitted.

I very much agree that the next models will be better, heck, I still suck at hobbyist training and could probably coax t5 to do better in Chess specifically, just need to get loads of data from Stockfish. Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models ha…

So what is the supposed leap? One agent per option to change, evaluating the board state that there move would create, by having a army evaluate the remaining piece options and average over that? Wee-Free-Man as a hierarchical army ? Pet-LLMs trained on one thing?

Honestly, for intelligence I don't know and I doubt anyone can claim to know. Maybe JEPA, there is potential concerning some shortcomings inherent to LLMs but it has its own, maybe scaling up the electron microscope stuff Google just did (though the connections are inferred), maybe future implementations of autoregressive and diffusion LLMs can at some point address its issues after all, maybe something else entirely.

All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn't loose sight of that fact, especially as "not being intelligent" does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn't start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.

If for leap you just mean more utility from LLMs as they are, then I'll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).

Re: Why I'm still bearish on LLMs after Navier-Stokes

#594

Earlier quoted context omitted.

They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board? The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.

> They can't possibly remember even a few positions. Sure they could, but that's irrelevant. A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to. However, that's not how LLMs work. They don't memorize inputs - they predict them, based on discover…

> prediction which is closer to memorization

> don't memorize inputs - they predict them

I feel some tension here.

> rice grains on a chess board? Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.

> just a list of 64 numbers

> remember even a few positions? Sure they could, but that's irrelevant.

I don't think you do. Or rather you do know the legend but for some funny reason seem to be unable to apply its lesson here, because you are talking about enormous terabytes of training data.

> Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent.

If for you it is about intelligence, I am out of this discussion.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#595

Earlier quoted context omitted.

But does that matter? If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker? Probably the only thing saving many jobs…

Tests are often conducted under restricted conditions. For example, elementary school students aren't given calculators in math class, or during an interview, you are asked what encapsulation is and aren't allowed to use Google. The chess test effectively demonstrates the reasoning capabilities of an LLM without relying on brute force, because a human is incapable of calculating trillions of combinations yet plays ch…

People might care about this for chess, but no one really cares if an LLM can command an army or manage production of a business without any tools. If it can do those tasks reliably when given access to tools (including any tools it autonomously creates for itself), then that's more than sufficient. No one cares if an LLM is doing reasoning the way humans do it, as long as it can get the job done.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#596

Earlier quoted context omitted.

State-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behav…

For me the useful intuition is that LLMs haven't somehow magickally learned to implement any of the algorithms we know that we have used to make strong chess engines: alpha-beta minimax and Monte-Carlo Tree Search on the one hand, and obviously the ability to learn accurate evaluation functions by self-play. I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to…

https://xxcancel.com/biobootloader/status/164051244495839641...

Re: Why I'm still bearish on LLMs after Navier-Stokes

#597

Earlier quoted context omitted.

Right now AI hasn't even managed to replace all the human workers taking orders at the fast food drive thru. That's a job often performed by literal children and companies are still waiting for AI to get good enough for even that. Maybe one day it will be good enough, maybe one day it will outperform humans at such a basic task, but that day is not today. If the hype were anything close to reality, we'd see it everyw…

Funny you mention this - a fast food restaurant in my town now has an LLM taking drive-thru orders. Though I highly doubt it has taken anyone's job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)

Last I heard McDonald's and Taco Bell were trialing AI again at a limited number of stores. It's the kind of job AI should be really good at and many fast food companies are using call center workers currently. They really want AI to work, so they keep trying every few years to make it happen, but so far all they get are embarrassing social media posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#598
post #542

Earlier quoted context omitted.

This has the smell of "Why don't I have a faster horse". Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard. Also: https://www.txdot.gov/about/newsroom/statewide/air-taxi-test...

You're making my point for me, surprised you don't realize that...

Because you don't fully understand your own point...

You look at science fiction and say "why didn't I get flying cars" and not "why didn't most science fiction predict a global always on network that put the furthest places away from you a few microseconds away from audio, video, or any other type of information that can be digitally encoded.

Trying to use flying cars as a gotcha is missing that flying cars aren't near as useful as one would think in relation to their costs. Moving information has become far more useful than moving objects long distances quickly, especially humans.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#599
post #567

Earlier quoted context omitted.

I've found the same, the rate of bugs has dropped pretty dramatically after switching to ai generated code. I think it's partly because ai will write 1000s of lines of unit tests without complaining. I also have a github workflow where claude runs the /code-review command on every PR.

In the world of web apps, I find the agent's ability to write good e2e tests to be a real game changer. Turns out with enough rigor you can write pretty stable mostly not flaky e2e tests. And even flaky ones are fixed quickly due to a fuck ton of assertions at every step. Test code looks like a mess, even more than the usual LLM code. takes a while to let it go. Test report looks beautiful though.

When the test code "looks like a mess", how do you get assurance it is testing the right properties?

Like you, I've found that LLMs can improve test coverage by decreasing the amount of developer time spent writing tests. But generally, it takes a lot of manual work to set up the initial testing framework, and even then, a lot of vigilance to ensure that what is actually tested corresponds to the description of the test.

Post reply on HN