Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

121–130 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#121

Earlier quoted context omitted.

Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck

If all you do is read chess books, you'll be a shit player. Training and practice is what it takes to be great.

WTF even is this post?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#122

Earlier quoted context omitted.

If humans were actually intelligent, they wouldn't need to train and practice to play good chess. I mean, what level do you think people without any practice or training are ?

Except all these LLMs were already trained with hundreds of chess book and game databases and they still suck

Contrary to popular belief, you need a lot of training on something for an LLM to be good and consistent with it.

People think that if one mention exists in the training set, then the LLM is perfect at it.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#123
post #96

Earlier quoted context omitted.

Thanks for saying this, feels like everyone has gone insane over this stuff.

Humans don’t code a $game engine to play $game, they can just play it. It seems like you are the one that has gone insane.

And how many years of direct play and study does it take for a human to get good at chess or any other game? Absolutely no human ever could be good at chess just by reading a few books, or even every book on chess. That's just not how the brain works. If LLMs could do that they would truly be superintelligence.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#124
post #93
post #81

Earlier quoted context omitted.

So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.

I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well. People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

I agree w/ this perspective. An agent with a harness that can run programs can solve a lot more than one without the harness. The AI system includes the harness, and it's not clear to me that AGI requires more than LLMs + code generation & execution are capable of.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#125
post #93

Earlier quoted context omitted.

I'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well. People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.

Then why respond at all for the sake of responding? We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own. You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess. I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the ot…

He keeps posting with a particular type of tone.

He definitely needs to touch grass.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#126

Earlier quoted context omitted.

This isn't how intelligence works. The LLM may not be able to play chess directly through inference, but it can write a program to do it and execute that program. Same as how human intelligence works. We can't fly, but we can build planes.

Human beings can play chess directly without coding up a tool.

Very poorly compared to the tools we have built. Similar to the LLM.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#127
post #20

Earlier quoted context omitted.

> Is this really any different to how humans learn yes.

being a bit more specific: the sample efficiency of humans is orders of magnitude larger for more abstract concepts. the same doesn't hold for memory-intensive tasks though (like any kind of trivia), but that only takes you so far.

We've had technology beating humans on memory for millennia, and we've had technology beating humans on computation for many decades now.

The tricky thing with LLMs is describing what they actually do. They are too clearly beating humans on some things, but what exactly? Memory – already done, they're bad at basic computation (all LLMs just write code for actual computation/calculation). And as you say, they do badly at more abstract concepts.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#128
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.

that is if it even a human commenter at all

Re: Why I'm still bearish on LLMs after Navier-Stokes

#129

Earlier quoted context omitted.

Human beings can play chess directly without coding up a tool.

Very poorly compared to the tools we have built. Similar to the LLM.

Poorly in what sense? I think human chess leagues are way more popular and fun than just playing a computer by yourself. Human oriented communities are always a vastly better experience than their digital counterparts.

There's more to games than simply winning you know.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#130

Not convinced by those points. In particular, I found this very misleading or irrelevant: a typical CPU project anecdotally has about three times as many specification and validation engineers as design engineers and a 5:1 ratio is not unheard of The reason silicon design has such verification to design ratio is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost an…

> The reason ... is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost and in schedule cost (it takes months ... and if you messed up and need to spin a fix, it costs tens of millions of dollars, not counting any design engineering cost).

Aren't you just describing waterfall? That's still very prevalent in software engineering, and pretty much any other type of engineering – civil, chemical, building, architecture, drug discovery.

It's typically true that software can fail faster and cheaper, but it's also true that the costs are still vastly higher to fix later in the process.

Post reply on HN