Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

241–250 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#241
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.

Back in 2001, our social medium was Slashdot and no one ever pretended to read the article. No one read the article either. It was slashdotted most of the time anyways.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#242
post #225

Earlier quoted context omitted.

That's quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).

So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don't see how this contradicts anything I said...

The _only_ illegal move a human might make as a beginner is a failed en passant or a bad castle. And yes, a few hundred games is all it takes to be better than any publicly available LLM at the moment.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#243

I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures. Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon ag…

I think automation is coming but it will be way more gnarly than frontier labs want public to believe. Value is just too big, when you can automate most of eg customer support it will create huge savings and same time customer satisfaction will get better.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#244

Earlier quoted context omitted.

I suck at chess. Are you saying I can't be intelligent?

If you read all chess tutorials, strategy documentation and game archives on the internet and then would still suck at chess: yes.

Declarative knowledge is not the same as procedural knowledge. You can read as many chess tutorials, strategy documentation and game archives as you like, they won't make you good at chess until you actually start practicing chess.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#245

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

" the lack of sentence capitalization makes it unnecessarily difficult to read."

If it had proper caps etc, people here would accuse it of written using LLMs.

You just can't win...

Re: Why I'm still bearish on LLMs after Navier-Stokes

#246
post #81
post #78

Earlier quoted context omitted.

The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…

So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.

An AGI doesn't stand for 'perfect intelligence' it stands for artificial general intelligence.

And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.

Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.

This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing

Re: Why I'm still bearish on LLMs after Navier-Stokes

#247
post #237

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

The gap in capabilities is mostly quantitative and not qualitative.

Is it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models.

Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#248

Earlier quoted context omitted.

>LLM can’t beat an avg chess player. Why should that matter?

If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this. So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

So we humans are not a general intelligence then?

And the stuff i'm using LLMs daily is just fake?

I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#249

Earlier quoted context omitted.

Llm systems are not really build for adhering to a grammar (other than "a string og tokens"). It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents. Certainly,a harness can easily correct for it.

It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.

How much support do we as humans need to get rules right?

I'm an expert in my field, read my comments, my gramma is shit.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#250

Earlier quoted context omitted.

Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there's a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn't make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge…

> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.

Thats just absolutly not true.

A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.

And there is a relevant and significant difference between the expectation of an AGI and an ASI system.

Post reply on HN