Earlier quoted context omitted.
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…
HN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.
Why I'm still bearish on LLMs after Navier-Stokes
241–250 of 646 posts
Re: Why I'm still bearish on LLMs after Navier-Stokes
#242Earlier quoted context omitted.
That's quite untrue. I taught my (adult) brother the moves, the only illegal move he ever tried against me (over his 6 first games) was a castle with a rook that already moved twice. Within a few hundred games (less than 500 for sure, he played 3 minutes blitz but always took at least 10 minutes analyzing his games) he was rated 1100 on lichess (which is like 1050 on chess.com and unranked in the real world).
So your brother tried to make illegal moves while learning the game and it took your brother hundreds of games to get to be a decent player? I don't see how this contradicts anything I said...
Re: Why I'm still bearish on LLMs after Navier-Stokes
#243I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures. Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon ag…
Re: Why I'm still bearish on LLMs after Navier-Stokes
#244Earlier quoted context omitted.
I suck at chess. Are you saying I can't be intelligent?
If you read all chess tutorials, strategy documentation and game archives on the internet and then would still suck at chess: yes.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#245Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
If it had proper caps etc, people here would accuse it of written using LLMs.
You just can't win...
Re: Why I'm still bearish on LLMs after Navier-Stokes
#246Earlier quoted context omitted.
The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…
So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
Re: Why I'm still bearish on LLMs after Navier-Stokes
#247Earlier quoted context omitted.
> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
The gap in capabilities is mostly quantitative and not qualitative.
Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#248Earlier quoted context omitted.
>LLM can’t beat an avg chess player. Why should that matter?
If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this. So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
And the stuff i'm using LLMs daily is just fake?
I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#249Earlier quoted context omitted.
Llm systems are not really build for adhering to a grammar (other than "a string og tokens"). It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents. Certainly,a harness can easily correct for it.
It seems absolutely crazy to me to expect an LLM to code a solution to a problem while also not expecting it to be able to adhere to a grammar.
I'm an expert in my field, read my comments, my gramma is shit.
Re: Why I'm still bearish on LLMs after Navier-Stokes
#250Earlier quoted context omitted.
Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. In fact there's a google paper on grandmaster level chess without search with a 270M transformer. Outside that, there was gpt-3.5-turbo instruct which was incidentally a 1800 lichess elo player that didn't make any illegal moves even after a few thousand moves. Frontier labs care deeply about automating knowledge…
> Frontier labs don't care about chess. If OpenAI cared, GPT-7 could be a grandmaster+ level chess player. If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.
A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.
And there is a relevant and significant difference between the expectation of an AGI and an ASI system.