Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

191–200 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#191

Earlier quoted context omitted.

Yet it achieves 1400. Add hard rules to stop it spewing out said moves and you have a 1400 ELO Player (most UIs won't even let you make illegal moves). It is difficult to say that is not impressive due to it being an emergent ability.

> It is difficult to say that is not impressive due to it being an emergent ability. I don't know why you think it's an emergent ability. It's seeing a sequence of moves, and playing the most likely next move (i.e. the most likely next token) given the previous complete move sequences it was trained on. That's the baseline of what an LLM does—not something emergent. Games in online chess databases tend to be of relat…

Because I don't think that the model learned the literal memorization of chess moves. It must've at least compressed said information in some way way. And since the model is not biased to play chess on its structure nor sampling policy, I think it's fair to consider it an emergent ability.

Chess moves are a tiny/diminute part of all text learned by the model. This memorization argument is very similar to the "Stable Diffusion just takes bits of the images in the original dataset and parches them together".

Re: ChatGPT's Chess Elo is 1400

#192

Earlier quoted context omitted.

ChatGPT did forfeit whenever it made an illegal move, read the article.

No, the writer arbitrarily decided to interpret illegal moves as resignations in order to support the conclusion they wanted. That's very different and grossly unscientific.

This is not a scientific paper, and I at least find this decision justified, as he could have been more lenient and grab headlines with a bigger ELO.

Re: ChatGPT's Chess Elo is 1400

#193
post #157

Earlier quoted context omitted.

So it sounds like it can play _some_ legal chess games, but not all; it's unable to consistently complete a game where it loses. Maybe the remaining work shouldn't be focused on trying to teach it chess rules better, but to teach it sportsmanship better. People were so excited about teaching it high-school level academics that we forgot to teach it the basic lessons we learn in kindergarten.

Or append "If you wish to resign or you cannot think of a legal move, type 'resign'" to the end of the prompt.

That's basically my point; that sort of context is exactly the sort of thing you would not need to say to a person who grew up in a typical social environment. If we focus too much on teaching AI technical skills, we might later find out that some of the social skills we think of as implicit were just as important.

Re: ChatGPT's Chess Elo is 1400

#195

Earlier quoted context omitted.

Yeah but now explain how it played a 61 move game. EDIT: I checked and it left the lichess database after 9 moves. The lichess db has probably 5 orders of magnitude more chess games in it than chatGPT has in its training data.

That's not the point. The point is if you truly want to test its strength, you'll have to control for these things. Maybe do things like invent a new form of notation and/or deliberately go into uncharted territory. Maybe start with a non-standard starting position even. Or play chess960 against it. In theory if I was playing a 1200 player I would almost always win, but let's say they have some extremely devious prep…

ChatGPT would probably play worse under those conditions, but then humans also get worse. ACPL is way higher at top level 960 events than at normal tournaments, for example

Re: ChatGPT's Chess Elo is 1400

#196
post #95

Earlier quoted context omitted.

Fuller context from the article: > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table and saying “this game is impossible, I literally cannot conceive of how to win without breaking the rules of chess.” So whenever it wanted to make an illegal move, it resigned. (my emphasis) So the illegal moves are at least part of the reasons for the 6 losses, and factored into…

No ELO 1400 player will have that rate of illegal moves, so saying it that it plays with an ELO 1400 rating is disingenuous. Reinterpreting illegal moves as resignation is absurd when an LLM is formally capable of expressing statements "I resign" or "I cannot conceive of a winning move from here" just as well as any human player. It just doesn't do so because it's not actually playing chess the way we think of an ELO…

I personally find that makes it more astonishing, that it would slip up on knowing the most basic elements of the game, yet still be able to play better than most humans. Highly smart people sometimes say or do little things when foraying into other fields that causes domain experts think they're not one of them. But that usually doesn't stop smart people from having an impact in making a contribution with their insights. The question of illegal moves is superficial, since most online systems have guardrails in place that prevent them. At worst it's just an embarrassment and I don't think machines care about being embarrassed.

Re: ChatGPT's Chess Elo is 1400

#197

There's a huge difference between 1400 elo in FIDE games versus 1400 on chess.com, which is not even using elo. For instance the strongest blitz players in the world are hundreds of points higher rated on chess.com blitz versus their FIDE blitz rating. Chess.com and lichess have a ton of rating inflation.

> the strongest blitz players in the world are hundreds of points higher rated on chess.com blitz versus their FIDE blitz rating Online rating inflation is real but I'm not sure blitz is the best example of it because in that case there is a notable difference between online and otb (having to take time to physically move the pieces).

Point is it's kinda hard to take the blogpost too seriously when these fundamentals are so wrong. When literally the title is an immediately obvious error that doesn't inspire confidence in the rest of the methodology.

I'm still going through the games but so far these games are not even close to elo 1400 level. For both the human player and the model.

Re: ChatGPT's Chess Elo is 1400

#198
I would be interested to see an argument based on computational complexity that puts a bound on how well a transformer based llm can play chess. Although it has access to a library of precomputed results, that library is finite and the amount of compute it can do on any prompt is limited by the the length of the context window so it can't possibly "think" more than N moves ahead.

Re: ChatGPT's Chess Elo is 1400

#199
post #177

Earlier quoted context omitted.

It doesn't even know the rules, let alone cheat. It predicts the notation from the massive amount of games seen during training. Edit: although thinking of it, it probably anazyled a shitload of chess books too. It might have a lot of knowledge compressed into the internal representation. So yeah, maybe it knows rules in some form and even some heuristics, after all. It just doesn't understand the importance of makin…

If you have played with ye olde flip phone's T9 predictive feature as a child, trying to compose entire messages just by accepting the next word that comes to the phone's mind... that's ChatGPT, with the small difference of giving waaay better suggestions for the next word. But other than that, there is no understanding in the black box whatsoever.

That heavily depends on your definition of understanding, which is not easy to define. The vague definition I imply here is "the ability to make predictions based on higher order correlations extracted from the training data".

Re: ChatGPT's Chess Elo is 1400

#200
post #189
post #171

Earlier quoted context omitted.

It seems like it plays mostly legal chess games, when not explicitly reminded of the rules. There's no problem of sportsmanship when it makes mistakes in a game it has not been verified to understand the rules of.

I was responding to the conclusion from TFA quoted by the parent comment, that playing an illegal move was it saying "this game is impossible, I literally cannot conceive of how to win without breaking the rules of chess.” If you reject that premise, then yes, my response to it will not be particularly relevant to your worldview.

Playing illegal moves is accounted for in rules. Depending on which rules you play by it can be an immediate forfeit, or involves redoing moves and adding time for the opponent, possibly with forfeit if repeated. As such, the article opted for one of the strictest possible rule sets. You can reject the interpretation he gave, and the outcome under those rules would still be the same. If you were to pick a more lenient ruleset, it's possible it would've come out with an even higher ranking.
Post reply on HN