Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

111–120 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#111
I just opened a random recent chess game on lichess ( https://lichess.org/YpxTUUbO/white#88 ) . I'm pretty sure ChatGPT can't be trained on games that were just played, so this ensures the game is not in its training data.

I gave the position before checkmate to ChatGPT to see if it would produce the checkmating move. It played an illegal move, replying with "Be5#" even there's no bishop of either color in the position.

Unfortunately I'm rate limited at the moment so I can't try other games, but this looks like a solid method to evaluate how often ChatGPT plays legal / good moves.

Re: ChatGPT's Chess Elo is 1400

#113
post #50

Most likely it has seen a similar sequence of moves in its training set. There are numerous chess sites with databases displayed in the form of web pages with millions of games in them. If it had any understanding of chess, it would never play an illegal move. It's not surprising that given a sequence of algebraic notation it can regurgitate the next move in a similar sequence of algebraic notation.

> Most likely it has seen a similar sequence of moves in its training set. Is this a joke making fun of the common way people dismiss other ChatGPT successes? This makes no sense with respect to chess, because every game is unique, and playing a move from a different game in a new game is nonsensical.

Sorry, but not every game is unique. The following game has been played millions of times.

1. e4 e5 2. Bc4 Bc5 3. Qh5? Nf6?? 4. Qxf7++

The game Go has a claim to every game being unique. But not chess. And particularly not if both players follow a standard opening which there is a lot of theory about. Opening books often have lines 20+ moves deep that have been played many times. And grandmasters will play into these lines in tournament games so that they can reveal a novel idea that they came up with even farther in than that.

Re: ChatGPT's Chess Elo is 1400

#114
post #92
post #50

Earlier quoted context omitted.

> Most likely it has seen a similar sequence of moves in its training set. Is this a joke making fun of the common way people dismiss other ChatGPT successes? This makes no sense with respect to chess, because every game is unique, and playing a move from a different game in a new game is nonsensical.

Is it though? I mean if you had data on millions of games what is the chance that you'd find one which has identical position that the one you're in (it's not like most moves are random..) I wonder how well it could perform in Go, there are way more permutations there so finding an identical state should be more difficult.

You could certainly test this by making completely random moves and seeing whether it's more likely to make illegal moves in those positions.

Though I think you're overestimating how many positions have occured. Frequently, by move 20-25 you have a unique position that's never been played before (unless you're playing a well known main line or something)

Re: ChatGPT's Chess Elo is 1400

#115

This may look low: ELO for mediocre players is 1500. But if it is obeying the rules of the game, then this is big. This is a signal that if it learns some expertise, like discovering how to use or create better search algorithms (like MCTS and heuristics to evaluate a state) and improve by itself (somewhat like alphazero did), then it may eventually reach superhuman level. It may then reach superhuman level in any ta…

According to https://chess.stackexchange.com/questions/2550/what-are-the-... median rating is 1148 (252,989 Players). So it's beating half of humanity at a mind sport and it wasn't even specifically trained for it.

Re: ChatGPT's Chess Elo is 1400

#116
post #86

Earlier quoted context omitted.

I think there is very low percentage of players at elo 1400 who can provide a valid next move after seeing just the list of moves and not the current board state.

I'm Elo 1400 and can beat literally everyone I know in the real world. I need to go online to find players at my skill level, or find tournament/competitive settings for a challenge. Yeah, I'm "class C", weak amateur chess player, but I think you're grossly underestimating the amount of study I put into this game. I'm not going to make an illegal move

I mean can you play just based on being provided the input of a series of moves without it being shown to you as a visual board?

I guess most players would mess up 20/30 moves in.

Re: ChatGPT's Chess Elo is 1400

#117
post #10

Good thing it's "incapable of reasoning"!

Is a normal chess program capable of reasoning?

The Monte Carlo analysis AlphaZero used functioned as a sort of multi-step reasoning for it. GPT can use its token buffer for some multi-step reasoning but that sort of interferes with providing a conversation with the user so it's much less effective.

Re: ChatGPT's Chess Elo is 1400

#118

Most likely it has seen a similar sequence of moves in its training set. There are numerous chess sites with databases displayed in the form of web pages with millions of games in them. If it had any understanding of chess, it would never play an illegal move. It's not surprising that given a sequence of algebraic notation it can regurgitate the next move in a similar sequence of algebraic notation.

I played chess against ChatGPT4 a few days ago without any special prompt engineering, and it played at what I would estimate to be a ~1500-1700 level without making any illegal moves in a 49 move game.

Up to 10 or 15 moves, sure, we're well within common openings that could be regurgitated. By the time we're at move 20+, and especially 30+ and 40+, these are completely unique positions that haven't ever been reached before. I'd expect many more illegal moves just based on predicting sequences, though it's also possible I got "lucky" in my one game against ChatGPT and that it typically makes more errors than that.

Of course, all positions have _some_ structural similarity or patterns compared to past positions, otherwise how would an LLM ever learn them? The nature of ChatGPT's understanding has to be different from the nature of a human's understanding, but that's more of a philosophical or semantic distinction. To me, it's still fascinating that by "just" learning from millions of PGNs, ChatGPT builds up a model of chess rules and strategy that's good enough to play at a club level.

Re: ChatGPT's Chess Elo is 1400

#119
post #92
post #50

Earlier quoted context omitted.

> Most likely it has seen a similar sequence of moves in its training set. Is this a joke making fun of the common way people dismiss other ChatGPT successes? This makes no sense with respect to chess, because every game is unique, and playing a move from a different game in a new game is nonsensical.

Is it though? I mean if you had data on millions of games what is the chance that you'd find one which has identical position that the one you're in (it's not like most moves are random..) I wonder how well it could perform in Go, there are way more permutations there so finding an identical state should be more difficult.

>I mean if you had data on millions of games what is the chance that you'd find one which has identical position that the one you're in (it's not like most moves are random..)

Very low. On lichess when you analyse your games you can see which positions have been reached before, and you almost always diverge in the opening.

The lichess db has orders of magnitude more games of chess than the chatGPT training data does, so there is absolutely no way that chatGPT could reach 1400 purely based off positions in its training data.

Re: ChatGPT's Chess Elo is 1400

#120

Earlier quoted context omitted.

> Most likely it has seen a similar sequence of moves in its training set. Wouldn't we expect a much higher rate of illegal moves if that was the case?

Doesn't ChatGPT indeed have a very high number of illegal moves? https://www.youtube.com/watch?v=kvTs_nbc8Eg In this example, ChatGPT's first few moves are reasonable (while it appears to be on-book), but then it goes off the rails and starts moving illegally, spawning pieces out of nowhere, deleting pieces for no reason, etc.

I think it was not given the whole game up to that point, just individual moves. That was the point of this article - if you include all of the moves in the prompt, it is less likely to make illegal moves.
Post reply on HN