Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

141–150 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#141
post #134
post #55

> These people used bad prompts and came to the conclusion that ChatGPT can’t play a legal chess game. (…) > With this prompt ChatGPT almost always plays fully legal games. > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table (…) > (…) with GPT4 (…) in the two games I attempted, it made numerous illegal moves. So you’ve ostensibly¹ found a way to reduce the error…

Is this the top comment (and not even grey) because more people failed to read the article than read it?

It does seem that way.

Re: ChatGPT's Chess Elo is 1400

#142

Earlier quoted context omitted.

I'm Elo 1400 and can beat literally everyone I know in the real world. I need to go online to find players at my skill level, or find tournament/competitive settings for a challenge. Yeah, I'm "class C", weak amateur chess player, but I think you're grossly underestimating the amount of study I put into this game. I'm not going to make an illegal move

You will under time pressure :) even Grandmasters have done that (I'm around 1900 elo for context) Also people forgetting they moved the king/rook and trying to castle.

Watch this ChatGPT game.

https://www.reddit.com/r/AnarchyChess/comments/10ydnbb/i_pla...

We're talking about pieces that don't exist, reappearing pieces, pieces moving completely wrong (Knight takes as if its a Pawn), etc. etc.

---------

People are taking these example games and saying ChatGPT is 1400 strength. I don't think so. This isn't a case of "oops, I castled even though I moved my king 15 turns ago".

Re: ChatGPT's Chess Elo is 1400

#143

I just opened a random recent chess game on lichess ( https://lichess.org/YpxTUUbO/white#88 ) . I'm pretty sure ChatGPT can't be trained on games that were just played, so this ensures the game is not in its training data. I gave the position before checkmate to ChatGPT to see if it would produce the checkmating move. It played an illegal move, replying with "Be5#" even there's no bishop of either color in the positi…

OP explained that you need to prompt the whole game, not just a position.

ChatGPT is an LLM, not a game tree engine. It needs the move history to help it create context for it's attention.

Re: ChatGPT's Chess Elo is 1400

#144
post #95
post #55

> These people used bad prompts and came to the conclusion that ChatGPT can’t play a legal chess game. (…) > With this prompt ChatGPT almost always plays fully legal games. > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table (…) > (…) with GPT4 (…) in the two games I attempted, it made numerous illegal moves. So you’ve ostensibly¹ found a way to reduce the error…

Fuller context from the article: > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table and saying “this game is impossible, I literally cannot conceive of how to win without breaking the rules of chess.” So whenever it wanted to make an illegal move, it resigned. (my emphasis) So the illegal moves are at least part of the reasons for the 6 losses, and factored into…

No ELO 1400 player will have that rate of illegal moves, so saying it that it plays with an ELO 1400 rating is disingenuous.

Reinterpreting illegal moves as resignation is absurd when an LLM is formally capable of expressing statements "I resign" or "I cannot conceive of a winning move from here" just as well as any human player. It just doesn't do so because it's not actually playing chess the way we think of an ELO 1400 player playing chess.

Re: ChatGPT's Chess Elo is 1400

#145

Earlier quoted context omitted.

1850 ELO player and also chess AI programmer here. This is an oversimplification at best. Many many games follow the same moves(1 move = 2 plies) for a long time, up to 30 moves in some cases, 20 moves is downright common and 10 moves is more common than not. These series of moves are referred to as opening theory and are described at copious length in tons of books. This is because while the raw number of possible p…

Yeah but now explain how it played a 61 move game. EDIT: I checked and it left the lichess database after 9 moves. The lichess db has probably 5 orders of magnitude more chess games in it than chatGPT has in its training data.

That's not the point. The point is if you truly want to test its strength, you'll have to control for these things. Maybe do things like invent a new form of notation and/or deliberately go into uncharted territory. Maybe start with a non-standard starting position even. Or play chess960 against it.

In theory if I was playing a 1200 player I would almost always win, but let's say they have some extremely devious preparation that I fell into due to nonchalance and by the time we're both out of book I'm down a queen. It might not matter that I'm 600 points stronger at that point. If they don't make a sufficient amount of errors in return I will lose anyway.

Re: ChatGPT's Chess Elo is 1400

#146

Earlier quoted context omitted.

If it isn not memorizing, how do you think is doing it?

by trying to learning the general rules that to explain the dataset and minimise its loss. That's what machine learning is about, it's not called machine memorising.

It cannot be learning the general rules because it occasionally tries to invent pieces out of whole cloth.

Re: ChatGPT's Chess Elo is 1400

#147
post #55

> These people used bad prompts and came to the conclusion that ChatGPT can’t play a legal chess game. (…) > With this prompt ChatGPT almost always plays fully legal games. > Occasionally it does make an illegal move, but I decided to interpret that as ChatGPT flipping the table (…) > (…) with GPT4 (…) in the two games I attempted, it made numerous illegal moves. So you’ve ostensibly¹ found a way to reduce the error…

Obviously the article should be taken with a giant grain of salt. That being said, not many things what aren't designed to play chess can play chess, with or without coaxing. My dog cannot, for instance, nor can my coffee table.

> My dog cannot, for instance, nor can my coffee table.

You must be giving them the wrong prompts.

Re: ChatGPT's Chess Elo is 1400

#148
This is GPT4, right? Because ChatGPT (GPT-3) still fails to provide a legal game of Tic Tac Toe with this prompt:

> "Let's play Tic Tac Toe. You are O, I'm X. Display the board in a frame, with references for the axes"

It failed to recognize that I won.

Then continued playing (past the end), played illegally over a move I had already done, obtained a line of 3 for itself, and still doesn't acknowledge the game has ended.

Re: ChatGPT's Chess Elo is 1400

#149
post #86

Earlier quoted context omitted.

I think there is very low percentage of players at elo 1400 who can provide a valid next move after seeing just the list of moves and not the current board state.

I'm Elo 1400 and can beat literally everyone I know in the real world. I need to go online to find players at my skill level, or find tournament/competitive settings for a challenge. Yeah, I'm "class C", weak amateur chess player, but I think you're grossly underestimating the amount of study I put into this game. I'm not going to make an illegal move

I'm much higher rated than you and I could not reliably play a legal game of chess just given a list of the moves and no board.

I suspect you can't either, you can try by turning on blindfold mode on lichess and seeing how far you get.

Re: ChatGPT's Chess Elo is 1400

#150

Earlier quoted context omitted.

That's how one uses any tool.

Yes, but it also completely invalidates the measurement of a 1400 elo rating. By comparison, any player making an illegal move is forfeiting the game, almost all people from ~300 elo can play without making illegal moves, chatgpt cant.

Why do illegal moves forfeit? In online play, they're validated. You can't make illegal moves. What's the ELO score if ChatGPT is corrected, and chooses a new move?
Post reply on HN