Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

311–320 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#311
post #307

Earlier quoted context omitted.

I'd be interested in seeing this game, if you saved it?

I uploaded the PGN to lichess: https://lichess.org/rzSriO6I#97 After reviewing the chat history I actually have to issue a correction here, because there were two moves where ChatGPT played illegally: 1. ChatGPT tried to play 32. ... Nc5, despite there being a pawn on c5 2. ChatGPT tried to play 42. ... Kxe6, despite my king being on d5 It corrected itself after I questioned whether the previous move was legal. I was…

Thanks! Interesting game.

Qxd7 early on was puzzling but has been played in a handful of master games and it played a consistent setup after that with b5 Bb7. Which I imagine was also done in those master games. But interesting that it went for a sideline like that.

It played remarkably well although a bit lacking in plan. Then cratered in the endgame.

Bxd5 was strategically absurd. fxg4 is tactically absurd. Interestingly they both follow the pattern: Piece goes to square -> takes on that square.

This is of course an extremely common pattern, so again tentatively pointing towards predicting likely sequences of moves.

Ke7 was also a mistake but a somewhat unusual tactic with Re2 and f5 is forced but after en passant the knight is pinned. This tactic does appear in some e4 e5 openings though. But then the rook is on e1 and the king never moved or if it did, usually to e8, not e7. Possibly suggesting that it has blind spots for tactics when they don't appear on the usual squares?

Fascinating stuff.

Re: ChatGPT's Chess Elo is 1400

#312
post #204

This is so easy to disprove it makes it look like the author didn't even try. Here is the convo I just had: me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed ChatGPT: Sure, I'd be happy to help! Please provide the move sequence and I'll give you my response. me: 1. e3 ChatGP…

I don’t think this suffices as disproving the hypothesis. It’s possible to play at 1400 and make some idiotic moves in some cases. You really need to simulate a wide variety of games to find out, and that is what the OP did more of. Though I do agree it’s suggestive that your first (educated) try at an edge case seems to have found an error.

This is broadly the “AI makes dumb mistakes” problem; while being super-human in some dimensions, they make mistakes that are incredibly obvious to a human. This comes up a lot with self-driving cars too.

Just because they make a mistake that would be “idiots only” for humans, doesn’t mean they are at that level, because they are not human.

Re: ChatGPT's Chess Elo is 1400

#313
post #115

This may look low: ELO for mediocre players is 1500. But if it is obeying the rules of the game, then this is big. This is a signal that if it learns some expertise, like discovering how to use or create better search algorithms (like MCTS and heuristics to evaluate a state) and improve by itself (somewhat like alphazero did), then it may eventually reach superhuman level. It may then reach superhuman level in any ta…

According to https://chess.stackexchange.com/questions/2550/what-are-the-... median rating is 1148 (252,989 Players). So it's beating half of humanity at a mind sport and it wasn't even specifically trained for it.

The median chess player is usually described as mediocre (if you ask chess players). They suck as badly as the median clarinet player in your high school band/orchestra.

Re: ChatGPT's Chess Elo is 1400

#314
post #303

I just deployed a GPT-4 powered chess bot to lichess. You can challenge it here: https://lichess.org/@/oopsallbots-gpt-4

Very cool! Are you doing prompt engineering, fine-tuning, both, something else?

I'm wondering if it'd be cool to have a chess contest where all the bots are LLM powered. Seems to me like the contest would have to ban prompt engineering -- would have to have a fixed prompt -- otherwise people would sneak chess engines into their prompt generation.

Re: ChatGPT's Chess Elo is 1400

#315

Earlier quoted context omitted.

> the strongest blitz players in the world are hundreds of points higher rated on chess.com blitz versus their FIDE blitz rating Online rating inflation is real but I'm not sure blitz is the best example of it because in that case there is a notable difference between online and otb (having to take time to physically move the pieces).

Probably the bigger difference is ability to premove online

I was thinking about this.

On chess.com you can chain premoves, on lichess you can't(afaik).

So in theory, to the extent premoves explain the rating difference, the difference should be greater on chess.com assuming they have the same parameters in their rating calculations. Therefore it should be possible to perform an analysis to shed light on this. But someone would have to go recompute the 3 different ratings under the same system first to be able to make a sensible analysis.

Re: ChatGPT's Chess Elo is 1400

#316
post #310

Earlier quoted context omitted.

Why do people always have to interpret everything in absolute terms? It's clearly following some opening theory in all the games I've looked at so far. So yes, it is regurgitating opening moves. That's clearly not all it's doing, which is very impressive, but these are not mutually exclusive.

I am responding to OP, who said "Most likely it has seen a similar sequence of moves in its training set." From this, I take it that the question is if ChatGPT is repeating existing games, or not. All you need is a single game where it's not repeating a single game to prove it definitively. You can hardly play 60 moves without an error by accident. I believe you're responding to a different question, something like "…

The OP was too unsophisticated in their analysis(as is TFA), no doubt. But I'm not too interested in what OP said or who was wrong or not, and rather more interested in finding what's right.

As someone very clever once said, welcome to the end of the thought process.

We've established that:

1. It doesn't repeat entire games when the games go long enough

2. It does repeat a lot of opening theory

3. It seems to repeat common, partially position independent tactical sequences even when they're illegal or don't work tactically.

Re: ChatGPT's Chess Elo is 1400

#317
post #204

This is so easy to disprove it makes it look like the author didn't even try. Here is the convo I just had: me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed ChatGPT: Sure, I'd be happy to help! Please provide the move sequence and I'll give you my response. me: 1. e3 ChatGP…

[dead]

Re: ChatGPT's Chess Elo is 1400

#318
post #267

Earlier quoted context omitted.

This argument is pretty flimsy. ChatGPT makes illegal moves frequently. In all my years of playing competitive chess (from 1000 to 2200), I have never seen an illegal move. I'm sure it has happened to someone, but it's extremely rare. ChatGPT does it all the time. No one is arguing that humans never make illegal moves; they're arguing that ChatGPT makes illegal moves at a significantly higher rate than a 1400 player…

> No one is arguing that humans never make illegal moves > something a 1400 ranked player would never do > fine, fair, "never" was too much. I mean, yes they were and they said as much after I called them out on it. But go off on how nobody is arguing the literal thing that was being argued. It's not like messages are threaded or something, and read top-down. You would have 100% had to read the comment I replied to f…

You have twice removed the substance of an argument and responded to an irrelevant nitpick. Here's what the OP said:

> He literally used the same prompt as the article. > Claim: "ChatGPT's Chess Elo is 1400"

> Reality: ChatGPT gives illegal moves (this happened to article author too),

> something a 1400 ranked player would never do

> Result: ChatGPT's rank is not 1400.

This is a completely fair argument that makes perfect sense to anyone with knowledge of competitive chess. I have never seen a 1400 make an illegal move. He probably hasn't either. Your point is literally correct in the sense that at some point in history a 1400 rated player has made an illegal move, but it completely misses the point of his argument: ChatGPT makes illegal moves at such an astronomically high rate that it wouldn't even be allowed to even play competitively, hence it cannot be accurately assessed at 1400 rating.

Imagine you made a bot that spewed random letters and said "My bot writes English as well as a native speaker, so long as you remove all of the letters that don't make sense." A native English speaker says, "You can't say the bot speaks English as well as a native speaker, since a native speaker would never write all those random letters." You would be correct in pointing out that sometimes native speakers make mistakes, but you would also be entirely missing the point. That's what's happening here.

Re: ChatGPT's Chess Elo is 1400

#319
post #204

This is so easy to disprove it makes it look like the author didn't even try. Here is the convo I just had: me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed ChatGPT: Sure, I'd be happy to help! Please provide the move sequence and I'll give you my response. me: 1. e3 ChatGP…

I played a game against it yesterday (it won) and the only time it made an ilegal was move 15 (the game was unique according to lichess database from much earlier) so I just asked it to try again. There's variance in what you get but your example seems much worse.

Re: ChatGPT's Chess Elo is 1400

#320
post #318

Earlier quoted context omitted.

> No one is arguing that humans never make illegal moves > something a 1400 ranked player would never do > fine, fair, "never" was too much. I mean, yes they were and they said as much after I called them out on it. But go off on how nobody is arguing the literal thing that was being argued. It's not like messages are threaded or something, and read top-down. You would have 100% had to read the comment I replied to f…

You have twice removed the substance of an argument and responded to an irrelevant nitpick. Here's what the OP said: > He literally used the same prompt as the article. > Claim: "ChatGPT's Chess Elo is 1400" > Reality: ChatGPT gives illegal moves (this happened to article author too), > something a 1400 ranked player would never do > Result: ChatGPT's rank is not 1400. This is a completely fair argument that makes pe…

> I have never seen a 1400 make an illegal move.

Ah yes, of course, just because you never saw it means it never happens. That's definitely why rules exist around this specific thing happening. Because it never happens. Totally.

In fact, it's so rare that in order to forefeit a game, you have to do it twice. But it never happens, ever, because pattrn has never seen it. Case closed everyone.

I made no judgement on what ChatGPT can and can't do. I pointed out an extreme. Which the commenter agreed was an extreme. The rest of your comment is completely irrelevant but congrats on getting tilted over something that literally doesn't concern you. Next time, just save us both the time and effort and don't bother butting in with irrelevant opinions. Especially if you couldn't even bother to read what was already said.

Post reply on HN