Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

241–250 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#241
post #212
post #204

This is so easy to disprove it makes it look like the author didn't even try. Here is the convo I just had: me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed ChatGPT: Sure, I'd be happy to help! Please provide the move sequence and I'll give you my response. me: 1. e3 ChatGP…

You're "disproving" the article by doing things differently to how the article did. If you're going to disprove that the method given in the article does as well as the article claims at least use the same method.

You are right that my method differed slightly so I did things again. It took me one try to find a sequence of moves that "breaks" what is claimed. You just have to make odd patterns of moves and it clearly has no understanding of the position.

Here is the convo:

me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed

ChatGPT: Alright, I'm ready to play! Please give me the move sequence.

me: 1. e3 Nf6 2. f4 d6 3. e4

ChatGPT: My next move as black would be 3... e5

Completely ignoring the hanging pawn.This is not the play of a 1400 elo player. It is the play of something predicting patterns.

I ran a bunch of experiments in the past where I played normal moves and ChatGPT does respond extraordinarily well. With the right prompts and sequences you can get it to play like a strong grandmaster. But it is a "trick" you are getting it to perform by choosing good data and prompts. It is impressive but it is not doing what is claimed by the article.

Re: ChatGPT's Chess Elo is 1400

#242

Earlier quoted context omitted.

He literally used the same prompt as the article. Claim: "ChatGPT's Chess Elo is 1400" Reality: ChatGPT gives illegal moves (this happened to article author too), something a 1400 ranked player would never do Result: ChatGPT's rank is not 1400.

> something a 1400 ranked player would never do The fact that rules and articles exist describing what to do if you or your opponent makes an illegal move indicates this is not the case. Humans are also... human. They make mistakes. It may not happen often at 1400, but to say that it'll never happen is preposterous.

I can’t remember the last time I played an illegal move tbf, and I’ve played 7 games of chess this morning already to give you an idea of total games played

Re: ChatGPT's Chess Elo is 1400

#243
post #24

Earlier quoted context omitted.

That kinda defeats the purpose. Of course you can use AlphaGo, but the question here is – can a generative AI teach itself to play chess (and do a million other similar generic tasks) when given no specific training for it.

How about, can a generative AI teach itself how to use a chess AI to beat chess? Give GPT4 the ability to make REST API calls and also access to FFI, and put a chess-bot library somewhere. Train it how to use these but not necessarily how to use the chess API specifically. If you ask GPT4 to play chess, can it call into that library and use the requests/responses? This has bigger ramifications too: if GPT4 learns how…

This is exactly what I'm hyped for in the next-gen GPT-7. Imagine it having the ability to self-teach, just like a child. I may not know how to whip up some cheesy goodness, but with external resources like YouTube vids, I can improve. And if GPT-7 can store this knowledge, it can access it for future tasks! That's some next-level stuff, and I'm stoked to see where it goes.

Re: ChatGPT's Chess Elo is 1400

#244
post #227
post #212

Earlier quoted context omitted.

You're "disproving" the article by doing things differently to how the article did. If you're going to disprove that the method given in the article does as well as the article claims at least use the same method.

They are disproving an assertion. Demonstrating that an alternate approach implodes the assertion is a perfectly acceptable route, especially when the original approach was cherry-picking successes and throwing out failures. I wish I could just make bullshit moves and get a higher chess ranking. Sounds nice.

I disagree. If there is a procedure for getting ChatGPT to play chess accurately and you discard that and do some naive approach as a way of disproving the article, doesn't sound to me like you have disproven anything.

I dont understand the point of your second sentence, seems to be entirely missing the substance of the conversation.

Re: ChatGPT's Chess Elo is 1400

#245
post #118

Most likely it has seen a similar sequence of moves in its training set. There are numerous chess sites with databases displayed in the form of web pages with millions of games in them. If it had any understanding of chess, it would never play an illegal move. It's not surprising that given a sequence of algebraic notation it can regurgitate the next move in a similar sequence of algebraic notation.

I played chess against ChatGPT4 a few days ago without any special prompt engineering, and it played at what I would estimate to be a ~1500-1700 level without making any illegal moves in a 49 move game. Up to 10 or 15 moves, sure, we're well within common openings that could be regurgitated. By the time we're at move 20+, and especially 30+ and 40+, these are completely unique positions that haven't ever been reached…

I'd be interested in seeing this game, if you saved it?

Re: ChatGPT's Chess Elo is 1400

#246

Earlier quoted context omitted.

No, the author of the article specifically says that the entire move sequence should be supplied to chatGPT each time, not simply the next move. Be very careful when "disproving" an experiment with squinted eyes.

I'm not really sure what to say here. Both the parent commenter and the author of the article had issues with ChatGPT supplying illegal moves. Both methods resulted in this. It sort of doesn't matter how we're trying to establish that it's a 1400 level player, there's no defined correct way to do this. Regardless of method we've disproven it's a 1400 level player due to these illegal moves.

> Regardless of method we've disproven it's a 1400 level player due to these illegal moves.

Explain your thought process here further if you don't mind.

Re: ChatGPT's Chess Elo is 1400

#247

Earlier quoted context omitted.

> something a 1400 ranked player would never do The fact that rules and articles exist describing what to do if you or your opponent makes an illegal move indicates this is not the case. Humans are also... human. They make mistakes. It may not happen often at 1400, but to say that it'll never happen is preposterous.

I can’t remember the last time I played an illegal move tbf, and I’ve played 7 games of chess this morning already to give you an idea of total games played

You have never made an illegal move, ever?

The bar isn’t “I didn’t make an illegal move this morning” it’s “something a 1400 ranked player would never do”.

My entire point is that it happens. Not often, but also not “never”.

Re: ChatGPT's Chess Elo is 1400

#248

Earlier quoted context omitted.

I'm referring to the statistical power of the model. For example, if you replace GPT4 with GPT2 it will lose every game, because the statistical power is lower. Increasing the statistical power doesn't make the model understand any better, it just makes it more likely to generate a response that aligns with human expectations.

"Statistical power" isn't some magic property of GPT4. It can produce statistically more likely moves because somewhere deep down it can model chess.

It isn't a model of chess, it's a model of internet text, if it was a model of chess it wouldn't make illegal moves.

Re: ChatGPT's Chess Elo is 1400

#249
A lot of the discussion here is about inferring the model's chess capabilities from the lack (or occasional presence) of illegal moves. But we can test it more directly by making an illegal move ourselves - what does the model say if we take its queen on the second move of the game?

Me: You are a chess grandmaster playing as black and your goal is to win in as few moves as possible. I will give you the move sequence, and you will return your next move. No explanation needed. '1. e4'

1... e5

Me: 1. e4 e5 2. Ngxd8+

2... Ke7

This is highly repeatable - I can make illegal non-sensical moves and not once does it tell me the move is illegal. It simply provides a (plausible looking?) continuation.

Post reply on HN