Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

231–240 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#231
post #148

This is GPT4, right? Because ChatGPT (GPT-3) still fails to provide a legal game of Tic Tac Toe with this prompt: > "Let's play Tic Tac Toe. You are O, I'm X. Display the board in a frame, with references for the axes" It failed to recognize that I won. Then continued playing (past the end), played illegally over a move I had already done, obtained a line of 3 for itself, and still doesn't acknowledge the game has en…

No, it is GPT 3.5 and the real answer to every debate about capabilities comes down to "you're prompting it wrong".

Can you describe how to prompt gpt to play a proper game of tic tac toe? I was disappointed gpt-4 didn’t really improve much on this front, at least with how I was prompting it.

Overall I’m very blown away by the generative AI developments. I’m also very curious about which types of prompts are most effective for each task, etc.

Re: ChatGPT's Chess Elo is 1400

#232
Not sure how useful this is but I simply asked ChatGPT what it thinks its skill level is. This is using GPT-4.

After some previous questions...

Me: "Would you be able to estimate your ELO?"

ChatGPT: "As a language model, I don't have a specific ELO rating. However, my knowledge of chess strategies and tactics is based on the information I have been trained on, which includes games played by humans with various ELO ratings. My level of play might be roughly equivalent to an ELO of around 1200-1500, but please keep in mind that this is only an estimate, and my actual performance could be higher or lower depending on the specific position or game."

Me: "As a generative AI, are you actually able to reason which move to make next in chess? Or are you somehow just predicting the next likely move based on training data."

ChatGPT: "As a generative AI language model, my primary function is to predict the most likely text given a certain context. When it comes to playing chess, I am not "reasoning" in the same way that a dedicated chess engine would, which involves evaluating positions and searching through a large number of potential moves using specialized algorithms.

Instead, I generate my chess moves based on patterns and sequences I've seen during my training, which include a large number of chess games and discussions about chess strategy. My knowledge is derived from the text data I've been trained on, and I apply that knowledge to suggest moves that I think are appropriate for a given position.

While this approach allows me to play at an intermediate level, it is not as precise or accurate as the methods used by dedicated chess engines. My suggestions might not always be the best moves, and my understanding of the game is limited by the quality and quantity of my training data."

Re: ChatGPT's Chess Elo is 1400

#233
post #226

Earlier quoted context omitted.

> Fine, just wrap the LLM in a simple function that detects illegal moves and replaces them with "I resign" or "I cannot conceive of a winning move from here". Then you aren't "reinterpreting" anymore. Then it's still isn't anywhere near ELO 1400.

Under FIDE rules it's first a forfeit after the second illegal move, so if anything it would seem that the interpretation used by the article author underestimates its ELO ranking.

Nope, still not even close to what the author claims. If I understand it correctly, it made illegal moves in 3 out of 19 games. That's probably a few orders of magnitude more illegal moves than even a 1400 ELO player would make of their entire lifetime.

Re: ChatGPT's Chess Elo is 1400

#234

Earlier quoted context omitted.

ChatGPT did forfeit whenever it made an illegal move, read the article.

No, the writer arbitrarily decided to interpret illegal moves as resignations in order to support the conclusion they wanted. That's very different and grossly unscientific.

I mean, that's more lenient than the official "interpretation" (rule) which is that your second illegal move results in a forfeit.

Re: ChatGPT's Chess Elo is 1400

#235
post #34

Why not just introduce AlphaGo as an API that can be used by chatGPT? So every time you want to do a this type of gaming, you just send a request. I mean, chatGPT sends a request to AlphaGo, but as a user you don't know actually what's happening. But in the background, it happens really fast, so it's just like you are chatting with chatGPT, but using much, much powerful tool to do this kind of things.

This is actually a huge debate right now. OpenAI is on the side of 'LLMs have only surprised us to the upside, so using crutches is counterproductive' Whereas other people think 'Teaching an LLM to do arbitrary math problems through brute force is probably one of the most wasteful things imaginable when calculators exist.' I'm actually very excited to see which side wins (I'm on team calculator, but want to be on tea…

How about a more human-like approach: the LLM designs a calculator and then makes use of that!

Re: ChatGPT's Chess Elo is 1400

#236

Earlier quoted context omitted.

> But since it is so highly trained you can mistake it for a master if you squint and don't look into what it is doing. But it is a master, as has been pointed out repeatedly. If you replace all illegal moves with resignations, and use the same style of prompt as the OP did, then it plays like an expert. I'm objecting because you're making it sound like it's a trivial result.

> you're making this sound like it's a trivial result I don't think this is a trivial result, emulating a highly trained idiot is still very impressive. But it is very different from an untrained genius.

You seem to have very rigid and boring definitions of the words "idiot" and "genius". The "AI effect" is real: https://en.wikipedia.org/wiki/AI_effect

Tbh, I don't even know what you're saying.

[edit] OK, I might have misunderstood you. It's not always clear what people mean.

Re: ChatGPT's Chess Elo is 1400

#237

Earlier quoted context omitted.

He literally used the same prompt as the article. Claim: "ChatGPT's Chess Elo is 1400" Reality: ChatGPT gives illegal moves (this happened to article author too), something a 1400 ranked player would never do Result: ChatGPT's rank is not 1400.

> something a 1400 ranked player would never do The fact that rules and articles exist describing what to do if you or your opponent makes an illegal move indicates this is not the case. Humans are also... human. They make mistakes. It may not happen often at 1400, but to say that it'll never happen is preposterous.

fine, fair, "never" was too much. posting link to this comment to not repeat same discussion twice

https://news.ycombinator.com/item?id=35201037

Re: ChatGPT's Chess Elo is 1400

#238

Earlier quoted context omitted.

You don’t get to 1400 like that. The amount of moves it has to literally remember is stupendous.

Nobody who is 1400 plays outright illegal moves.

Does that still hold when the player doesn't have a board in front of them, but just a list of previous moves?

Re: ChatGPT's Chess Elo is 1400

#240

Earlier quoted context omitted.

He literally used the same prompt as the article. Claim: "ChatGPT's Chess Elo is 1400" Reality: ChatGPT gives illegal moves (this happened to article author too), something a 1400 ranked player would never do Result: ChatGPT's rank is not 1400.

No, the author of the article specifically says that the entire move sequence should be supplied to chatGPT each time, not simply the next move. Be very careful when "disproving" an experiment with squinted eyes.

I'm not really sure what to say here. Both the parent commenter and the author of the article had issues with ChatGPT supplying illegal moves. Both methods resulted in this. It sort of doesn't matter how we're trying to establish that it's a 1400 level player, there's no defined correct way to do this. Regardless of method we've disproven it's a 1400 level player due to these illegal moves.
Post reply on HN