Live data from Hacker News

ChatGPT's Chess Elo is 1400

dkb.blog

351–360 of 361 posts

Re: ChatGPT's Chess Elo is 1400

#351
post #227

Earlier quoted context omitted.

They are disproving an assertion. Demonstrating that an alternate approach implodes the assertion is a perfectly acceptable route, especially when the original approach was cherry-picking successes and throwing out failures. I wish I could just make bullshit moves and get a higher chess ranking. Sounds nice.

I disagree. If there is a procedure for getting ChatGPT to play chess accurately and you discard that and do some naive approach as a way of disproving the article, doesn't sound to me like you have disproven anything. I dont understand the point of your second sentence, seems to be entirely missing the substance of the conversation.

The gymnastics you GPT True Believers go through to make this stuff "work" are really something else.

By the way - definitely read the article. But once again - I thought the methodology was bad, and thus the conclusion was bad.

Re: ChatGPT's Chess Elo is 1400

#352

Earlier quoted context omitted.

Personally I think the illegal moves are irreverent, the fact that it doesn't play exactly like a typical 1400 doesn't mean it can't have a 1400 rating. Rating is purely determined by wins and losses against opponents, it doesn't matter if you lose a game by checkmate, resignation, or playing an illegal move. That's not to say ChatGPT can play at 1400, just that that playing in an odd way doesn't determine its rating…

This is like saying I play at a 2900 level if you just ignore all the times I lose.

The article does not ignore the losses. In fact, it used a rule stricter than FIDE rules to trigger losses on illegal moves.

Re: ChatGPT's Chess Elo is 1400

#353
post #351

Earlier quoted context omitted.

I disagree. If there is a procedure for getting ChatGPT to play chess accurately and you discard that and do some naive approach as a way of disproving the article, doesn't sound to me like you have disproven anything. I dont understand the point of your second sentence, seems to be entirely missing the substance of the conversation.

The gymnastics you GPT True Believers go through to make this stuff "work" are really something else. By the way - definitely read the article. But once again - I thought the methodology was bad, and thus the conclusion was bad.

I don’t think this is any crazy level of gymnastics.

But not going to keep replying, you engage online in a way that will turn lots of people you talk to away.

Re: ChatGPT's Chess Elo is 1400

#354

Earlier quoted context omitted.

I'm not really sure what to say here. Both the parent commenter and the author of the article had issues with ChatGPT supplying illegal moves. Both methods resulted in this. It sort of doesn't matter how we're trying to establish that it's a 1400 level player, there's no defined correct way to do this. Regardless of method we've disproven it's a 1400 level player due to these illegal moves.

The #1 misconception when working with large language models is thinking that a capability is a property of the model, rather than the model + input. It may be simultaneously true that ChatGPT has an elo of 100 when given a conversational message and an elo of 1400 when given an optimized message (e.g., strings that resemble chess games, with many examples present in the conversation). Understanding this concept is c…

[deleted]

Re: ChatGPT's Chess Elo is 1400

#355
post #212

Earlier quoted context omitted.

You're "disproving" the article by doing things differently to how the article did. If you're going to disprove that the method given in the article does as well as the article claims at least use the same method.

It’s super scary how ChatGPT brings out people who are veeeery good at seeing the Emperor’s clothes.

You know, I didn't remember the story very well so I checked wikipedia. Here's what it says about the (start of) the plot:

>> Two swindlers arrive at the capital city of an emperor who spends lavishly on clothing at the expense of state matters. Posing as weavers, they offer to supply him with magnificent clothes that are invisible to those who are stupid or incompetent. The emperor hires them, and they set up looms and go to work. A succession of officials, and then the emperor himself, visit them to check their progress. Each sees that the looms are empty but pretends otherwise to avoid being thought a fool.

So everyone "pretends otherwise to avoid being thought a fool".

Huh. I guess that explains it. Good metaphor.

Re: ChatGPT's Chess Elo is 1400

#356
post #350

Earlier quoted context omitted.

He claims he was forfeiting every time he got an illegal move. Does no one on this website actually read the article? Whether any of it is actually true is a different question.

And it has already been stated elsewhere in the thread: an illegal move is not technically a forfeiture, so this is some heavy "giving the benefit of the doubt".

It would be interesting to see how ChatGPT would play after making the first illegal move. Would it go off the rails completely, playing an impossible game? Would it be able to play well if its move was corrected (I'm not sure how illegal moves are treated in chess; are they allowed to be taken back if play hasn't progressed?). Could it figure out it made an illegal move, if it was told it did, without specifying which one, or why it was illegal? By stopping the game as soon as an illegal move is made, the author is missing the chance to understand an important aspect of ChatGPT's ability to play chess.

I got the impression the author did this because they thought they were being fair with ChatGPT, but they're much more likely to be letting it off the hook than they seem to realise.

(Sorry about the "they"'s; I think the author is a guy but wasn't sure).

Re: ChatGPT's Chess Elo is 1400

#357

Earlier quoted context omitted.

Personally I think the illegal moves are irreverent, the fact that it doesn't play exactly like a typical 1400 doesn't mean it can't have a 1400 rating. Rating is purely determined by wins and losses against opponents, it doesn't matter if you lose a game by checkmate, resignation, or playing an illegal move. That's not to say ChatGPT can play at 1400, just that that playing in an odd way doesn't determine its rating…

This is like saying I play at a 2900 level if you just ignore all the times I lose.

No it's not, we're not ignoring losses or illegal moves at all, they are counted as losses and that's how you arrive at 1400.

It's a (theoretically) 1400 player which plays significantly better then 1400 when it knows the lines, but makes bad or illegal moves when it doesn't, and that play averages out to be around your typical 1400 player. Functionally is just what a 1400 player already is, but with higher extremes and lower lows.

Re: ChatGPT's Chess Elo is 1400

#358
post #351

Earlier quoted context omitted.

The gymnastics you GPT True Believers go through to make this stuff "work" are really something else. By the way - definitely read the article. But once again - I thought the methodology was bad, and thus the conclusion was bad.

I don’t think this is any crazy level of gymnastics. But not going to keep replying, you engage online in a way that will turn lots of people you talk to away.

I'll admit to having mistook your reply with another (hence the non-sequitur second half of my comment.) Apologies for my brusque tone.

Re: ChatGPT's Chess Elo is 1400

#359

Better than me then. But does it give credit to who taught it. These models are basically a scrape of the best of humankind and a claim that it's their own.

Do you give credit to people you've played in the past when you play a game of chess?

It would be difficult to remember or keep record of them but sure, if I'm learning from someone- I'll remember that.

Re: ChatGPT's Chess Elo is 1400

#360

Earlier quoted context omitted.

Do you give credit to people you've played in the past when you play a game of chess?

It would be difficult to remember or keep record of them but sure, if I'm learning from someone- I'll remember that.

That doesn't seem to be any more noteworthy than saying OpenAI knows what's in the corpus.
Post reply on HN