Live data from Hacker News

Claude Plays Pokémon

twitch.tv

1–10 of 28 posts

Re: Claude Plays Pokémon

#3
This is truly tremendous to watch. Eleven years from TPP, and we're watching the current best-in-class AI try its best at the same. Who'll get there first, the historical gestalt of Twitch users or the just-shy-of-10^26 FLOPS [0] AI model?

Now here's a concept for anyone with more money than sense: ClaudePlaysTwitchPlaysPokemon, where it's TPP but every participant is Claude. Would hivemind AI consensus perform better than a single AI? Anthropic's certainly looking into it! [1]

[0]: https://www.oneusefulthing.org/p/a-new-generation-of-ais-cla...

[1]: https://www.anthropic.com/news/visible-extended-thinking

Re: Claude Plays Pokémon

#6
post #4

It's run by Anthropic! https://x.com/AnthropicAI/status/1894419011569344978

This administrative request for a reset is incredible - I can't help but feel that this is intended as the equivalent of a prompt injection for the person running it. Time to rewatch Ex Machina.

https://x.com/AnthropicAI/status/1894419017756029427?t=xDXk6...

Re: Claude Plays Pokémon

#9
This is neat but watching a reasoning model that stops to consider "I have read half of a dialogue block, time to press A to get the rest of the text" gets old really quick. I think I'd rather watch a model try to play pokemon against human opponents on a simulator like pokemon showdown (which I understand is a bit further in an IP rights grey area than emulating a 30 year old game). In that case you would get to see how it handles unknown information and updates its reasoning based on the success/failure of its predictions.
Post reply on HN