Live data from Hacker News

Claude 4

anthropic.com

31–40 of 1001 posts

Re: Claude 4

#31

Sooo... it can play Pokemon. Feels like they had to throw that in after Google IO yesterday. But the real question is now can it beat the game including the Elite Four and the Champion. That was pretty impressive for the new Gemini model.

That Google IO slide was somewhat misleading as the maintainer of Gemini Plays Pokemon had a much better agentic harness that was constantly iterated upon throughout the runtime (e.g. the maintainer had to give specific instructions on how to use Strength to get past Victory Road), unlike Claude Plays Pokemon.

The Elite Four/Champion was a non-issue in comparison especially when you have a lv. 81 Blastoise.

Re: Claude 4

#33
post #19

Allegedly Claude 4 Opus can run autonomously for 7 hours (basically automating an entire SWE workday).

Which sort of workday? The sort where you rewrite your code 8 times and end the day with no marginal business value produced?

Re: Claude 4

#37

Sooo... it can play Pokemon. Feels like they had to throw that in after Google IO yesterday. But the real question is now can it beat the game including the Elite Four and the Champion. That was pretty impressive for the new Gemini model.

Gemini can beat the game?

Gemini has beat it already, but using a different and notably more helpful harness. The creator has said they think harness design is the most important factor right now, and that the results don't mean much for comparing Claude to Gemini.

Re: Claude 4

#38
post #19

Allegedly Claude 4 Opus can run autonomously for 7 hours (basically automating an entire SWE workday).

Easy, I can also make a nanoGPT run for 7 hours when inferring on a 68k, and make it produce as much value as I usually do.

Re: Claude 4

#40
Can't wait to hear how it breaks all the benchmarks but have any differences be entirely imperceivable in practice.
Post reply on HN