Live data from Hacker News

Claude 4

anthropic.com

41–50 of 1001 posts

Re: Claude 4

#41
post #19

Allegedly Claude 4 Opus can run autonomously for 7 hours (basically automating an entire SWE workday).

Which sort of workday? The sort where you rewrite your code 8 times and end the day with no marginal business value produced?

Well Claude 3.7 definitely did the one where it was supposed to process a file and it finally settled on `fs.copyFile(src, dst)` which I think is pro-level interaction. I want those $0.95 back.

But I love you Claude. It was me, not you.

Re: Claude 4

#43
My mind has been blown using ChatGPT's o4-mini-high for coding and research (it knowledge of computer vision and tools like OpenCV are fantastic). Is it worth trying out all the shiny new AI coding agents ... I need to get work done?

Re: Claude 4

#45
post #11

> Try Claude Sonnet 4 today with Claude Opus 4 on paid plans. Wait, Sonnet 4? Opus 4? What?

Claude names their models based on size/complexity:

- Small: Haiku

- Medium: Sonnet

- Large: Opus

Re: Claude 4

#46
I'm curious what are others priors when reading benchmark scores. Obviously with immense funding at stakes, companies have every incentive to game the benchmarks, and the loss of goodwill from gaming the system doesn't appear to have much consequences.

Obviously trying the model for your use cases more and more lets you narrow in on actually utility, but I'm wondering how others interpret reported benchmarks these days.

Re: Claude 4

#47

Sooo... it can play Pokemon. Feels like they had to throw that in after Google IO yesterday. But the real question is now can it beat the game including the Elite Four and the Champion. That was pretty impressive for the new Gemini model.

Gemini can beat the game?

2 weeks ago

Re: Claude 4

#48
> Users requiring raw chains of thought for advanced prompt engineering can contact sales

So it seems like all 3 of the LLM providers are now hiding the CoT - which is a shame, because it helped to see when it was going to go down the wrong track, and allowing to quickly refine the prompt to ensure it didn't.

In addition to openAI, Google also just recently started summarizing the CoT, replacing it with an, in my opinion, overly dumbed down summary.

Re: Claude 4

#50
Nice to see that Sonnet performs worse than o3 on AIME but better on SWE-Bench. Often, it's easy to optimize math capabilities with RL but much harder to crack software engineering. Good to see what Anthropic is focusing on.
Post reply on HN