Live data from Hacker News

Claude 4

anthropic.com

51–60 of 1001 posts

Re: Claude 4

#51
I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

Re: Claude 4

#52
post #19

Allegedly Claude 4 Opus can run autonomously for 7 hours (basically automating an entire SWE workday).

I can write an algorithm to run in a loop forever, but that doesn't make it equivalent to infinite engineers. It's the output that matters.

Re: Claude 4

#53
> Extended thinking with tool use (beta): Both models can use tools—like web search—during extended thinking, allowing Claude to alternate between reasoning and tool use to improve responses.

I'm happy that tool use during extended thinking is now a thing in Claude as well, from my experience with CoT models that was the one trick(tm) that massively improves on issues like hallucination/outdated libraries/useless thinking before tool use, e.g.

o3 with search actually returned solid results, browsing the web as like how i'd do it, and i was thoroughly impressed – will see how Claude goes.

Re: Claude 4

#54
I can't think of more boring than marginal improvements on coding tasks to be honest.

I want GenAI to become better at tasks that I don't want to do, to reduce the unwanted noise from my life. This is when I'll pay for it, not when they found a new way to cheat a bit more the benchmarks.

At work I own the development of a tool that is using GenAI, so of course a new better model will be beneficial, especially because we do use Claude models, but it's still not exciting or interesting in the slightest.

Re: Claude 4

#55
When will structured output be available? Is it difficult for anthropic because custom sampling breaks their safety tools?

Re: Claude 4

#56
post #51

I've found myself having brand loyalty to Claude. I don't really trust any of the other models with coding, the only one I even let close to my work is Claude. And this is after trying most of them. Looking forward to trying 4.

I've been initially fascinated by Claude, but then I found myself drawn to Deepseek. My use case is different though, I want someone to talk to.

Re: Claude 4

#57
Anthropic might be scammers. Unclear. I canceled my subscription with them months ago after they reduced capabilities for pro users and I found out months later that they never actually canceled it. They have been ignoring all of my support requests.. seems like a huge money grab to me because they know that they're being out competed and missed the ball on monetizing earlier.

Re: Claude 4

#58
How long will the VScode wrapper (cursor, windsurf) survive?

Love to try the Claude Code VScode extension if the price is right and purchase-able from China.

Re: Claude 4

#59
Anyone with access who could compare the new models with say O1 Pro Mode? Doesn't have to be a very scientific comparison, just some first impressions/thoughts compared to the current SOTA.

Re: Claude 4

#60
post #13

I'll look at it when this shows up on https://aider.chat/docs/leaderboards/ I feel like keeping up with all the models is a full time job so I just use this instead and hopefully get 90% of the benefit I would by manually testing out every model.

Are these just leetcode exercises? What I would like to see is an independent benchmark based on real tasks in codebases of varying size.
Post reply on HN