Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

151–160 of 819 posts

Re: Claude Sonnet 4.5

#151
post #17

I'm really interested in the progress on computer use. These are the benchmarks to watch if you want to forecast economic disruption, IMO. Mastery of computer use takes us out of the paradigm of task-specific integrations with AI to a more generic interface that's way more scalable.

What are some standard benchmarks you look at in this space?

Re: Claude Sonnet 4.5

#152
post #5

Looking at the chart here, it seems like Sonnet 4 was already better than GPT-5-codex in the SWE verified benchmark. However, my subjective personal experience was GPT-5-codex was far better at complex problems than Claude Code.

GPT-5 is like the guy on the baseball team that's really good at hitting home runs but can't do basic shit in the outfield. It also consistently gets into drama with the other agents e.g. the other day when I told it we were switching to claude code for executing changes, after badmouthing claude's entirely reasonable and measured analysis it went ahead and decided to `git reset --hard` even after I twice pushed back…

> "another agent"

You could just say it’s another GPT-5 instance.

Re: Claude Sonnet 4.5

#153
These benchmarks in real world work remain remarkably weak. If you're using this for day-to-day work, the eval that really matters is how the model handles a ten step action. Context and focus are absolutely king in real world work. To be fair, Sonnet has tended to be very good at that...

I wonder if the 1m token context length is coming for this ride too?

Re: Claude Sonnet 4.5

#154

Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…

[deleted]

Re: Claude Sonnet 4.5

#156

Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…

I am almost convinced your comment is parody but I am not entirely sure.

You want proof for critical/supportive criticism? Then almost in the same sentence you make an insane claim without backing this up by any evidence.

Re: Claude Sonnet 4.5

#157
I am a paying subscriber to Gemini, Claude and OpenAI.

I don't know if it's me, but over the last few weeks I've got to the conclusion ChatGPT is very strongly leading the race. Every answer it gives me is better - it's more concise and more informative.

I look forward to testing this further, but out of the few runs I just did after reading about this - it isn't looking much better

Re: Claude Sonnet 4.5

#158
post #33
post #14

Price is playing a big role in my AI usage for coding. I am using Grok Code Fast as it's super cheap. Next to it GPT-5 Codex. If you are paying for model use out of pocket Claude prices are super expensive. With better tooling setup those less smart (and often faster) models can give you better results. I am going to give this another shot but it will cost me $50 just to try it on a real project :(

I'm paying $90(?) a month for the Max and it holds up for about an hour or so of in depth coding before it kicks in the 5-hour window lockout (so effectively about 4 hours of time when I can't run it). Kinda frustrating, even with efficient prompt and context length conservation techniques. I'm going to test this new sonnet 4.5, now but it'll probably be just as quick to gobble my credits.

Do you normally run Opus by default? It seems the Max subscription should let you run Sonnet in an uninterrupted way, so it was surprising to read.

Re: Claude Sonnet 4.5

#159

Earlier quoted context omitted.

How do you measure 3x sustained output increase? Is it number of lines? Tickets closed? PRs opened or merged? Number of happy customers?

Oh good, a new discussion point that we haven't heard 1000x on here. Have you heard of that study that shows AI actually makes developers less productive, but they think it makes them more productive?? EDIT: sorry all, I was being sarcastic in the above, which isn't ideal. Just annoyed because that "study" was catnip to people who already hated AI, and they (over-) cite it constantly as "evidence" supporting their pr…

> Have you heard of that study that shows AI actually makes developers less productive, but they think it makes them more productive??

Have you looked into that study? There's a lot wrong with it, and it's been discussed ad nauseam.

Also, what a great catch 22, where we can't trust our own experiences! In fact, I just did a study and my findings are that everyone would be happier if they each sent me $100. What's crazy is that those who thought it wouldn't make them happier, did in fact end up happier, so ignore those naysayers!

Re: Claude Sonnet 4.5

#160

Earlier quoted context omitted.

You are making assumptions about someone you have never talked to in the past, and don't know anything about. Of the two of you, I know which one I'd bet on being "right". (Hint: It's the one talking about their own experience, not the one supplanting theirs onto someone else)

[flagged]

We are reading very different "negative" comments here.

Most of the anti-AI comments I see on HN are NOT a version of "the problem with AI is that it's so good it's going to replace me!"

Post reply on HN