Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

231–240 of 819 posts

Re: Claude Sonnet 4.5

#231

Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…

> based on your actual work, include some specifics of the task and prompt.

Can't show prompts and actual, real work, because, well, it's confidential, and I'd like to get a paycheck instead of a court summons sometime in the next two weeks.

Generally, 'I can't show you the details of my work' isn't a barrier in communicating about tech, because you can generalize and strip out the proprietary bits, but because LLM behavior is incredibly idiosyncratic, by the time you do that, you're no longer accurately communicating the problem that you're having.

Re: Claude Sonnet 4.5

#232
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

> I told it to Give me a zip file of everything you have done so far—you can explore the contents of the file it made me in this Gist.

For those who don't have time to dig into the gist, did it work and do a good job? I assume yes to at least nominally working or you would have mentioned that, but any other thoughts on the solution it produced?

Re: Claude Sonnet 4.5

#233

Interesting, in the new 2.0.0 claude code they got rid of the "Plan with Opus then switch to Sonnet" feature. I hope they're correct in Sonnet being good enough to Plan too, because i quite preferred Opus planning. It wasn't necessarily "better", just more predictable in my experience. Also as a Max $200 user, feels weird to be paying for an Opus tailored sub when now the standard Max $100 would be preferred since th…

I'm also a max user and I just _leave_ it on Opus 4.1 - I've never hit a rate limit.

Same, it very quickly says "approaching rate limits", and then just keeps going forever.

Re: Claude Sonnet 4.5

#234
post #74

Is "parallel test time compute" available in claude code or the api? Or is it something they built internally for benchmark scores?

I'm wondering if it's not just: spawn multiple time the same prompt and take the best

Re: Claude Sonnet 4.5

#235

Interesting, in the new 2.0.0 claude code they got rid of the "Plan with Opus then switch to Sonnet" feature. I hope they're correct in Sonnet being good enough to Plan too, because i quite preferred Opus planning. It wasn't necessarily "better", just more predictable in my experience. Also as a Max $200 user, feels weird to be paying for an Opus tailored sub when now the standard Max $100 would be preferred since th…

In the same boat and ready to downgrade. But this must be on their radar, or they were/are losing money with opus...

Re: Claude Sonnet 4.5

#236

Earlier quoted context omitted.

> I'm infinity times more productive, as I just wouldn't even start projects without the LLM to sidestep my ADHD. 1 is not infinitely greater than 0.

It... literally is? Or otherwise, can you share what you think the ratio is?

No, 1 is 1 more than 0. There’s a certain sense in which you could say that 1 is infinitely greater than 0, but only in an abstract, unquantifiable way. In this case, it doesn’t make sense to say you’re “infinitely more productive” because you’re producing something rather than nothing.

Re: Claude Sonnet 4.5

#238
post #164

Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…

this is a great copypasta

I was thinking the same. Way too perfect to not be spammed around forever.

Re: Claude Sonnet 4.5

#239

Earlier quoted context omitted.

How do you measure 3x sustained output increase? Is it number of lines? Tickets closed? PRs opened or merged? Number of happy customers?

All these are useless metrics. It doesn't say anything meaningful on the quality of your life. I would be more interested in knowing if he can now retire in next 5 years instead of waiting another 15? Or do he now just just get to work for 2 hours and enjoy the remaining 6 hours doing meaningful things apart from staring at a screen.

Not everyone hates their job and gets no satisfaction from it. Some of us relish doing something useful and getting paid for it.

Re: Claude Sonnet 4.5

#240

I just ran this through a simple change I’ve asked Sonnet 4 and Opus 4.1, and it fails too. It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I worry everyone is chasing benchmarks to the detriment of general performance. Or the next token weight for the incorrect c…

> It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I don't understand why this kind of thing is useful. Do the thing yourself and move on. For every one problem like this, AI can do 10 better/faster than I can.

You dont understand how complete unreliability is a problem?

So instead of just "doing things" you want a world where you try it ai-way, fail, then "do thing" 47 times in a row, then 3 ai-way saved you 5 minutes. Then 7 ai-way fail, then try to remember hmm did this work last time or not? ai-way fails another 3 times. "do thing" 3 times. How many ai-way failed today? oh it wasted 30% of the day and i forget which ways worked or not, i better start writing that all down. Lets call it the MAGIC TOME of incantations. oh i have to rewrite the tome again the model changed

Post reply on HN