Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

391–400 of 819 posts

Re: Claude Sonnet 4.5

#391
post #380

Does 4.5 still answer everything with "You're absolutely right!" or is it now able to communicate like a real programmer?

I won’t be satisfied until I get a Linus Torvalds mode.

“Your idea is shit because you are so fucking stupid”

“Please stop talking, it hurts my GPUs thinking down to your level”

“I may seem evil but at least I’m not incompetent”

Re: Claude Sonnet 4.5

#392
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

Why did you have access to a preview?

They are an AI evangelist that told me I can replace any technical book created with an LLM.

They are a nice person.

Re: Claude Sonnet 4.5

#393
post #256

Earlier quoted context omitted.

Why did you have access to a preview?

I get access to previews from OpenAI, Anthropic and Gemini pretty often. They're usually accompanied by an NDA and an embargo date - in this case the embargo was 10am Pacific this morning. I won't accept preview access if it comes with any conditions at all about what I can say about the model once the embargo has lifted.

[flagged]

Re: Claude Sonnet 4.5

#394
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

I am curious how the sandbox handles potentially malicious code. For example, what would happen if someone tried to run something like a crypto miner or a DDoS script?

Re: Claude Sonnet 4.5

#395
post #324

Anecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s de…

Did you use plan mode?

Yes, I did.

I ran the test again, took Claude ~4mins this time. There was no error now with the auth, but the functionality was totally broken. It could not even find the most basic stuff that matches perfectly.

Re: Claude Sonnet 4.5

#396

Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…

"Please include ACTUAL EVIDENCE!"

"I tripled my output (I provide no evidence for this claim)"

Never change, HN.

Re: Claude Sonnet 4.5

#397

Earlier quoted context omitted.

You definitely need some context management like Serena.

Even with Serena and detailed plans crafted by Gemini that lay out file-by-file changes, Claude will sometimes go off the rails. Claude is very task-completion driven, and it's willing to relax the constraints of the task to complete in the face of even slight adversity. I can't tell you the number of times I've had Claude try to install a python computational library, get an error, then either try to hand-roll the a…

[deleted]

Re: Claude Sonnet 4.5

#398

Earlier quoted context omitted.

I had similar experience, not good enough yet to come back for the Claude max plan. Sticking with ChatGPT pro sub and gpt5 codex on high.

do you ever hit your pro quota?

Never hit pro quota yet, huge repo. Have multiple projects on the go locally and in cloud.

Feel like this is going to be thr $1000 plan soon

Re: Claude Sonnet 4.5

#399
post #256

Earlier quoted context omitted.

I get access to previews from OpenAI, Anthropic and Gemini pretty often. They're usually accompanied by an NDA and an embargo date - in this case the embargo was 10am Pacific this morning. I won't accept preview access if it comes with any conditions at all about what I can say about the model once the embargo has lifted.

[flagged]

Why do you even care?

Re: Claude Sonnet 4.5

#400

Earlier quoted context omitted.

Why did you have access to a preview?

Simonw is a cheerful and straightforward AI journalist who likes to show and not just tell. He has done a good job aggregating and documenting the progress of LLM tools and models. As I understand it, OpenAI and Anthropic have both wisely decided to make sure he has up to date info because they know he'll write about it. Thanks for all your work, Simon! You're my favorite journalist in this space and I really appreci…

Simon has a popular blog, but he's also co-creator of Django and very well-known in the Python community.
Post reply on HN