Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

751–760 of 819 posts

Re: Claude Sonnet 4.5

#751
post #595

> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left…

> Apparently they told it to build a Slack clone and left it unattended for 30 hours, and it built a Slack clone using 11,000 lines of code it's going to be an issue I think, now that lots of these agents support computer use, we are at the point where you can install an app, tell the agent you want something that works exactly the same and just let it run until it produces it. The software world may find it's got mo…

It has been trivial to build a clone of most popular services for years, even before LLMs. One of my first projects was Miguel Grinberg's Flask tutorial, in which a total noob can build a Twitter clone in an afternoon.

What keeps people in are network effects and some dark patterns like vendor lock-in and data unportability.

Re: Claude Sonnet 4.5

#752
So far the only thing I’ve noticed is that it made me confirm that it should do a 10 minute task “manually” because it “would take 2 or 3 hours”

It was a context merging task for my unorganized collection of agents… it sort of made sense, but was the exact reason I was asking it to do it… like you’re the bot, lol

Re: Claude Sonnet 4.5

#753
post #752

So far the only thing I’ve noticed is that it made me confirm that it should do a 10 minute task “manually” because it “would take 2 or 3 hours” It was a context merging task for my unorganized collection of agents… it sort of made sense, but was the exact reason I was asking it to do it… like you’re the bot, lol

Excited to try Claude agents sdk though

Re: Claude Sonnet 4.5

#754

Earlier quoted context omitted.

> I worry everyone is chasing benchmarks to the detriment of general performance. I've been worried about this for a while. I feel like Claude in particular took a step back in my own subjective performance evaluation in the switch from 3.7 to 4, while the benchmark scores leaped substantially. To be fair, benchmarking has always been the most difficult problem to solve in this space, so it's not surprising that benc…

Not that it was better at programming, but I really miss Sonnet 3.5 for educational discussions. I've sometimes considered that what I actually miss was the improvement 3.5 delivered over other models at that time. Though since my system message for Sonnet since 3.7 has been primarily instructing it to behave like a human and have a personality, I really think we lost something.

I still use 3.5 today in Cursor. It's still the best model they've produced for my workflow. It's twice as fast as 4 and doesn't vomit pointless comments all over my code.

Re: Claude Sonnet 4.5

#755
post #411

Earlier quoted context omitted.

Why is this getting downvoted? It was hilarious! I actually added a fun thing to my user-wide CLAUDE.md, basically saying that it should come up with a funny insult every time I come up with an idea that wasn't technically sound (I got the prompt from someone else). It seems to be disobeying me, because I refuse to believe that I don't have bad ideas. Or some other prompt is overriding it.

This is brilliant! Can you give me some pointers? Ie : if I make a request that seems dumb tell me custom instruction?

This is the prompt (I copied it verbatim either from Reddit or HN, don't remember, sorry to the original author for the misattribution):

> Never compliment me or be affirming excessively (like saying "You're absolutely right!" etc). Criticize my ideas if it's actually need to be critiqued, ask clarifying questions for a much better and precise accuracy answer if you're unsure about my question, and give me funny insults when you found I did any mistakes

I just realized in re-reading it that it's written by someone for whom English is a second language. I'll try to rewrite it and see if it works better.

I have it in my ~/.claude/CLAUDE.md. But it still has never done that.

Re: Claude Sonnet 4.5

#756

I haven't shouted into the void for a while. Today is as good a day as any other to do so. I feel extremely disempowered that these coding sessions are effectively black box, and non-reproducible. It feels like I am coding with nothing but hopes and dreams, and the connection between my will and the patterns of energy is so tenuous I almost don't feel like touching a computer again. A lack of determinism comes from m…

> A lack of determinism comes from many places, but primarily: 1) The models change 2) The models are not deterministic... models themselves are deterministic, this is a huge pet peeve of mine, so excuse the tangent, but the appearance of nondeterminism comes from a few sources, but imho can be largely attributed to the probabilistic methods used to get appropriate context and enable timely responses. here's an examp…

The previous poster is correct for a very slightly different definition of the word "model". In context, I would even say their definition is the more correct one.

They are including the random sampler at the end of the LLM that chooses the next token. You are talking about up to, but not including, that point. But that just gives you a list of possible output tokens with values ("probabilities"), not a single choice. You can always just choose the best one, or you could add some randomness that does a weighted sample of the next token based on those values. From the user's perspective, that final sampling step is part of the overall black box that is running to give an output, and it's fair to define "the model" to include that final random step.

Re: Claude Sonnet 4.5

#757
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

Your notes on 4.5 were very interesting, but you asked it a question that only you/someone who already knows the code could ask. I don't though, so I asked it at a higher level: Claude, add tree-structured conversations to https://github.com/simonw/llm. Claude responded with a whole design doc, starting with database schema change (using the same column name even!). https://claude.ai/share/f8f0d02a-3bc1-4b48-b8c7-aa75d6f55021 As I don't know your code, that design doc looks cromulent, but you'd have to read it for yourself to decided how well it did with that higher level of ask.

Re: Claude Sonnet 4.5

#758
post #588

Earlier quoted context omitted.

> As I understand it, OpenAI and Anthropic have both wisely decided to make sure he has up to date info because they know he'll write about it. And the wisest part is if he writes something they don't like, they can cut off that advanced access. As is the longstanding tradition in games journalism, travel journalism, and suchlike.

If they do that I'll go back to writing about them after they ship. Not a big loss for me at all.

I get it, you would trust yourself if you said that, but it doesn't really matter whether you say that or not, what counts for your ongoing credibility if you will preface every future blog post with, whether you got special access, a special deal, sponsorship, or the fact that you didn't get any of those things.

You're a reviewer. This is how reviewers stay credible. If you don't disclose your relationship with the thing or company you're reviewing, I'm probably better off assuming you're paid.

And if your NDA says you can't write that in your preface, then logically, it is impossible to write a credible review in the first place.

Re: Claude Sonnet 4.5

#759
post #743

Earlier quoted context omitted.

Kinda pointless listening to the opinions of people who've used previews because it's not gonna be the same model you'll experience once it gets downgraded to be viable under mass use and the benchmarks influencers use are all in the training data now and tested internally so any sort of testing like pelicans on bikes is just PR at this point.

I learned that lesson from GPT-5, where the preview was weeks long and the models kept changing during that period. This Claude preview lasted from Friday to Monday so I was less worried about major model changes. I made sure to run the pelican benchmark against the model after 10am on Monday (the official release date) just to be safe. The only thing I published that I ran against the preview model was the Claude co…

Well, if they produced a really really really good image for pelicans on bicycles and nothing else, then their cheating would be obvious, so it makes sense to cheat just a little bit, across the board (if we want to assume they're cheating).
Post reply on HN