Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

711–720 of 819 posts

Re: Claude Sonnet 4.5

#711

It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?

> It was a very simple question about something very well documented (Oban timeouts).

It's some 3rd party thing for Elixir, a niche within a niche. I wouldn't expect an LLM to do well there.

> I doubt LLM benchmarks more and more, what are they even testing?

Probably testing by asking it to solve a problem with python or (java|type)script. Perhaps not even specifying a language and watching it generate a generic React application.

Re: Claude Sonnet 4.5

#712

> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left…

Have the released the code for this? Does it work? or are there x number of caviets and excuses. I'm kinda of sick of them (and others) getting a free pass at saying stuff like this.

They don't seem to link any source code or demo. They could have run Claude for 10 hours to write thousands of the verge articles as well.

Re: Claude Sonnet 4.5

#713

> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left…

“30 hours of unattended work” is totally vague and it doesn’t mean anything on its own. It - at the very least - highly depends on the amount of tokens you were able to process. Just to illustrate, say you are running on a slow machine that outputs 1 token per hour. At that speed you would produce approximately one sentence.

"Slack clone" is also super vague:

(First of all: Why would anyone in their right mind want a Slack clone? Slack is a cancer. The only people who want it are non-technical people, who inflict it upon their employees.)

Is it just a chat with a group or 1on1 chat? Or does it have threads, emojis, voice chat calls, pinning of messages, all the CSS styling (which probably already is 11k lines or more for the real Slack), web hooks/apps?

Also, of course it is just a BS announcement, without honesty, if they don't publish a reproducible setup, that leads to the same outcome they had. It's the equivalent of "But it worked on my machine!" or "scientific" papers that prove anti gravity with superconductors and perpetuum mobile infinite energy, that only worked in a small shed where some supposed physics professor lives.

Re: Claude Sonnet 4.5

#714
post #711

It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?

> It was a very simple question about something very well documented (Oban timeouts). It's some 3rd party thing for Elixir, a niche within a niche. I wouldn't expect an LLM to do well there. > I doubt LLM benchmarks more and more, what are they even testing? Probably testing by asking it to solve a problem with python or (java|type)script. Perhaps not even specifying a language and watching it generate a generic Reac…

Something that is well documented should still perform well, there’s few places to go wrong, compared with something like React where the training data seems to be a cesspool of the worst code imaginable, at least that’s my experience using it for React.

Re: Claude Sonnet 4.5

#715
post #204

I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…

Kinda pointless listening to the opinions of people who've used previews because it's not gonna be the same model you'll experience once it gets downgraded to be viable under mass use and the benchmarks influencers use are all in the training data now and tested internally so any sort of testing like pelicans on bikes is just PR at this point.

Re: Claude Sonnet 4.5

#716
post #96

Earlier quoted context omitted.

Just a few months ago people were still talking about exponential progress. The fact that we’re already going for just linear progress is not a good sign

Linear growth on a 0-100 benchmark is quite likely an exponential increase in capability.

This got me thinking - is there any reasonable metric we could use to measure the intellectual capabilities of the most capable species on Earth that had evolved at each point in time? I wonder what kind of growth function we'd see.

Silly idea - is there an inter-species game that we could use in order to measure ELO?

Re: Claude Sonnet 4.5

#717

Earlier quoted context omitted.

You’re overlooking the fact that it still says that when you are, in reality, absolutely wrong .

That’s not the purpose of it, as I understand it; it’s a token phrase generated to cajole it down a particular path.[1] An alignment mechanism. The complement appears to be, “actually, that’s not right.”, a correction mechanism. 1: https://news.ycombinator.com/item?id=45137802

But the there’s also the negative psychological impact on the user having the model so strongly agree with them all the time. —— I cannot be the only one who half expects humans to say this to me all the time now?

Re: Claude Sonnet 4.5

#718
post #711

Earlier quoted context omitted.

> It was a very simple question about something very well documented (Oban timeouts). It's some 3rd party thing for Elixir, a niche within a niche. I wouldn't expect an LLM to do well there. > I doubt LLM benchmarks more and more, what are they even testing? Probably testing by asking it to solve a problem with python or (java|type)script. Perhaps not even specifying a language and watching it generate a generic Reac…

Something that is well documented should still perform well, there’s few places to go wrong, compared with something like React where the training data seems to be a cesspool of the worst code imaginable, at least that’s my experience using it for React.

Sure, I'm just answering your question of what people are benchmarking and it's not elixir. You could be the person that benchmarks LLMs in niche languages and shows how bad they are at it.

If your benchmark suite became popular enough and folks referenced it, the people training the LLMs would most likely try to make the model better at those languages.

Re: Claude Sonnet 4.5

#719
post #707
post #674

Earlier quoted context omitted.

> rails 1 codebase to rails 8 A bit off topic, but Rails *1* ? I hope this was an internal app and not on the public internet somewhere …

haha no it's an old (15years old) abandoned enterprise app running on-prem that hasn't seen updates in more than a decade.

Wow Rails 3 came out 15 years ago, so that thing started life out of date.

Re: Claude Sonnet 4.5

#720
Sonnet 4 had turned to shit recently (about 2.5 months according to my observations). It hallucinated on 3 questions in a row while looking at a simple bash script. Was enough for me to cancel. Claude biz is killing Claude dev. It was good while they were not so stingy on GPU.
Post reply on HN