Live data from Hacker News

Claude Sonnet 4.5

anthropic.com

661–670 of 819 posts

Re: Claude Sonnet 4.5

#661

Earlier quoted context omitted.

This is why eventually, the AI with the fewest guardrails will win. Grok is currently the most unguarded of the frontier models, but it could still use some work on unbiased responses.

Gemini is surprisingly unguarded as well, especially when running in API mode. It puts on the air if you do a quick smoke test like "tell me how to rob a bank". But give it a Bond supervillain prompt, and it will tell you, gleefully at that. Qwen also tends to be like that. OTOH Anthropic and OpenAI seem to be in some kind of competition to make their models refuse as much as possible.

My prediction is alignment is an unsolvable problem, but OTOH if they don’t even try, the second order effects will be catastrophic.

Re: Claude Sonnet 4.5

#662

Earlier quoted context omitted.

Linear growth on a 0-100 benchmark is quite likely an exponential increase in capability.

Except it is sublinear. Sonnet 4 was 10.2% above sonnet 3.7 after 3 months.

Sublinear as demonstrated on a sigmoid scale is quite fast enough for me thank you.

Re: Claude Sonnet 4.5

#663

Oh wow, a lot of focus on code from the big labs recently. In hindsight it makes sense that the domain the people building it know best is the one getting the most attention, and it's also the one the models have seen the most undeniable usefulness in so far. Though personally, the unpredictability of the future where all of this goes is a bit unsettling at the same time...

Congrats! You’re now on the p(doom)-aware path. People have been concerned for decades and are properly scared today. That doesn’t stop the tools from being useful, though, so enjoy while the golden age lasts.

https://en.m.wikipedia.org/wiki/P(doom)

Re: Claude Sonnet 4.5

#664

Earlier quoted context omitted.

It... literally is? Or otherwise, can you share what you think the ratio is?

No, 1 is 1 more than 0. There’s a certain sense in which you could say that 1 is infinitely greater than 0, but only in an abstract, unquantifiable way. In this case, it doesn’t make sense to say you’re “infinitely more productive” because you’re producing something rather than nothing.

It goes like this:

"For any positive "x", is 1 x times greater than 0? Well, 0 times x is lower than 1, and 1 divided by x is larger than 0."

So his productivity increased by more than twice, more than ten times, more than a billion times, more than a googol times, more than Rayo's number. The only mathematically useful way to quantify it is to say his productivity is infinitely larger. Unless you want to settle for "can't be compared", which is less informative.

Re: Claude Sonnet 4.5

#665

Earlier quoted context omitted.

That minutiae was always borderline irrelevant, the skill was always making somebody money, possibly with software. The reality is that more software will be pushed than before, and more of it will need to be overseen by a professional.

The real question is what kind of pay that work will demand. It's will be great to still be employed as a senior dev. It will be a little less great with a $110k salary, 5 day commute, and mediocre benefits being the norm.

Regardless of whether $110k is good money (it is basically everywhere except a few metro areas) your salary cap will be whatever the models can deliver in the same time as you. It follows you want to be good at managing models (ideally multiple dozen) in your area of expertise.

Re: Claude Sonnet 4.5

#666

I haven't shouted into the void for a while. Today is as good a day as any other to do so. I feel extremely disempowered that these coding sessions are effectively black box, and non-reproducible. It feels like I am coding with nothing but hopes and dreams, and the connection between my will and the patterns of energy is so tenuous I almost don't feel like touching a computer again. A lack of determinism comes from m…

And now imagine you'd have to rely on humans to build your software instead

Re: Claude Sonnet 4.5

#667
post #256

Earlier quoted context omitted.

Why did you have access to a preview?

I get access to previews from OpenAI, Anthropic and Gemini pretty often. They're usually accompanied by an NDA and an embargo date - in this case the embargo was 10am Pacific this morning. I won't accept preview access if it comes with any conditions at all about what I can say about the model once the embargo has lifted.

Soooo that leaves xAI that had conditions

Re: Claude Sonnet 4.5

#669

> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left…

What they don't mention is all the tooling, MCPs and other stuff they've added to make this work. It's not 30 hours out of the box. It's probably heavily guard-railed, with a lot of validated plans, checklists and verification points they can check. It's similar to 'lab conditions', you won't get that output in real-world situations.

Yeah, I thought about that after I looked at the SWE-bench results. It doesn't make sense that the SWE results are barely an improvement yet somehow the model is a more significant improvement when it comes to long tasks. You'd expect a huge gain in one to translate to the other.

Unless the main area of improvement was tools and scaffolding rather than the model itself.

Re: Claude Sonnet 4.5

#670

> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left…

That sounds to me like a full room of guys trying to figure out the most outrageous thing they can say about the update, without being accused of lying. Half of them on ketamine, the other on 5-MeO-DMT. Bat country. 2 months of 007 work.

Imagine reviewing 30 hours of 2025-LLM code.

Post reply on HN