Live data from Hacker News

Claude Opus 4.5

anthropic.com

421–430 of 525 posts

Re: Claude Opus 4.5

#421

Earlier quoted context omitted.

The models are non-deterministic. You can't just assume that because it did better before that it was on average better than before. And the variance is quite large.

No one talked about determinism. First it was able to do a task, second time not. It’s not that the implementation details changed.

There are many, many tasks that a given LLM can successfully do 5% of the time.

Feeling lucky?

Re: Claude Opus 4.5

#422

Earlier quoted context omitted.

Not a skeptic, I use AI for coding daily and am working on a custom agent setup because, through my experience for more than a year, they are not up to hard tasks. This is well known I thought, as even the people who build the AIs we use talk about this and acknowledge their limitations.

I'm pretty sure at this point more than half of Anthropic's new production code is LLM-written. That seems incompatible with "these agents are not up to the task of writing production level code at any meaningful scale".

how are you pretty sure? What are you basing that on?

If true, could this explain why Anthropics APIs are less reliable than Gemini's? (I've never gotten a service overloaded response from Google like I did from Anthropic)

Re: Claude Opus 4.5

#423

Earlier quoted context omitted.

This is not the way these agents are not up to the task of writing production level code at any meaningful scale looking forward to high paying gigs to go in and clean up after people take them too far and the hype cycle fades --- I recommend the opposite, work on custom agents so you have a better understanding of how these things work and fail. Get deep in the code to understand how context and values flow and get…

You cannot clean up the code, it is too verbose. That said, you can produce production ready code with AI, you just need to put up very strong boundaries and not let it get too creative. Also, the quality of production ready code is often highly exaggerated.

I have AI generated, production quality code running, but it was isolated, not at scale or broad in view / spanning many files or systems

What I mean more is that as soon as the task becomes even moderately sized, these things fail hard

Re: Claude Opus 4.5

#424
post #292

Earlier quoted context omitted.

> played around with You'll never get an accurate comparison if you only play We know by now that it takes time to "get to know a model and it's quirks" So if you don't use a model and cannot get equivalent outputs to your daily driver, that's expected and uninteresting

I rotate models frequently enough that I doubt my personal access patterns are so model specific that they would unfairly advantage one model over another; so ultimately I think all you're saying is that Claude might be easier to use without model-specific skilling than other models. Which might be true. I certainly don't have as much time on Gemini 3 as I do on Claude 4.5, but I'd say my time with the Gemini family…

yeah, this generally vibes with my experience, they aren't that different

As I've gotten into the agentic stuff more lately, I suspect a sizeable part of the different user experiences comes down to the agents and tools. In this regard, Anthropic is probably in the lead. They certainly have become a thought leader in this area by sharing more of their experience and know hows in good posts and docs

Re: Claude Opus 4.5

#425

Earlier quoted context omitted.

This is not the way these agents are not up to the task of writing production level code at any meaningful scale looking forward to high paying gigs to go in and clean up after people take them too far and the hype cycle fades --- I recommend the opposite, work on custom agents so you have a better understanding of how these things work and fail. Get deep in the code to understand how context and values flow and get…

> these agents are not up to the task of writing production level code at any meaningful scale I think the new one is. I could be the fool and be proven wrong though.

It's marginally better, no where close to game changing, which I agree will require moving beyond transformers to something we don't know yet

Re: Claude Opus 4.5

#426

Earlier quoted context omitted.

I thought AI safety was dumb/unimportant until I saw this dataset of dangerous prompts: https://github.com/mlcommons/ailuminate/blob/main/airr_offic... I don't love the idea of knowledge being restricted... but I also think these tools could result in harm to others in the wrong hands

I once heard a devils advocate say, “if child porn can be fully AI generated and not imply more exploitation of real children, and it’s still banned then it’s about control not harm.” Attack away or downvote my logic.

CG CSAM can be used to groom real kids, by making those activities look normal and acceptable.

Re: Claude Opus 4.5

#427

Earlier quoted context omitted.

https://apps.apple.com/us/app/claude-by-anthropic/id64737536... Has a section for code. You link it to your GitHub, and it will generate code for you when you get on the bus so there's stuff for you to review after you get to the office.

Thanks. Still looking for some kind of total code by phone thing.

The app version is iPhone only, you don’t get Code in the Android app, you have to use a web browser.

I use it every day. I’ll write the spec in conversation with the chatbot, refining ideas, saying “is it possible to …?” Get it to create detailed planning and spec documents (and a summary document about the documents). Upload them to Github and then tell Code to make the project.

I have never written any Rust, am not an evangelist, but Code says it finds the error messages super helpful so I get it to one shot projects in that.

I do all this in the evenings while watching TV with my gf.

It amuses me we have people even this thread claiming what it already does is something it can’t do - write working code that does what is supposed to.

I get to spend my time thinking of what to create instead of the minutiae of “ok, I just need 100 more methods, keep going”. And I’ve been coding since the 1980 so don’t think I’m just here for the vibes.

Re: Claude Opus 4.5

#428

The burying of the lede here is insane. $5/$25 per MTok is a 3x price drop from Opus 4. At that price point, Opus stops being "the model you use for important things" and becomes actually viable for production workloads. Also notable: they're claiming SOTA prompt injection resistance. The industry has largely given up on solving this problem through training alone, so if the numbers in the system card hold up under a…

Just on Claude Code, I didn't notice any performance difference from Sonnet 4.5 but if it's cheaper then that's pretty big! And it kinda confuses the original idea that Sonnet is the well rounded middle option and Opus is the sophisticated high end option.

It does, but it also maps to the human world: Tokens/Time cost money. If either is well spent, then you save money. Thus, paying an expert ends up costing less than hiring a novice, who might cost less per hour, but takes more hours to complete the task, if they can do it at all.

It's both kinda neat and irritating, how many parallels there are between this AI paradigm and what we do.

Re: Claude Opus 4.5

#429
post #154

This is gonna be game-changing for the next 2-4 weeks before they nerf the model. Then for the next 2-3 months people complaining about the degradation will be labeled “skill issue”. Then a sacrificial Anthropic engineer will “discover” a couple obscure bugs that “in some cases” might have lead to less than optimal performance. Still largely a user skill issue though. Then a couple months later they’ll release Opus 4…

There are two possible explanations for this behavior: the model nerf is real, or there's a perceptual/psychological shift. However, benchmarks exist. And I haven't seen any empirical evidence that the performance of a given model version grows worse over time on benchmarks (in general.) Therefore, some combination of two things are true: 1. The nerf is psychologial, not actual. 2. The nerf is real but in a way that…

https://www.youtube.com/watch?v=DtePicx_kFY

"There's something still not quite right with the current technology. I think the phrase that's becoming popular is 'jagged intelligence'. The fact that you can ask an LLM something and they can solve literally a PhD level problem, and then in the next sentence they can say something so clearly, obviously wrong that it's jarring. And I think this is probably a reflection of something fundamentally wrong with the current architectures as amazing as they are."

Llion Jones, co-inventor of transformers architecture

Re: Claude Opus 4.5

#430

Earlier quoted context omitted.

> only Claude Code has been able to really solve; Sonnet 4.5 in there consistently performs better than Sonnet 4.5 anywhere else. I think part of it is this[0] and I expect it will become more of a problem. Claude models have built-in tools (e.g. `str_replace_editor`) which they've been trained to use. These tools don't exist in Cursor, but claude really wants to use them. 0 - https://x.com/thisritchie/status/1944038…

This feels like a dumb question, but why doesn't Cursor implement that tool? I built my own simple coding agent six months ago, and I implemented str_replace_based_edit_tool ( https://platform.claude.com/docs/en/agents-and-tools/tool-us... ) for Claude to use; it wasn't hard to do.

Is the code to your agent and its implementation of "str_replace_based_edit_tool" public anywhere? If not, can you share it in a Gist?
Post reply on HN