Live data from Hacker News

Opus 4.5 is not the normal AI agent experience that I have had thus far

burkeholland.github.io

11–20 of 1001 posts

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#14
post #8

Opus 4.5 has become really capable. Not in terms of knowledge. That was already phenomenal. But in its ability to act independently: to make decisions, collaborate with me to solve problems, ask follow-up questions, write plans and actually execute them. You have to experience it yourself on your own real problems and over the course of days or weeks. Every coding problem I was able to define clearly enough within th…

> In the traditional sense, I haven’t really coded privately at all in recent weeks. Instead, I’ve been guiding and directing, having it write specifications, and then refining and improving them.

This is basically all my side projects.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#15

[flagged]

This is a natural response to software enshittification. You can hardly find an iOS app that is not plagued by ads, subscriptions, or hostile data collection. Now you can have your own small utilities that can work for you. This sort of personal software might be very valuable in the world where you are expected to pay 5$ to click any button.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#18

"Opus 4.5 feels to me like" The article is fine opinion but at what point are we going to either: a) establish benchmarks that make sense and are reliable , or b) stop with the hypecycle stuff?

>establish benchmarks that make sense and are reliable

How aren't current LLM coding benchmarks reliable?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#19
It's also the feeling I have, opus is not a ground-breaking model by any means.

However, Opus 4.5 is incredible when you give it everything it needs, a direction, what you have versus what you want and it will make it work, really, it will work. The code might me ugly, undesirable, would only work for that one condition, but with futher prompting you can evolve it and produce something that you can be proud of.

Opus is only as good as the user and the tools the user gives to it. Hmm, that's starting to sound kind-of... human...

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#20
See also: a post from a couple days ago which came to the same conclusion that Opus 4.5 is an inflection point above Sonnet 4.5 despite that conclusion being counterintuitive: https://news.ycombinator.com/item?id=46495539

It's hard to say if Opus 4.5 itself will change everything given the cost/latency issues, but now that all the labs will have very good synthetic agentic data thanks to Opus 4.5, I will be very interested to see what the LLMs release this year will be able to do. A Sonnet 4.7 that can do agentic coding as well as Opus 4.5 but at Sonnet's speed/price would be the real gamechanger: with Claude Code on the $20/mo plan, you can barely do more than one or two prompts with Opus 4.5 per session.

Post reply on HN