Live data from Hacker News

Opus 4.5 is not the normal AI agent experience that I have had thus far

burkeholland.github.io

811–820 of 1001 posts

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#811
post #591

Earlier quoted context omitted.

Opus 4.5 is writing code that Opus 5.0 will refactor and extend. And Opus 5.5 will take that code and rewrite it in C from the ground up. And Opus 6.0 will take that code and make it assembly. And Opus 7.0 will design its own CPU. And Opus 8.0 will make a factory for its own CPUs. And Opus 9.0 will populate mars. And Opus 10.0 will be able to achieve AGI. And Opus 11.0 will find God. And Opus 12.0 will make us a time…

Just one more OPUS bro.

Honestly the scary part is that we don’t really even need one more Opus. If all we had for the rest of our lives was Opus 4.5, the software engineering world would still radically change.

But there’s no sign of them slowing down.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#812
post #650
post #601

Earlier quoted context omitted.

LLM's are good at making stuff from scratch and perfect when you don't have to worry about the codes future. 'Research' can be a great tool. But LLMs are horrible in big codebases and multiple micro services. Also at making decision, never let it make a decision for you. You need to know what's happening and you can't ship straight AI code. It can save time, but it's not a lot and it won't replace anyone.

Are you saying this from experience? We have a large monorepo at my company. You're right that for adding entirely new core concepts to an existing codebase we wouldn't give an AI some vague requirements and ask it to build something – but we wouldn't do that for a human engineer either. Typically we would discuss as a team and then once we've agreed on technologies and an approach someone will implement it relying h…

> Are you saying this from experience?

Yes. I mostly work on Quarkus microservices and use cursor with auto agent mode.

> we wouldn't give an AI some vague requirements and ask it to build something > we would discuss as a team

seems like a reasonable workflow. It's the polar opposite of what was written in the blog post. That is the usual, easy way people use agents and what I think is the wrong path. May I also ask what language and/or framework you work with where so much context works good enough?

> Asking AI to explain code and help me learn how it works means I can pick up new systems significantly quicker.

Summarization is generaly a great task for LLMs

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#813

Most software engineers are seriously sleeping on how good LLM agents are right now, especially something like Claude Code. Once you’ve got Claude Code set up, you can point it at your codebase, have it learn your conventions, pull in best practices, and refine everything until it’s basically operating like a super-powered teammate. The real unlock is building a solid set of reusable “skills” plus a few agents for th…

I made a similar comment on a different thread, but I think it also fits here: I think the disconnect between engineers is due to their own context. If you work with frontend applications, specially React/React Native/HTML/Mobile, your experience with LLMs is completely different than the experience of someone working with OpenGL, io_uring, libev and other lower level stuff. Sure, Opus 4.5 can one shot Windows utilit…

We have an in-house, Rust-based proxy server. Claude is unable to contribute to it meaningfully outside of grunt work like minor refactors across many files. It doesn't seem to understand proxying and how it works on both a protocol level and business logic level.

With some entirely novel work we're doing, it's actually a hindrance as it consistently tells us the approach isn't valid/won't work (it will) and then enters "absolutely right" loops when corrected.

I still believe those who rave about it are not writing anything I would consider "engineering". Or perhaps it's a skill issue and I'm using it wrong, but I haven't yet met someone I respect who tells me it's the future in the way those running AI-based companies tell me.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#814
I've only started but I mostly use Claude Code for building out code that has been done a million times. So its good at setting up a project to get all the boiler plate crap out of the way.

When you need to build out specific feature or logic, it can fail hard. And the best is when you have something working, and it fixes something else and deletes the old code that was working, just in a different spot.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#815

I'm tired of constantly debating the same thing again and again. Where are the products? Where is some great performing software all LLM/agent crafted? All I see is software bloatness and decline. Where is Discord that uses just a bunch of hundreds megs of ram? Where is unbloated faster Slack? Where is the Excel killer? Fast mobile apps? Browsers and the web platform improved? Why Cursor team don't use Cursor to get…

Could someone explain this to me? I have the same question: why Cursor team don't use Cursor to get rid of vscode base and code its super duper code editor?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#816

Every time I see a post like this on HN I try again and every time I come to the same conclusion. I have never see one agent managing to pull something off that I could instantly ship. It still ends up being very junior code. I just tried again and ask Opus to add custom video controls around ReactPlayer. I started in Plan mode which looked overal good (used our styling libs, existing components, icons and so on). I…

Back in the day when you found a solution to your problem on Stackoverflow, you typically had to make some minor changes and perhaps engage in some critical thinking to integrate it into your code base. It was still worth looking for those answers, though, because it was much easier to complete the fix starting from something 90% working than 0%. The first few times in your career you found answers that solved your p…

> it was much easier to complete the fix starting from something 90% working than 0%.

As an expert now though, it is genuinely easier and faster to complete the work starting from 0 than to modify something junky. The realplayer example above I could do much faster, correctly, than I could figure out what the AI code was trying to do with all the effects and refactor it correctly. This is why I don't use AI for programming.

And for the cases where I'm not skilled, I would prefer to just gain skill, even though it takes longer than using the AI.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#817

Earlier quoted context omitted.

Yes agreed, and tbh even if that thesis is wrong, what does it matter?

in my experience, what happens is the code base starts to collapse under its own weight. it becomes impossible to fix one thing without breaking another. the coding agent fails to recognize the global scope of the problem and tries local fixes over and over. progress gets slower, new features cost more. all the same problems faced by an inexperienced developer on a greenfield project! has your experience been otherwi…

No that has certainly been my experience, but what is going to be the forcing function after a company decides it needs less engineers to go back to hiring?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#818

Earlier quoted context omitted.

UBI (from taxing big tech) and retraining. In the U.S they'll have enough money to do this and it will still suck and many people won't recover the extreme loss of status and income (after we've been told our income and status are the most important things in life it's gonna be very hard for people to adapt to the loss of it). Countries like India and Philipines and Ukraine which are basically knowledge support hub w…

Also, time to tax for AI use. Introduce AI usage disclosures for corporations. If a company's AI usage is X, they should pay Y tax because that effectively means they didn't employ Z people instead and the society has to take care of them via unemployment benefits and what not. The more the AI usage, higher the tax percentage on a sliding scale.

You're right. But you know what they'll do - they'll offshore those "jobs" e.g token usage to countries that are A.I friendly or that can be bribed easily and do whatever they have to do to fight it out in courts for a decade or as long as it takes. Or am I being pessimistic here?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#819
post #755

It’s a bit strange how anecdotes have become acceptable fuel for 1000 comment technical debates. I’ve always liked the quote that sufficiently advanced tech looks like magic, but its mistake to assume that things that look like magic also share other properties of magic. They don’t. Software engineering spans over several distinct skills: forming logical plans, encoding them in machine executable form(coding), making…

> It’s a bit strange how anecdotes have become acceptable fuel for 1000 comment technical debates.

Progress is so fast right now anecdotes are sometimes more interesting than proper benchmarks. "Wow it can do impressive thing X" is more interesting to me than a 4% gain on SWE Verified Bench.

In early days of a startup "this one user is spending 50 hours/week in our tool" is sometimes more interesting than global metrics like average time in app. In the early/fast days, the potential is more interesting than the current state. There's work to be done to make that one user's experience apply to everyone, but knowing that it can work is still a huge milestone.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#820

Most software engineers are seriously sleeping on how good LLM agents are right now, especially something like Claude Code. Once you’ve got Claude Code set up, you can point it at your codebase, have it learn your conventions, pull in best practices, and refine everything until it’s basically operating like a super-powered teammate. The real unlock is building a solid set of reusable “skills” plus a few agents for th…

> (used voice to text then had claude reword, I am lazy and not gonna hand write it all for yall sorry!)

take my downvote as hard as you can. this sort of thing is awfully off-putting.

Post reply on HN