Live data from Hacker News

Opus 4.5 is not the normal AI agent experience that I have had thus far

burkeholland.github.io

481–490 of 1001 posts

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#481

Earlier quoted context omitted.

I think we're entering a world where programmers as such won't really exist (except perhaps in certain niches). Being able to program (and read code, in particular) will probably remain useful, though diminished in value. What will matter more is your ability to actually create things, using whatever tools are necessary and available, and have them actually be useful. Which, in a way, is the same as it ever was. Ther…

Isn't there more indirection as long as LLMs use "human" programming languages?

If you think of the training data, e.g. SO, github etc, then you have a human asking or describing a problem, then the code as the solution. So I suspect current-gen LLMs are still following this model, which means for the forseeable future a human like language prompt will still be the best.

Until such time, of course, when LLMs are eating their own dogfood, in which case they - as has already happened - create their own language, evolve dramatically, and cue skynet.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#482

What bothers me about posts like this is: mid-level engineers are not tasked with atomic, greenfield projects. If all an engineer did all day was build apps from scratch, with no expectation that others may come along and extend, build on top of, or depend on, then sure, Opus 4.5 could replace them. The hard thing about engineering is not "building a thing that works", its building it the right way, in an easily unde…

It just one shots bug fixes in complex codebases.

Copy-paste the bug report and watch it go.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#483
post #191

Earlier quoted context omitted.

>A JavaScript interpreter written in Python? I'm assuming this refers to the python port of Bellard's MQJS [1]? It's impressive and very useful, but leaving out the "based on mqjs" part is misleading. [1] https://github.com/simonw/micro-javascript ?

That's why I built the WebAssembly one - the JavaScript one started with MQJS, but for the WebAssembly one I started with just a copy of the https://github.com/webassembly/spec repo. I haven't quite got the WASM one into a share-able shape yet though - the performance is pretty bad which makes the demos not very interesting.

Isn’t that telling though?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#484
post #371
post #248

Earlier quoted context omitted.

It could be that the people who are focused on building monetizable products with LLMs don't feel the need to share what they are doing - they're too busy quietly getting on with building and marketing their products. Sharing how you're using these tools is quite a lot of work!

What would be more likely, That people making startups is too bussy working to share it on HN or that AI is useless in real projects.

The former.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#485

Earlier quoted context omitted.

> I really think a lof of people tried AI coding earlier, got frustrated at the errors and gave up. That's where the rejection of all these doomer predictions comes from. It's not just the deficiencies of earlier versions, but the mismatch between the praise from AI enthusiasts and the reality. I mean maybe it is really different now and I should definitely try uploading all of my employer's IP on Claude's cloud and…

You enter some text and a computer spits out complex answers generated on the spot Right or wrong - doesn’t matter. You typed in a line of text and now your computer is making 3000 word stories, images, even videos based on it How are you NOT astounded by that? We used to have NONE of this even 4 years ago!

Because I want correct answers.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#486
post #473

Earlier quoted context omitted.

I find Opus 4.5 very, very strong at matching the prevailing conventions/idioms/abstractions in a large, established codebase. But I guess I'm quite sensitive to this kind of thing so I explicitly ask Opus 4.5 to read adjacent code which is perhaps why it does it so well. All it takes is a sentence or two, though.

> ask Opus 4.5 to read adjacent code which is perhaps why it does it so well. All it takes is a sentence or two, though. People keep telling me that an LLM is not intelligence, it's simply spitting out statistically relevant tokens. But surely it takes intelligence to understand (and actually execute!) the request to "read adjacent code".

I used to agree with this stance, but lately I'm more in the "LLMs are just fancy autocomplete" camp. They can just autocomplete increasingly more things, and when they can't, they fail in ways that an intelligent being just wouldn't. Rather that just output a wrong or useless autocompletion.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#487
post #411
post #133

Opus 4.5 ate through my Copilot quota last month, and it's already halfway through it for this month. I've used it a lot, for really complex code. And my conclusion is: it's still not as smart as a good human programmer. It frequently got stuck, went down wrong paths, ignored what I told it to do to do something wrong, or even repeat a previous mistake I had to correct. Yet in other ways, it's unbelievably good. I ca…

>So my verdict is that it's great for code analysis, and it's fantastic for injecting some book knowledge on complex topics into your programming, but it can't tackle those complex problems by itself. I don't think you've seen the full potential. I'm currently #1 on 5 different very complex computer engineering problems, and I can't even write a "hello world" in rust or cpp. You no longer need to know how to write co…

How are you qualified to judge its performance on real code if you don't know how to write a hello world?

Yes, LLMs are very good at writing code, they are so good at writing code that they often generate reams of unmaintainable spaghetti.

When you submit to an informatics contest you don't have paying customers who depend on your code working every day. You can just throw away yesterday's code and start afresh.

Claude is very useful but it's not yet anywhere near as good as a human software developer. Like an excitable puppy it needs to be kept on a short leash.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#488

What bothers me about posts like this is: mid-level engineers are not tasked with atomic, greenfield projects. If all an engineer did all day was build apps from scratch, with no expectation that others may come along and extend, build on top of, or depend on, then sure, Opus 4.5 could replace them. The hard thing about engineering is not "building a thing that works", its building it the right way, in an easily unde…

I totally agree. And welcome to disposable software age.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#489

I had an app I wanted for over a decade. I even wrote a prototype 10 years ago. It was fine but wasn't good enough to use, so I didn't use it. This weekend I explained to Claude what I wanted the app to do, and then gave it the crappy code I wrote 10 years ago as a starting point. It made the app exactly as I described it the first time. From there, now that I had a working app that I liked, I iterated a few times to…

This reminds me of how much screensavers on Mac are a PITA. But yes, such a boon for us doodad makers.

And dads who just don't have time to make doodads like we used to!

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#490

Cool. Please check back in with us after they’ve raised the price 50x and you can no longer build anything because you are alienated from your tools.

I’ve said many times, I’d still pay even $1,000 a month for CC.

But I’m a business owner so the calculus is different.

But I don’t think they’ll raise prices uncontrollably because competition exists. Even just between OpenAI and Anthropic.

Post reply on HN