Live data from Hacker News

Opus 4.5 is not the normal AI agent experience that I have had thus far

burkeholland.github.io

951–960 of 1001 posts

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#951

Earlier quoted context omitted.

I'll second this. I'm making a fairly basic iOS/Swift app with an accompanying React-based site. I was able to vibe-code the React site (it isn't pretty, but it works and the code is fairly decent). But I've struggled to get the Swift code to be reliable. Which makes sense. I'm sure there's lots of training data for React/HTML/CSS/etc. but much less with Swift, especially the newer versions.

I hate "vibe code" as a verb. May I suggest "prompt" instead? "I was able to prompt the React site…."

You aren't prompting the React site, you're prompting the LLM.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#952

Earlier quoted context omitted.

One problem with the idea of making businesses out of this kind of application is actually mentioned in passing in the article "I decided to make up for my dereliction of duties by building her another app for her sign business that would make her life just a bit more delightful - and eliminate two other apps she is currently paying for" OP used Opus to re-write existing applications that his wife was paying for. So…

I think you're misunderstanding my point. If you can crank out a custom app this quickly, you don't make a commercial app and then try to sell it on an app store. Customers pay you to make apps for their specific usecase. One app, one customer. And if a week later they want some new features, they pay you (or another freelancer) to add it. Put another way, we programmers have the luxury of being able to write custom…

Why do they pay you though, why not just do it themselves? With improving models and surrounding tooling the barrier to creating apps is lowered, and it's easier for a user just to create their own app, no 3rd party person needed.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#953

Earlier quoted context omitted.

The venn diagram of engineering and prompting is two circles, maybe a tiny overlap with integrated environments like claude code. A program, by definition, is analyzable and repeatable, whereas prompting is anything but that.

As long as your program is large and multi-threaded (most programs that matter commercially), it is not very analyzable or repeatable. You replace those qualities with QA and tests, the same is true with prompting.

Sounds like it's time for your LLM daddy to have the Coq talk with you..

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#954

Earlier quoted context omitted.

This was me. I was a huge AI coding detractor on here for a while (you can check my comment history). But, in order to stay informed and not just be that grouchy curmudgeon all the time, I kept up with the models and regularly tried them out. Opus 4.5 is so much better than anything I've tried before, I'm ready to change my mind about AI assistance. I even gave -True Vibe Coding- a whirl. Yesterday, from a blank dire…

> Oh, and I did this all without ever opening a single source file or even looking at the proposed code changes while Opus was doing its thing. I don't even know Kotlin and still don't know it. ... says it all.

What exactly does it say, in your opinion? I can imagine 4-5 different takes on that post.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#955

Earlier quoted context omitted.

- I cloned a project from GitHub and made some minor modifications. - I used AI-assisted programming to create a project. Even if the content is identical, or if the AI is smart enough to replicate the project by itself, the latter can be included on a CV.

Do people really see a CV and read "computer mommy made me a program" and think it's impressive

Unfortunately, it is happening. I remember an old post on HNs, it mentioned that a "prompt engineer for article generating" can find more jobs than a columnist writer. And op just wrote articles by himself but declared that all artices were generated by AI.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#956

I've been on a small adventure of posting more actively on HN since the release of Gemini 3, trying to stir debate around the more “societal” aspects of what's going on with AI. Regardless of how much you value Cloud Code technically, there is no denying that it has/will have huge impact. If technology knowledge and development are commoditised and distributed via subscription, huge societal changes are going to happ…

The tokens cost the same in Bangalore as they do in San Francisco. The robots will be able to make stuff in San Francisco just as well as they do in Bangalore. The only thing that will matters is natural resource availability and who has more fierce NIMBYs.

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#957

I’ve been saying this a countless time, LLM are great to build toy and experimental projects. I’m not shaming but I personally need to know if my sentiment is correct or not or I just don’t know how to use LLMs Can vibe coder gurus create operating system from scratch that competes with Linux and make it generate code that basically isn’t Linux since LLM are trained on said the source code … Also all this on $20 plan…

> Can vibe coder gurus create operating system from scratch that competes with Linux and make it generate code that basically isn’t Linux since LLM are trained on said the source code …

No.

Vibe-coding, in the original sense where you don't bother with code reviews, the code quality and speed are both insufficient for that.

I experimented with them just before Christmas. I do not think my experiments were fully representative of the entire range of tasks needed for replacing Linux: Having them create some web apps, python scripts, a video game, a toy programming language, all beat my expectations given the METR study. While one flaw with the METR study is the small number of data points at current 50% successful task length, I take my success as evidence I've been throwing easy tasks at the LLM, not that the LLM is as good as it looks like to me.

However, assume for the moment that they were representative tasks:

For quality, what I saw just before Christmas was the equivalent of someone with a few years' experience under their belt, the kind of person who is just about to stop being a junior and get a pay rise. For speed, $20 of Claude Code will get you around 10 sprints' equivalent to that level of human's output.

"Junior about to get a pay rise" isn't high enough quality to let loose unchecked on a project that could compete with Linux, and even if it was, 10 sprints/month is too slow. Even if you increase the spend on LLMs to match the cost of a typical US junior developer, you're getting an army of 1500 full-time (40h/week) juniors, and Linux is, what, some 50-100 million developer-hours, so it would still take something like 16-32 years of calendar time (or, equivalently, order-of 1.2-2.5 million dollars) even if you could perfectly manage all those agents.

If you just vibe code, you get some millions of dollars worth of junior grade technical debt. There's cases where this is fine, an entire operating system isn't one of them.

> Also all this on $20 plan. Free and self host solution will be best

IMO unlikely, but not impossible.

A box with 10x the resources of your personal computer may be c. 10x the price, give or take.

While electricity is negligible (which today, hah!): If any given person is using that server only during a normal 40 hour work week, that's 25% utilisation rate, therefore if it can be rented out to people in other timezones or where the weekend is different, the effective cost for that 10x server is only 2.5x.

When electricity price is a major part of the cost, and electricity prices vary a lot from one place to another, then it can be worth remote-hosting even when you're the only user.

That said, energy efficiency of compute is still improving, albeit not as rapidly as Moore's Law used to, and if this trend continues then it's plausible that we get performance equivalent to current SOTA hosted models running on high-end smartphones by 2032. Assuming WW3 doesn't break out between the US and China after the latter tries to take Taiwan and all their chip factories, or whatever

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#959
the author asks one interesting question and then glides right by it. If the agents only need their own code, what should that code look like? If all their learning has come from old human code, how will that change in the future as the ecosystem fills up with agent code?

Re: Opus 4.5 is not the normal AI agent experience that I have had thus far

#960
post #77
post #8

Opus 4.5 has become really capable. Not in terms of knowledge. That was already phenomenal. But in its ability to act independently: to make decisions, collaborate with me to solve problems, ask follow-up questions, write plans and actually execute them. You have to experience it yourself on your own real problems and over the course of days or weeks. Every coding problem I was able to define clearly enough within th…

I find my sweet spot is using the Claude web app as a rubber duck as well as feeding it snippets of code and letting it help me refine the specific thing I'm doing. When I use Claude Code I find that it *can* add a tremendous amount of ability due to its ability to see my entire codebase at once, but the issue is that if I'm doing something where seeing my entire codebase would help that it blasts through my quota to…

I've had similar experiences but I've been able to start using Claude Code for larger projects by doing some refactoring with the goal of making the codebase understandable by just looking at the interfaces. This along with instructions to prefer looking at the interface for a module unless working directly on the implementation of the module seems to allow further progress to be made within session limits.
Post reply on HN