Live data from Hacker News

Vibe engineering

simonwillison.net

511–520 of 759 posts

Re: Vibe engineering

#511
post #305
post #258

Earlier quoted context omitted.

Its reasonable to stay away from something one considers dystopian considering the industry is not even sure about the usefulness of coding agents in professional environments. When the tractors replaced the horses, everyone could agree they outperform horses. The result was easily measurable. Its not that simple with LLM agents owned by big corporations.

Sure, it's not yet clear what impact LLMs will have on software development, but the impact it will have will not depend on if developers like to use it or not. If it is going to make software development 10x faster, companies will adopt it, whether devs like it or not.

Sadly true. Most companies don’t even care if the software is sloppy, slow, and ridden with errors that cause data loss or privacy breaches. They care about exploiting workers and extracting value.

Is it ethical? Probably not. It took a few bridges falling and buildings caving in before traditional engineering became a profession.

In this post-Reagan world I’m not sure software has the right context to make that happen. I’m pretty sure we’ll stay the course where the big tech companies like it: very little regulation, loose liability, and terrible software for everyone.

Re: Vibe engineering

#512
post #392

Earlier quoted context omitted.

> How can anyone intellectually honest not see that? The idea that they can only solve problems that they've seen before in their training data is one of these things that seems obviously true, but doesn't hold up once you consistently use them to solve new problems over time. If you won't accept my anecdotal stories about this, consider the fact that both Gemini and OpenAI got gold medal level performance in two ext…

> consider the fact that both Gemini and OpenAI got gold medal level performance Yet ChatGPT 5 imagines API functions that are not there and cannot figure out basic solutions even when pointed to the original source code of libraries on GitHub.

Yes. I don't see why these have to be mutually exclusive.

Re: Vibe engineering

#513

Earlier quoted context omitted.

I really don't get the idea that LLMs somehow create value. They are burning value. We only get useful work out of them because they consume past work. They are wasteful and only useful in a very contrived context. They don't turn electricity and prompts into work, they turn electricity, prompts AND past work into lesser work. How can anyone intellectually honest not see that? Same as burning fossil fuels is great an…

It's not about being honest. It's about Joe Bullshit from the Bullshit-Department having it easier in his/her/theirs Bullshit Job. Because you see, Joe decided two decades ago to be an "office worker", to avoid the horrors of working honestly with your hands or mind in a real job, like electrician, plumber or surgeon. So his day consists of preparing powerpoints, putting together various Excel sheets, attending whate…

Jeez. Brutal but true.

Re: Vibe engineering

#514
post #408

Earlier quoted context omitted.

Of course the devil is in the details. What you say and the skills needed make sense. It's unfortunately also the easiest aspects to dismiss either under pressure as there is often little immediate payoff, or because it's simply the hard part. My experience with llms in general is that sadly, they're mostly good bullshitters. (current google search is the epitome of worthlessness, the AI summary so hard tries to make…

Google's "AI overviews" are one of the worst LLM-powered features on the market today, they're genuinely damaging the reputation of the whole industry. Meanwhile I've started using ChatGPT GPT-5 search as my default search engine! A year ago I would have laugher at the idea: https://simonwillison.net/2025/Sep/6/research-goblin/ And Google themselves have an "AI mode" which is a different league of quality from "AI ov…

It might actually be in Googles best interest to damage the interest in LLMS by showing those crappy AI Mode stuff, because it materially impacts their business model.

The perception of LLMs in the gen pop is what matters, not in the eyes of techies.

Re: Vibe engineering

#515
post #480
post #53

Earlier quoted context omitted.

That's only true if you don't put effort into figuring out how best to use them. If using LLMs makes you slower or reduces the quality of your output, your professional obligation is to notice that and change how you use them. If you can't figure out how to have them increase both the speed and the quality of your work, you should either drop them or try and figure out why they aren't working by talking to people who…

GGP's sentiment resonates with me. I invest a fair bit of time into LLMs to keep up on how †hings are evolving and I do throw both small and large tasks at them. I'm seeing great results with some small task but with anything that is remotely close to actual engineering I just can't get satisfactory results. My largest project is a year old, it's full-stack JavaScript, and I consciously use patterns, structures, and…

Which model and tools are you using it that repo?

Re: Vibe engineering

#516

Earlier quoted context omitted.

And empirical studies on informal code review show that humans have a very small impact on error rates. It disappears when they read more than roughly 200 SLOC in an hour.

Interesting, do you have a link to the study? Our experience is different, at least when reviewing LLM generated code, we find quite a few errors, especially beyond 200 LOC. It also depends on what you're reviewing, 200 LOC != 200 LOC. A boilerplate 200 LOC change? A security sensitive 200 LOC change? A purely algorithmic and complex 200 LOC change?

https://rebels.cs.uwaterloo.ca/papers/emse2016_mcintosh.pdf

Re: Vibe engineering

#517

Earlier quoted context omitted.

> that is the kind of mentality that pushed subpar products on the web for so many years Famously, some of those subpar products are now household names who were able to stake out their place in the market because of their ability to move quickly and iterate. Had they prioritized long-maintainable code quality rather than user journey, it's possible they wouldn't be where they are today. "Move fast and break things"…

Facebook didn't solved any real problem, "Move fast and break things" is for investors not hackers. Famously gmail was very good quality web app and code(can't say the same today) surely not the product of today's "fast iteration" culture

Meta has a 1.79T market cap, they definitely solved some very real problems to get there.

There are lots of companies doing well producing high quality products out of the gate today though, look at Linear.

Both approaches are valid for building sustainable enduring businesses.

Re: Vibe engineering

#518

Earlier quoted context omitted.

Have you used the tools to their full potential?

Another non existing argument, if the agent fails to give the same answer twice i can't even explore his full potential

Hope you've never tried training a dog!

Re: Vibe engineering

#519
post #435

Earlier quoted context omitted.

Very curious to hear responses about this too

The problem with this is that software engineering is a very unorganized and fashion/emotion driven domain. We don't have reliable productivity numbers for basically... anything. I that I'm more productive with statically typed languages but I haven't seen large scale, reliable studies. Same with unit tests, integration tests, etc. And then there are all the types of software engineering: web frontend, web API, mobil…

The other problem is the perennial, how much of what we do actually has value?

Churning out 5x (or whatever - I’m deliberately being a bit hyperbolic) as much code sounds great on the face of it but what does it matter if little to none of it is actually valuable?

You correctly identify that software development is often driven by fashion and emotion but the much much bigger problem is that product and portfolio management is driven by fashion and emotion. How much stuff is built based on the whims of CEOs or other senior stakeholders without any real evidence to back it up?

I suppose the big advantage of being more “productive” is that you can churn through more wrong ideas more quickly and thus perhaps improve your chances of stumbling across something that is valuable.

But, of course, as I’ve just said: if that’s to work it’s absolutely predicated on real (and very substantial) productivity gains.

Perhaps I’m thinking about this wrong though: it’s not about production where standards, and the need to be vigilant, are naturally high, but really the gains should be seen mostly in terms of prototyping and validating multiple/many solutions and ideas.

Re: Vibe engineering

#520
post #506

Earlier quoted context omitted.

> GPT-3 three years ago was over 1,000x the price of much better models today. right, so only another 27 years of moores law continuing left > I'm not yet ready to bet against that trend holding for a while longer. I wouldn't expect an industry evangelist to say otherwise

I'm a pretty bad "industry evangelist" considering I won't shut up about how prompt injection hasn't had any meaningful improvements in the last three years and I doubt that a robust solution is coming any time soon. I expect this industry might prefer an "evangelist" who hasn't written 126 posts about that: https://simonwillison.net/tags/prompt-injection/ (And another 221 posts about ethical concerns with how this s…

yes, enough "concern" to provide plausible deniability
Post reply on HN