Live data from Hacker News

2x, not 10x: coding with LLMs in 2026

obryant.dev

221–230 of 260 posts

Re: 2x, not 10x: coding with LLMs in 2026

#221
post #178

Earlier quoted context omitted.

I don't think this is the right framing, because I think it confuses the means and the ends. Its like making a jig in woodworking. The measure of the jig's success is not whether it gets re-used or widespread adoption, its whether it made it easier to achieve some actual objective. Because the jig is a means to some other end. Lots of these "we wouldn't have done it otherwise" applications are means, not ends.

That is true, but that suggests that the value of software is reverting to the value of a jig. Software that gets deployed to millions of people gets a large valuation because it provides small convenience times a million. In the era of bespoke software, you’re basically 100x-ing the production of a millionth of the value of commercial software. That’s why it isn’t all that impactful in the end and I think it’s worth…

It's the underpants gnomes again.

  - Vibe code a bunch of small projects that we couldn't justify ROI before
  - ???
  - Profit

Re: 2x, not 10x: coding with LLMs in 2026

#222

I have a weird issue with using AI for coding. I can code something entirely by myself at my baseline speed; call it 1x. Or I can use Claude to do it, and it does it in 1/10 - 1/4 of the time. The problem, however, is that to review Claude’s code properly takes 2-3x the amount of time it would have taken me to write it all by hand. So my two choices are basically “YOLO, LGTM” and hope I can revert if it breaks someth…

The models are good enough these days that I see it like managing a team of juniors. The LLMs need guidance and oversight, and about the same amount of time I'd spend reviewing code from a junior I spend on the LLM output. I think generally it's best practice to work in a way that "commit and hope it works (because it passes all the CI/CD automated testing and verification, and it's a change gated behind feature flag…

I feel like these days it's more like managing a team of seniors who are technically very competent, but may not always have the full domain context of the problem they're solving. And/or they forget parts, if that context is too large. And/or they have a poor temporal dimension (deprecated badly maintained information polluting the context).

They also tend to overengineer solutions if not guided well.

If you think about it, in that respect it's not that different from managing actual human senior engineers.

I think this is why focussing one's limited human attention more the input (defining clear requirements) as well as on validating the output (good CI/CD including end-to-end automated tests) is far more important than manual code reviews and micro-managing the development process.

Re: 2x, not 10x: coding with LLMs in 2026

#223
post #89

Earlier quoted context omitted.

In 10 years llms will handle it all from start to finish.

We have passed peak LLM already. In 10 years, something might be handling "it all" from start to finish, but it won't be LLMs.

Saying we're at the Peak of LLMS is like saying we never went further than everest, these are probabilistic state machines sure, but the real value is the effort put into tooling and training these models in RL situations.

Re: 2x, not 10x: coding with LLMs in 2026

#224

I have a weird issue with using AI for coding. I can code something entirely by myself at my baseline speed; call it 1x. Or I can use Claude to do it, and it does it in 1/10 - 1/4 of the time. The problem, however, is that to review Claude’s code properly takes 2-3x the amount of time it would have taken me to write it all by hand. So my two choices are basically “YOLO, LGTM” and hope I can revert if it breaks someth…

The models are good enough these days that I see it like managing a team of juniors. The LLMs need guidance and oversight, and about the same amount of time I'd spend reviewing code from a junior I spend on the LLM output. I think generally it's best practice to work in a way that "commit and hope it works (because it passes all the CI/CD automated testing and verification, and it's a change gated behind feature flag…

Reviewing LLM code is quite different from junior code.

With juniors you tread carefully and give feedback only on important points to encourage growth.

With LLMs you channel the inner sailor and nitpick so much that even a senior would start to cry.

Re: 2x, not 10x: coding with LLMs in 2026

#226
This has been my experience on a recent small project for a client. I took effort to understand and prepare the spec before feeding it to Claude Opus. Yet, the result was slop full of defects, and one that I did not understand well at that. I just did not have a mental model of it. I just thought, there's no way I could ever feel confident about this code, something is missing. It did not take me long to write the important parts myself, from scratch and develop that mental model.

Then I would use Claude to make a small fix, write a unit test (write a unit test, not suggest what they unit test should be) or write some less important UI code. It just felt like old school development, with some assistance. I think we may see a shift back to a "copilot" mode and by the way, I still use GitHub Copilot. The $19 subscription includes $30 worth of credits and the autocomplete in VS code that uses some cheap model is completely free. This autocomplete is very good and I would like the market to focus more on IDE integrations. The agentic thing is either ahead of its time or it will never have time. Time will tell. I think the LLM can only be as good as the training data, so it will always excel at small chunks, but to implement entire projects of which there's such variety, I would question that. And also, how using it to implement entire projects means the developer does not hold the program model in the mind, which leads to more issues long term.

By the way, in this way of working, I see no difference between Sonnet and Opus. Sonnet is good enough to use as an aid.

Re: 2x, not 10x: coding with LLMs in 2026

#227

Earlier quoted context omitted.

> gave me a checklist that helped me quell my travel anxiety. it's truly fascinating how many positive descriptions of AI gesture at emotional management. I think that's the killer feature of this technology -- it makes people feel good, capable, reassured -- without the risk and vulnerability of interacting with another human. I use Google Weather for my forecasts btw, no need to vibecode an app for that

Gathering forecasts for each leg of a trip only on the particular days for the particular location is tedious. Anyone can do it with google weather but why spend the time when you don't have to. I can keep track of my expenses on a napkin but i'd much rather use a spreadsheet or dedicated app especially when that app is effectively free.

I asked Google Gemini to generate me a basic expense estimator to paste in ... Google Sheets. Every single formula needed editing. At least it generated the basic layout for me.

Re: 2x, not 10x: coding with LLMs in 2026

#228

Even if that 2x were a representative number (although another HN story today says 1.1x - 10% improvement), what really matters is whether companies are overall seeing any AI spending return on the bottom line. There is a lot of reason to suspect that at most companies, especially more established non-startup ones, there will close to zero bottom line benefit, because these companies already have free paid-for develo…

I agree the bottom line ROI largely isn't there and won't be in the future either. The upside for companies is not the bottom line but the potential savings from hiring fewer developers or cutting jobs. As AI based development processes continue to improve, the confidence level employers will have in safely cutting jobs will increase. The longer term impact of AI is HR savings rather than introducing new capabilities…

> The longer term impact of AI is HR savings rather than introducing new capabilities into business.

I tend to agree, although this goes against the Dwarkesh narrative (apparently matching current SV zeitgeist) that there is some insatiable demand for "warehouses full of genius coders". If there really was demand for more coders (especially at the high prices the AI companies are hoping for), then companies would be hiring the unemployed developers available right now, not laying more off.

I wonder how many CEOs or CTOs appreciate the massive functional gap between an AI coder like Fable and a human developer - full general intelligence, with continual learning, theory of mind (so they understand what the boss wants, not just what s/he says), etc... all available now, not some 5-10 year AGI/ASI stretch goal (that may in fact take much longer - artificial brain, not just "AGI") ...

Re: 2x, not 10x: coding with LLMs in 2026

#229
post #92

Earlier quoted context omitted.

Not every engineering effort is a product in search of market fit or a community. Some things are already very useful as just a one shot. I just made a quick app to help me pack for a trip, it updated forecasts every day, let me know when rain entered the forecast at one of my stops and gave me a checklist that helped me quell my travel anxiety. The greatest thing that LLMs have done is allow many to achieve things t…

> Not every engineering effort is a product in search of market fit or a community I think you misunderstood my point: by "legitimate stand-out products", I meant exactly that, with no connotation of commercialization. Maybe you can agree that having a high signal-to-noise ratio for (open source) projects is a desirable goal? > Some things are already very useful as just a one shot. I agree. I too have made or forked…

> having a high signal-to-noise ratio for (open source) projects is a desirable goal?

It’s not obviously true. A higher number of attempts, a larger talent pool, typically doesn’t change the average much (or it might even make the average go down), but tends to produce higher peak outcomes.

We see this everywhere (science, startups, sports, chess, etc).

If you want the best spreadsheet, game, or whatever app you want, you’re only interested in the few highest peaks.

So you do actually get better signal to noise with a larger wasteland of discarded attempts. The higher peaks make it easier to filter out the noise.

The goal you’re intrinsically motivated by seems different than this. That seems to be the whole disagreement.

Post reply on HN