Live data from Hacker News

Some thoughts on LLMs and software development

martinfowler.com

411–420 of 422 posts

Re: Some thoughts on LLMs and software development

#411
post #8

Earlier quoted context omitted.

> Prompting (overseas) humans [...] After over a decade of trying that, we learned that had... flaws. I think this is a US-centric point of view, and seems (though I hope it's not!) slightly condescending to those of us not in the US. Software engineering is more than what happens to US-based businesses and their leadership commanding hundreds or thousands of overseas humans. Offshoring in software is certainly a US…

The problem isn't that there aren't high quality offshore developers - far from it. Or even high quality AI models. The problems are inherent with outsourcing to a 3rd party and having little oversight. Oversight is, in both cases, way harder than it appears.

The problems are actually not the oversight but the motivation: the point of outsourcing is cost-cutting, not finding equally talented developers elsewhere.

This is especially evident in manufacturing: China can produce extremely sophisticated technology with excellent quality and well thought-out design but that's not what foreign customers look for when outsourcing to China. They want cheap, so they get cheap.

"Oversight" in this case only solves the problem of the quality problems being intentionally hidden initially to get the sale. This might be where the actual parallels with AI models lie: quality of results going down after the initial launch because maintaining the quality is too expensive.

Re: Some thoughts on LLMs and software development

#412

Earlier quoted context omitted.

> _sigh_. Really dude? Just because people overestimate them on average doesn’t mean every person does. In the study, every single person overestimated time saved on nearly every single task they measured. Some people saved time, some didn’t. Some saved more time, some less. But every single person overestimated time saved by a large margin. I’m not saying you aren’t saving time, but it’s very unlikely that if you ar…

I’ll admit it’s possible my estimates are off a bit. What isn’t up for debate though is that it’s made a huge difference in my life and saved me a ton of time. The fact that people overestimate its usefulness is somewhat of a “shrug” for me. So long as it _is_ making big differences, that’s still great whether people overestimate it or not.

If people overestimate time saved by huge margins, we don’t know whether it’s making big differences or not. Or more specifically whether the boost is worth the cost (both monetary and otherwise).

Re: Some thoughts on LLMs and software development

#413
post #396

Earlier quoted context omitted.

This is true, but I've never heard of a use case. To which you might reply, "doesn't mean there isn't one," which you would be also right about. Maybe you know one.

I presume your definition of use case is something that doesn't include what people normally use it for. And I presume me using it for coding every day is disqualified as well.

I didn't mean to suggest it has no utility at all. That's obviously wrong (same for crypto). I meant a use case in line with the projections the companies have claimed (multiple trillions). Help with basic coding (of which efficiency gains are still speculative) is not a multi-trillion dollar business.

Re: Some thoughts on LLMs and software development

#414

Earlier quoted context omitted.

I’ll admit it’s possible my estimates are off a bit. What isn’t up for debate though is that it’s made a huge difference in my life and saved me a ton of time. The fact that people overestimate its usefulness is somewhat of a “shrug” for me. So long as it _is_ making big differences, that’s still great whether people overestimate it or not.

If people overestimate time saved by huge margins, we don’t know whether it’s making big differences or not. Or more specifically whether the boost is worth the cost (both monetary and otherwise).

Only if we’re only using people’s opinions as data. There are other ways to do this.

Re: Some thoughts on LLMs and software development

#415

Earlier quoted context omitted.

If people overestimate time saved by huge margins, we don’t know whether it’s making big differences or not. Or more specifically whether the boost is worth the cost (both monetary and otherwise).

Only if we’re only using people’s opinions as data. There are other ways to do this.

Sure and if we look at data, the. only independent studies we have show either small productivity gains or a reduction in productivity for everything but small greenfield projects.

Re: Some thoughts on LLMs and software development

#416
post #262

Earlier quoted context omitted.

How could a tool be at fault? If an airplane crashes is the plane at fault or the designers, engineers, and/or pilot?

Designers, engineers, and/or pilots aren't tools, so that's a strange rhetorical question. At any rate, it depends on the crash. The NTSB will investigate and release findings that very well may assign fault to the design of the plane and/or pilot or even tools the pilot was using, and will make recommendations about how to avoid a similar crash in the future, which could include discontinuing the use of certain tool…

My point is that the tool (the airplane in this case) is not at at fault, but rather the humans in the loop.

Re: Some thoughts on LLMs and software development

#417

Earlier quoted context omitted.

Only if we’re only using people’s opinions as data. There are other ways to do this.

Sure and if we look at data, the. only independent studies we have show either small productivity gains or a reduction in productivity for everything but small greenfield projects.

Studies plural? Can you link them?

Re: Some thoughts on LLMs and software development

#418

> My former colleague Rebecca Parsons, has been saying for a long time that hallucinations aren’t a bug of LLMs, they are a feature. Indeed they are the feature. All an LLM does is produce hallucinations, it’s just that we find some of them useful. This is an example of my least favorite style of feigned insight: redefining a term into meaninglessness just so you can say something that sounds different while not actu…

The way you have expressed this, I am borrowing it for myself. Many times i run into these kind of situations and I fail to explain why doing something like this is frustrating and actually useless. Thank you.

Re: Some thoughts on LLMs and software development

#419

Earlier quoted context omitted.

Sure and if we look at data, the. only independent studies we have show either small productivity gains or a reduction in productivity for everything but small greenfield projects.

Studies plural? Can you link them?

Google for the Stanford study by Yegor Denisov-Blanch. You might have to pay to access the paper, but you can watch the author’s synopsis on YouTube.

For low complexity greenfield projects (best case) they found a 30% to 40% productivity boost.

For high-complexity brownfield projects (worst case) they found a -5% to 10% productivity boost.

The METR study from a few weeks ago showed an average productivity drop around 20%.

That study also found that the average developer believed AI had made them 20% more productive. The difference in perception and reality was on average 40 percentage points.

Re: Some thoughts on LLMs and software development

#420

Earlier quoted context omitted.

Studies plural? Can you link them?

Google for the Stanford study by Yegor Denisov-Blanch. You might have to pay to access the paper, but you can watch the author’s synopsis on YouTube. For low complexity greenfield projects (best case) they found a 30% to 40% productivity boost. For high-complexity brownfield projects (worst case) they found a -5% to 10% productivity boost. The METR study from a few weeks ago showed an average productivity drop around…

The devil is always in the details with these studies. What did they measure, how did they measure it, are they counting learning the new tool as unproductive time, etc etc etc. I’ll have to read them myself. Regardless, I’ll be sad if it makes most people less productive on average if that’s the scientific truth, but it won’t change the fact that for my specific use case there is a clear time save.
Post reply on HN