Live data from Hacker News

95% of generative AI pilots at companies are failing – MIT report

fortune.com

11–20 of 174 posts

Re: 95% of generative AI pilots at companies are failing – MIT report

#11
> Despite the rush to integrate powerful new models, about 5% of AI pilot programs achieve rapid revenue acceleration; the vast majority stall, delivering little to no measurable impact on P&L.

This summer, I built two very sophisticated pieces of software. A financial ledger to power accrual accounting operations and a code generation framework that scaffolds a database from a defined data model to the frontend components and everything in between.

I used ChatGPT substantially. I'm not sure how long it would have taken without generative AI, but in reality, I would have just given up out of frustration or exhaustion. From the outside, it would appear to any domain expert that at least three other people worked on these giving the pace at which they got completed.

The completion of those two were seminal moments for me. I can't imagine how anyone, in any field of information systems, is not multiples more effective than they were five years ago. That directly affects a P&L and I can't think of anything in my career that is even remotely close to having that magnitude.

I don't know what encapsulates an AI pilot in these orgs, and I'm sure they are massively more complex than anything I've done. But to hear 95% of these efforts don't have a demonstrable effect is just wild.

Re: 95% of generative AI pilots at companies are failing – MIT report

#12

What's the failure rates if technology pilots in general for comparison? For example, I heard that SAP has an 80-90% deployment failure rate back in the day, but don't have a citable source for it.

Depends on industry I would think. In my previous industry it was something like 25 %, in my current industry it is closer to 80 %.

Re: 95% of generative AI pilots at companies are failing – MIT report

#16
I'm arriving at the conclusion that deployments of LLMs is most suitable in areas where the cost of false positives and, crucially, false negatives are low.

If you cannot tolerate false negatives I don't see how you get around the inaccuracy of LLMs. As long as you can spot false positives and their rate is sufficiently low they are merely an annoyance.

I think this is a good consideration before starting a project leveraging LLMs

Re: 95% of generative AI pilots at companies are failing – MIT report

#17

Why so bad?

Any consumer facing AI project has to contend with the fact that GenAI is predominantly associated with "slop." If you're not actively using an AI tool, most of your experience with GenAI is seeing social media or Youtube flooded with low quality AI content, or having to deal with useless AI customer support. This gives the impression that AI is just cheap garbage, and something that should be actively avoided.

Re: 95% of generative AI pilots at companies are failing – MIT report

#20
> The data also reveals a misalignment in resource allocation. More than half of generative AI budgets are devoted to sales and marketing tools, yet MIT found the biggest ROI in back-office automation—eliminating business process outsourcing, cutting external agency costs, and streamlining operations.

Makes sense. The people in charge of setting AI initiatives and policies are office people and managers who could be easily replaced by AI, but the people in charge not going to let themselves be replaced. Salesmen and engineers are the hardest to replace, yet they aren't in charge so they get replaced the fastest.

Post reply on HN