Live data from Hacker News

95% of generative AI pilots at companies are failing – MIT report

fortune.com

81–90 of 174 posts

Re: 95% of generative AI pilots at companies are failing – MIT report

#81

> The data also reveals a misalignment in resource allocation. More than half of generative AI budgets are devoted to sales and marketing tools, yet MIT found the biggest ROI in back-office automation—eliminating business process outsourcing, cutting external agency costs, and streamlining operations. Makes sense. The people in charge of setting AI initiatives and policies are office people and managers who could be…

I think this is being overly complimenting to AI. I think the most obvious reason is that for almost all business use cases its not very helpful. All these initiatives have the same problem. Staff asking 'how can this actually help me,' because they can't get it to help them other than polishing emails, polishing code, and writing summaries which is not what most people's jobs are. Then you have to proofread all of this because AI makes a lot of mistakes and poor assumptions, on top of hallucinations.

I dont think Joe and Jane worker are purposely not using to protect their jobs, everyone wants ease at work, its just these LLM-based AI's dont offer much outside of some use cases. AI is vastly over-hyped and now we're in the part of the hype cycle where people are more comfortable saying to power, "This thing you love and think will raise your stock price is actually pretty terrible for almost all the things you said it would help with."

AI has its place, but its not some kind of universal mind that will change everything and be applicable in significant and fundamentally changing ways outside of some narrow use cases.

I'm on week 3 of making a video game (something I've never done before) with Claude/Chat and once I got past the 'tutorial level' design, these tools really struggle. I think even where an LLM would naturally be successful (structured logical languages), its still very underwhelming. I think we're just seeing people push back on hype and feeling empowered to say "This weird text autogenerator isn't helping me."

Re: 95% of generative AI pilots at companies are failing – MIT report

#82

Nobody actually wants half the useless tools companies are coming up with because most of the solutions are not really novel. They are just wrapping an LLM. It's kinda like what I realized with the meta Ray-Bans: I can have these things on my face, they can tell me the answer to virtually any question in 10 seconds or less. But I, as a human, rarely have questions to ask. When you walk in to your local grocery store…

I think those kind of glasses may be really useful for blind people. I have seen similar glasses targeted at blind people, that at least in theory, seemed to me like a good idea.

I recall the glasses also can write on the screen inside the lens, which makes me think they may be good for deaf people as well.

It's just that these use-cases seem uncool, and big companies seem to have to be cool in order to keep either their status or their profits. But I have a feeling the technology may be really useful for some really vulnerable people.

Re: 95% of generative AI pilots at companies are failing – MIT report

#83
post #47
post #16

I'm arriving at the conclusion that deployments of LLMs is most suitable in areas where the cost of false positives and, crucially, false negatives are low. If you cannot tolerate false negatives I don't see how you get around the inaccuracy of LLMs. As long as you can spot false positives and their rate is sufficiently low they are merely an annoyance. I think this is a good consideration before starting a project l…

Has inaccuracies been an issue for any of the systems you have developed using LLMs? I hear your complaint quite a bit but it does not align with my experience. Definitely one shotting a chatbot around an esoteric problem introduces possible inaccuracies. If I get an LLM to interrogate a pdf or other document that error rate drops significantly and is mostly on the part of the structuring process and not the LLM. Gen…

Yes I've seen issues with both, but in part what's tricky about false negatives is also that you don't necessarily realise they are there. In the systems I've worked on we've made it simple for operators to verify the work the LLM has done, but this only guards against false positives, which are less problematic.

I've had pretty good success using LLMs for coding and in some ways they are perfect for that. False positives are usually obvious and false negatives don't matter because as long as the LLM finds a solution, it's not a huge deal if there was a better way to do it. Even when the LLM cannot solve the problem at all, it usually produces some useful artifacts for the human to build on.

Re: 95% of generative AI pilots at companies are failing – MIT report

#84

> Despite the rush to integrate powerful new models, about 5% of AI pilot programs achieve rapid revenue acceleration; the vast majority stall, delivering little to no measurable impact on P&L. This summer, I built two very sophisticated pieces of software. A financial ledger to power accrual accounting operations and a code generation framework that scaffolds a database from a defined data model to the frontend comp…

“AI pilots” in the article refers to developing AI-based tools, not to using AI for software development. These projects have a 95% failure rate of successfully deploying the AI tool being developed into production.

Regarding use of AI in software development (which is not what the article is about), the proof of the pudding isn’t in greenfield projects, it’s in longer-term software evolution and legacy code. Few disagree that AI saves time for prototyping or creating a first MVP.

Re: 95% of generative AI pilots at companies are failing – MIT report

#86
post #74

Earlier quoted context omitted.

But do you need AI for those answers? I sometimes do the same thing, but Google/DDG/whatever works fine for most, and a niche app works for others (IDing a bird = Merlin app, for example).

Not the OP, but I ask way more questions now than I used to. Before, I’d sometimes wonder about things, but not enough to actually go and research them. Now, it’s as simple as asking the AI, and more often than not, I get a satisfying answer.

What was the last thing you asked about? What was the answer?

Re: 95% of generative AI pilots at companies are failing – MIT report

#87

Nobody actually wants half the useless tools companies are coming up with because most of the solutions are not really novel. They are just wrapping an LLM. It's kinda like what I realized with the meta Ray-Bans: I can have these things on my face, they can tell me the answer to virtually any question in 10 seconds or less. But I, as a human, rarely have questions to ask. When you walk in to your local grocery store…

> There is like one or two really clever uses I've seen - disappointingly, one of them was Jira. The internal jargon dictionary tool was legitimately impressive. Will it make any more money? Probably not. Sounds like Microsoft 365 Copilot at my org. Sucks at nearly everything, but it actually makes a fantastic search engine for emails, teams convos, sharepoint docs, etc. Much better that Microsoft's own global search…

My favorite copilot use is when I join a MS Teams meeting a few minutes late I can ask copilot: what have I missed? It does a fantastic job of summarizing who said what.

Re: 95% of generative AI pilots at companies are failing – MIT report

#88

Nobody actually wants half the useless tools companies are coming up with because most of the solutions are not really novel. They are just wrapping an LLM. It's kinda like what I realized with the meta Ray-Bans: I can have these things on my face, they can tell me the answer to virtually any question in 10 seconds or less. But I, as a human, rarely have questions to ask. When you walk in to your local grocery store…

> I, as a human, rarely have questions to ask This is an eye-opening sentence. It's quite hard to imagine how to live one's daily life with "few questions to ask." Perhaps this is a neurodivergent thing?

I'm autistic and I probably ask many more questions than most people.

I would also argue that ND people seem to be the heavier AI users, at least in my experience. Its a bit like the stereotypical 'wikipedia deep dive' but 10x.

Re: 95% of generative AI pilots at companies are failing – MIT report

#89
At this rate, how is it better than pure random chance?

The article mentions 19-20 year old founders, focused on solving single user problems, were the successes.

The sample size is 300 public AI deployments and an undisclosed number of private in-house AI projects. And the survey seems to only consider business applications, as compared with end-user applications like media and software. That's significant but not definitive.

Isn't it more likely that existing problems with low hanging fruit, perhaps unpopular answers, that could be solved by leaning on "AI". And perhaps "AI" wasn't the key to success?

Post reply on HN