How many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for everything that went right and nothing for what could go wrong.
> The post has counts for everything that went right and nothing for what could go wrong. That's AI for you. At Amazon we have many forums to share our AI wins, but none to share AI failures or disappoinments. No wonder execs make bad decisions regarding AI, they only hear completely one-sided stories.
If you were the CEO of Amazon, would you be setting up channels for people to talk about their AI failures? The general way technology is deployed is that we try to find ways to make it work, because those are the most interesting. We're not as interested in all the ways it doesn't work.
From my perspective, some people are trying to use AI in the same way somebody might use a laptop to paddle a canoe. Sure, you can do it, but it's not a good idea. The fact that it doesn't work well is not particularly interesting.
If I steelman your position, I guess the ideal repository would be a set of cases where it's known to work well and a set of cases where it's known to not work well. A little bit like ProtonDB or SteamDB, perhaps.