Live data from Hacker News

How to recognize AI snake oil [pdf]

cs.princeton.edu

321–330 of 364 posts

Re: How to recognize AI snake oil [pdf]

#321
Easiest way to cut through AI BS: ask “what’s their dataset?” If it’s not obvious, there’s a problem. AI is only as powerful as the data it learns from.

There is one exception to this: if an exhaustive simulation of the problem exists. This is why AI is so successful at sandboxed games like chess and Go. It can generate its own data with zero ambiguity.

So: what’s your dataset? What simulation are you inverting? If neither, you’re just writing an expert system based on heuristics.

Re: How to recognize AI snake oil [pdf]

#322

Earlier quoted context omitted.

We won’t ever have an AI winter like in the 70s again. A lot of ML is already very useful across many domains (computer vision, NLP, advertising, etc). Back then, there was almost no personal computing, almost no internet, smol data, and so on. Stuff you need for ML to be useful and used. So what if some corporate hack calls linear regression “AI”? The results speak for themselves. The ML genie is too profitable to g…

Didn't linear regression used to be called "AI" as recently as a decade ago?

I personally don’t see a problem with this. Where do you draw the line at model simplicity? Are decision trees too simple to be AI? What about random forest? Are deep neural nets the only model sophisticated enough to be “AI”? It’s not the model, it’s how you use it.

Re: How to recognize AI snake oil [pdf]

#323
This is good. I co-founded Futurescaper[1], a company that does work in strategic foresight systems -- collective mapping of complex systems, and analytical tools to help people and organisations to understand them. This put us in a prime position to be vendors of AI snake oil. Due to a stubborn overabundance of ethics, we've refused to do so, which has undoubtedly cost of a lot of business. People want to buy snake oil. It goes down much smoother than hard truths.

To amplify what this presentation says, here's the hard truths about predicting social outcomes: either they

1. Are simple and obvious, and can be easily understood and predicted via regressions and trendlines.

2. Are complex, non-obvious, and neither can nor should be predicted. In fact, attempting to predict them is often dangerously wrong. But this doesn't mean that they can't be understood.

In the first instance, you don't need any fancy software, so there's no snake oil to be sold. The second instance, however, gets caught in a cognitive bias: we think that the future can be predicted. This is because, in simple systems, it can: drop a glass above a hard floor, and you can accurately predict that it will fall and shatter. That is a simple system. We think that complex systems must be similar... just more complex.

But complex systems are fundamentally different. Consider a double-pendulum. Its movement can't be predicted for more than the next few swings. Even if you know the exact starting configuration of the pendulum -- and by exact, I mean not just every single sub-atomic particle in the pendulum arms, but the gravitational influence of literally every object in the universe -- you still wouldn't be able to greatly extend your foreknowledge of its movements. This is because the feedback loops are driven by chaos, and chaos is baked directly into the mathematical fabric of the universe itself.

For the mathematically inclined, consider the Mandelbrot set: it is, essentially, the equivalent of a double pendulum. It asks a question: "by starting with this number and iteratively exponentiating it, will in trend towards zero or infinity?". When you ask this question on a simple number line, then the answer is obvious: below 1, it trends towards zero; above 1, it trends towards infinity. However when you ask this question on the complex number plane -- with two numbers feeding back to each other as they exponentiate -- then the only way to answer the question is to keep iterating the numbers to find out. There's no shortcut. In some places, the answer is found quite quickly. In others, it takes thousands of iterations. In still others, it takes infinite iterations: you could build a computer the size of the galaxy and you still wouldn't be able to answer the question of whether such-and-such coordinates trends towards zero or infinity. That's the complexity you get from just two number lines and a very simple feedback loop.

The real world of psychological and social cause-and-effect is far more complex than a double pendulum or a Mandelbrot set, and attempting to predict it is even more futile. In fact, it's dangerous.

What makes this dangerous is that you can throw statistics at complex futures, and make statements like "Future A has a 40% probability; futures B, C, and D all have a 20% probability". But the human mind is terrible at making good judgements based on this kind of information. Hearing it, people tend to think: "right, that's settled then: Future A is twice as likely as any other scenario, so that's what we'll plan for." We fixate on what looks like the most likely future, and disregard the rest.

The problem is that if the preparations for Future A are contrary to what you'd want to do in Futures B, C, and D, then betting everything on Future A means that 60 % of the time, you'll lose.

What's interesting is that increasing the accuracy of those percentages doesn't necessarily help, and in many cases can hurt, since it only reinforces our tendency towards target fixation. In a 40/20/20/20 scenario, we might still make some concessions towards planning for the "20s". But in a 70/10/10/10 scenario, those "10s" will simply be discarded. Which means that 30% of the time you won't just lose: you'll be blindsided. Utterly fucked.

Unfortunately, most of the predictive AI that I've seen is focused on either increasing the "certainty" via extremely dubious means, or simply hiding the non-dominant answers altogether. So the AI does the target fixation for you. That's not a good thing.

The entire discipline of Scenario Planning[2] evolved to help people and companies "un-predict" the future. Rather than putting percentages on probabilities of outcomes, people need to understand the possibilities of outcomes. Even if complex systems can't be predicted, they can be better understood, and that can be very valuable for helping to navigate them in real-time. Rather than giving people the easy (and usually wrong) answer of "here's what is going to happen", you can give them a range of futures that could happen, and understanding those possibilities can aid navigation and lead to better outcomes.

It may even be that AI has a legitimate role to play in this process -- but it won't be in falsely predicting the outcome of complex systems, nor will it be in absolving people of the responsibility of thinking for themselves. This, unfortunately for my company's ability to raise capital, is a much less sexy sales pitch than the AI snake oil. (Although it does get us smarter and less annoying clients, so there's that!)

1: https://www.futurescaper.com/

1: https://en.wikipedia.org/wiki/Scenario_planning

Re: How to recognize AI snake oil [pdf]

#324
post #240
post #129

Earlier quoted context omitted.

This reminds me of a job interview I was on. I was asked about how I would use AI/machine learning for their problem space. Since they seemed to be smart and level-headed, I answered honestly, "Pick something unimportant, use a machine learning algorithm just to get familiar with the tools, ignore the result unless it happens to work, then put machine learning in your marketing materials. But keep track of it, and if…

Speaking of jobs and interviews. I am yet to find a job board which does not show JavaScript jobs when searching for Java jobs. Some of them claim to use AI. :)

Let's not forget the job application systems that make you chronologically list every previous employer, their location, your title, and dates of employment, your education history including dates and degrees received, skillsets/technologies you have experience with, etc. All information that is on your CV and/or LinkedIn profile and they want you to manually re-type it into their late 1990's era job application system rather than using some basic NLP to extract it. Personally, the moment I start applying and the system asks me for more than my email/mobile and a CV upload, I bail on the application.

Re: How to recognize AI snake oil [pdf]

#325

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

Doesn’t it depend on the size and variety of your dataset?

I can see “random” performing well in a set of <1000 videos, all on similar subject spaces (eg “memes”, or “python”), but recommending relevant stuff gets much harder as the amount of content grows...

Re: How to recognize AI snake oil [pdf]

#326

What I dislike far more than the idea of using such systems to predict social outcome is that the usage of such systems is done behind closed doors. I would be much more willing to accept such systems if the law required any system to be fully accessible online, including the current neural network, how it was trained, and training data used to train it (if the training data cannot be shared online, then the neural n…

Using an association between features to make a prediction about something, rather than measuring the thing itself, is exactly what’s meant by “prejudice.” Even when the associations are real and the model is built with perfect mathematical rigor. ML is categorically unsuitable for government decisions affecting lives.

You seem to think this level of prejudice for prediction is wrong. Why?

If someone has killed 12 people, being prejudice about their chance of killing another and using that to determine the length of a sentence seems reasonable.

Even with something like a health inspection. Measuring how they store and cook raw chicken is about predicting the health risks to the public eating it, not about measuring the actual number of outbreaks of salmonella. And even if they were to measure the previous outbreaks of salmonella and use it to prediction the future outbreaks, that is still two different things.

Re: How to recognize AI snake oil [pdf]

#327
post #69

What I dislike far more than the idea of using such systems to predict social outcome is that the usage of such systems is done behind closed doors. I would be much more willing to accept such systems if the law required any system to be fully accessible online, including the current neural network, how it was trained, and training data used to train it (if the training data cannot be shared online, then the neural n…

The problem with leaving this to independent companies is that some of the most natural application areas are dominated by independent companies operating as a cartel -- think about credit scoring. The fact that the data and models are a natural "moat" (in the sense of Warren Buffett) is all the more worrying.

The difference is the level of threat between a private company denying you a loan and the government deciding you should spend 10 years in prison. Please don't misunderstand, I'm not saying the former is good, only that it isn't nearly as bad as the latter.

Re: How to recognize AI snake oil [pdf]

#328

What I dislike far more than the idea of using such systems to predict social outcome is that the usage of such systems is done behind closed doors. I would be much more willing to accept such systems if the law required any system to be fully accessible online, including the current neural network, how it was trained, and training data used to train it (if the training data cannot be shared online, then the neural n…

A judge using his experience and judgement to subjectively set a jail sentence is as opaque as a proprietary algorithm. He or she may cite reasons for the sentence, but nobody is verifying that judges' sentences are consistent with the criteria they cite.

You are right and I see that as a flaw in the current system that should be fixed.

Re: How to recognize AI snake oil [pdf]

#329
post #129

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

This reminds me of a job interview I was on. I was asked about how I would use AI/machine learning for their problem space. Since they seemed to be smart and level-headed, I answered honestly, "Pick something unimportant, use a machine learning algorithm just to get familiar with the tools, ignore the result unless it happens to work, then put machine learning in your marketing materials. But keep track of it, and if…

The exact same applies to blockchains.

Blockchain is mostly a marketing tool, not something you would want to use in production for anything.

Re: How to recognize AI snake oil [pdf]

#330

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

IMH (and biased) O, a lot of great coders are implimentors, or let's say applied computers scientists. That is great for building incredible open source software and a lot of other things that I would not be able to do given a 1000 years. However (again IMHBO) a specific ML and any other specific application of stastistics or mathematics becomes really tricky once your use case is explicitly defined. You then need in…

I could not agree more about self driving cars. They will disrupt and cause us to actually look at our terrible transportation infastructure, not learn to survive in it.
Post reply on HN