Live data from Hacker News

How to recognize AI snake oil [pdf]

cs.princeton.edu

241–250 of 364 posts

Re: How to recognize AI snake oil [pdf]

#241
Once upon a time, I used to do research in the field of Wireless Sensor Networks, which used to be abbreviated into WSN. You can still find lots of research papers using the abbreviation in the title, circa 10-12 years ago.

It went through a hype cycle, and then people sort of moved on - into IoT. And IoT, naturally, just sounds better and is more accessible, plus the tech did catch up, so it became a lot more popular than WSN (the term).

I told a friend of mine that no one paid attention when the field was called WSN, and now it is the exact same thing, but IoT is taking off like crazy. (At that point, I had left research altogether). He said "Yes, and that's perfectly reasonable. If you don't invent new terminology every few years and manufacture some kind of new hype, the funding agencies stop sending you money."

This is a community which is based on the mantra of "making something users want". Once you realize that the upstream user here is the funding agency, and what they really want is to make bets on "cool stuff for tomorrow" rather than boring old 3-5 year old tech (such as WSN), the hype actually makes perfect sense.

Sure, there is still a need to separate the snake oil from the reality. But that is true of tech in general. I am not sure if AI/ML is particularly bad in some way.

Re: How to recognize AI snake oil [pdf]

#242
post #14

Earlier quoted context omitted.

You should emphasize that this is Organic AI. It's low carbon and overall greener.

Or keep calling it AI, and concede that AI stands for "actual intelligence" if someone asks you directly.

Call it Al, and only hire Alans, Alberts, Alphonses, Alis, etc

Re: How to recognize AI snake oil [pdf]

#243
I tend to find the following few questions a quick way to evaluate an ML/AI pitch in an elevator:

Ask yourself: could a human given the inputs reasonably produce the outputs you are looking for? — this helps avoid/identify the pure magic pitches.

Then ask the person pitching: 1. What’s your training data and how is it collected? 2. What’s your validation data and how is that constructed, and how does your system perform on that set? 3. What are the blind spots and biases in your model and how are you mitigating them?

If they don’t have succinct and competent answers or major red flags like no validation or claiming no biases then run away.

Re: How to recognize AI snake oil [pdf]

#244

Earlier quoted context omitted.

Or keep calling it AI, and concede that AI stands for "actual intelligence" if someone asks you directly.

AI now is like Cyber was in the 1990s it's seems to be nothing but a buzzword for many organizations to throw around. The term AI is used as if humanity now has figured out general AI or artificial general intelligence (AGI). It's quite obvious organizations and people use the term AI to fool the less tech inclined into thinking it's AGI - a real thinking machine.

Has been like that for quite a while, it has up and lows with the 90s marking a winter season for AI, and now the hype machine is on full steam again, until people find out again a lot of it is marketing BS to get funding. Then the research that is worthwhile gets mixed up with that, go into a lack of funding and in 30 years or so it's back on full hype again.

Re: How to recognize AI snake oil [pdf]

#245

Earlier quoted context omitted.

I think the point is labeling itself is very difficult except for special and limited domains. Manually constructed labels, like feature engineering, are not robust and do not advance the field in general.

That makes sense. I'm coming from the angle of applied ML where solutions need to solve a business problem rather than advance the field of ML. In consulting many problems can't be solved well without a labeled dataset and in lieu of one, less credible data scientists will claim they can solve it in an unsupervised manner.

For sure. There are counter-examples however - fully unsupervised machine translation for resource poor languages comes to mind and is increasingly getting business applications.

I think that in the future, more and more clever unsupervised approaches will be the path forward in huge AI advances. We've essentially run out of labeled data for a large variety of tasks.

Re: How to recognize AI snake oil [pdf]

#246

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

Since we are sharing anecdotes, I can report it's been 20 years of buying stuff on the internet and the combined billions of ad tracking research dollars spent by Amazon and Google have not yet come up with a better algorithm than to bombard me with ads for the exact same thing I just bought .

I have been looking for a shelf that’s as close to 50” wide and 10” deep as possible. No search on any site that allows you to search for such a thing as far as I can tell.

Re: How to recognize AI snake oil [pdf]

#247
post #215

Earlier quoted context omitted.

I'm aware of moderation. What else?

Recaptcha is probably one you've actually interacted with, but even then you're mostly reinforcing existing predictions. But other applications within Google Maps are things like street number recognition, logo recognition, etc. Waymo contract out object detection from vehicle cameras and LIDAR point clouds. Google even sell a data labeling as a service. I believe Google Maps has a lot of humans who tidy up the autom…

Yes you can report problems with the road network and people update GMaps manually. And up until a couple of years ago users could do it themselves, but they took it down for some reason.

Changes to other types of places can still be done manually by GMaps users themselves, and other users can evaluate that, and I guess if it's a "controversial" (low rep user did the change or people voted against it) a Google employee evaluates it. And if you're beyond certain level as a GMaps user you can get most changes published immediately.

Re: How to recognize AI snake oil [pdf]

#248
post #20

Top textual feature predicting snake oil: calling the product AI rather than ML.

Sometimes you just need to do it that way, because the people buying do not know what machine learning is - even though they've heard about AI. For example, I was networking for some jobs in data science - and was approached by some energy company. Struck up a conversation with the guy (older exec), and he said "so I hear you have a background from AI, correct?" to which I replied "I have a degree in Machine Learning…

Seems like there was a time when "machine" was the term grant committees wanted to hear, they were perhaps fed up with the theory and wanted more tangible stuff. So the theorists branded their stuff as machines. See support vector machine, kernel machines...

Similar to how "dynamic programming" was coined to please funding agencies.

Re: How to recognize AI snake oil [pdf]

#249
post #238

Earlier quoted context omitted.

This can happen when a retargeting campaign doesn't have a 'burn pixel' or conversion event trigger. It's a common oversight, which can cause a re-targeting program to kick-off unnecessarily (or cause ads that have followed you around to become obvious)

And yet I haven't seen one ad that has a 'not interested' feature similar to YouTube. I mean, if I could stop seeing washing machines or whatever, I'd probably click it.

do you clear your cookies and html5 storage? That should wipe any personalization that's happening, but the ads will become very generic then.

You can also block ads (Ublock/Umatrix)

Re: How to recognize AI snake oil [pdf]

#250

The author says "AI is already at or beyond human accuracy in all the tasks on this slide and is continuing to get better rapidly" and one of his examples is "Medical diagnosis from scans". That is an example of precisely the sort of snake oil hype he's berating in the social prediction category. In an extremely narrow sense of pattern recognition of some "image features", i.e. 5% of what a radiologist actually does,…

Author here. I appreciate your criticism. What I had in mind was more along the lines of Google's claims around diabetic retinopathy. I received feedback very similar to yours, i.e. that those claims are based on an extremely narrow problem formulation: https://twitter.com/MaxALittle/status/1196957870853627904 I will correct this in future versions of the talk and paper.

Thanks for the .pdf and the research in general, great stuff!

One thing I'd love is a look at 'noise' in these systems, specifically injecting noise into them. Addons like Noiszy [0] and trackmenot [1] claim to help, but I'd imagine that doing so with your GPS location is a bit tougher. I'd love to know more on such tactics, as it seems that opt-ing out of tracking isn't super feasible anymore (despite the effectiveness of the tracking).

Again, great work, please keep it up!

[0] https://noiszy.com/

[1] https://trackmenot.io/

Post reply on HN