Live data from Hacker News

How to recognize AI snake oil [pdf]

cs.princeton.edu

211–220 of 364 posts

Re: How to recognize AI snake oil [pdf]

#211

Earlier quoted context omitted.

Well to be fair that's how all employers also hire. If you did a good job at the last company you'll probably do a good job here. If you did a good job yesterday, you'll probably do a good job today. For the most part they are usually correct.

I hope you are joking since the industry collectively knows how much employers/interviewers value algorithms-based coding interview, which doesn't correlate strongly with performance. Even if you are talking about senior positions where they don't matter, then you should know that people hire someone they know+like who did decently well, rather than the truly best on the market.

The coding interview comes usually after they vet your resume, any profiles, get you talk to an HR, and potentially check your references.

Even the coding interviews are just a signal against overall performance given a short amount of time. It's only a sample of data, but if you did your interview right you should be able to protect some against bad people getting very lucky. Just like driving a car; bad drivers tend to stay pretty bad, and good drivers tend to stay safe. Even though there's a lot of ways to define what is a good driver, there's clearer ways to define what is a bad driver, and if someone was a bad driver yesterday they are still probably a bad driver.

Re: How to recognize AI snake oil [pdf]

#212
This seems contradictory:

> AI can't predict social outcomes

> In most cases, manual scoring rules are just as accurate

So manual scoring rules don't work either for predicting social outcomes? There is some magic sauce that humans use for prediction that we haven't cracked yet? Nothing can predict social outcome?

AI is perfectly capable of predicting social outcomes, and only in very few cases are manual scoring rules as accurate as black box AI. The ethical concern is not about accuracy, but about our sensibilities when it comes to protected classes. The author cherry picked examples where simpler approaches also worked, but says nothing of practical feasability or increase in variance. Try actually doing face recognition or spam detection with manual rules.

Face recognition being way more accurate is just as much an ethical concern as a gun that is way more accurate. It all depends on who you point it at. Accurate face recognition at the border helps save lives as much as equipping the police with more accurate hand guns.

The talk of AGI is misguided. Everybody can see that the economy will be increasingly automated with narrow AI. Just because "big data" was a hype word, does not mean companies haven't been monitizing their big data (and were thus right to collect it).

We can predict probabilities about the future. The author is attacking these systems for not being 100% sure. Predictive policing is automated resource management. Militaries have been doing this for decades. It has its drawbacks, but also benefits (wiser usage of tax money, protecting low-income neighborhoods from falling in the hands of gangs).

The author also claims that algorithms automatically turn away people at the border for posting or liking or being connected to terrorist propaganda. But these systems just give a score and a human border guard makes the (more informed) decision.

A system not being 100% accurate is not an ethical concern, as long as we not treat those systems as 100% accurate and give proper recourse.

Just a spelling check can and does weed out poor candidates. Why does HR want to automate? Because they get 1000+ resumes for a single position. The manual glance they give them pale in comparison to what an automated system can do.

What is more likely? That these HR systems show promise? Or that the VC market has completely lost it (despite working with software and automation for decades, and have AI experts on staff) and is pumping billions into tealeaf reading, because now its called "AI"?

If you cheat the system by adding "Cambridge" or "Oxford" in white letters to your CV, is that ethical? Why not add it to your education section in black letters? Would you hire a good potential candidate, if you knew they acted like 90s search engine spammers? Maybe a candidate from Oxford or Cambridge really deserves to be on the top of the pile, or is it now unethical to look at education when hiring?

This presentation likes to mix ethics with technical success. Just say that a HR system is unethical, without calling it bogus with zero proof other than "some AI experts agree that this is impossible".

Yes, there is a lot of snake oil AI, and this will only increase. But these systems can and do work. I am sure there are AI experts building these systems right now.

Re: How to recognize AI snake oil [pdf]

#213
post #209
post #129

Earlier quoted context omitted.

This reminds me of a job interview I was on. I was asked about how I would use AI/machine learning for their problem space. Since they seemed to be smart and level-headed, I answered honestly, "Pick something unimportant, use a machine learning algorithm just to get familiar with the tools, ignore the result unless it happens to work, then put machine learning in your marketing materials. But keep track of it, and if…

That sounds odd, like they don't really need machine learning (unless it is to snare investors?).

Snaring customers too. I swear, people are obsessed with "machine learning" even when the domain really isn't suited for it.

Re: How to recognize AI snake oil [pdf]

#214
post #213
post #209

Earlier quoted context omitted.

That sounds odd, like they don't really need machine learning (unless it is to snare investors?).

Snaring customers too. I swear, people are obsessed with "machine learning" even when the domain really isn't suited for it.

I sometimes wonder if management and engineers don’t have more in common than acknowledged. Publications such as Harvard Business Review have huge coverage of things like managing AI, and being able to say you managed an “AI project” might mean something.

Re: How to recognize AI snake oil [pdf]

#215

Earlier quoted context omitted.

Even companies like Facebook, Apple and Google employee humans to do work that people believe is done by "computers" and non of the companies seem keen on informing the public that they do in fact have humans scanning through massive amounts of data. So perhaps it is in fact cheaper, or the problems they face remains to hard for current types of AI. Given the number of people Facebook employees to censor content and…

I'm aware of moderation. What else?

Recaptcha is probably one you've actually interacted with, but even then you're mostly reinforcing existing predictions. But other applications within Google Maps are things like street number recognition, logo recognition, etc. Waymo contract out object detection from vehicle cameras and LIDAR point clouds. Google even sell a data labeling as a service.

I believe Google Maps has a lot of humans who tidy up the automated mapping algorithms (such as adjusting roads).

Annotation is time consuming and therefore extremely expensive if you have a $100k engineer doing it.

https://www.forbes.com/sites/korihale/2019/05/28/google-micr...

Re: How to recognize AI snake oil [pdf]

#216
post #129

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

This reminds me of a job interview I was on. I was asked about how I would use AI/machine learning for their problem space. Since they seemed to be smart and level-headed, I answered honestly, "Pick something unimportant, use a machine learning algorithm just to get familiar with the tools, ignore the result unless it happens to work, then put machine learning in your marketing materials. But keep track of it, and if…

I think generally that is how many products and features work.

Re: How to recognize AI snake oil [pdf]

#217
post #187
post #156

Earlier quoted context omitted.

> Supervised learning in machine learning is nothing remotely like a human teaching anyone anything. I disagree, I think it's exactly the same. As an example, a human teaching a human how to use an orbital sander to smooth out the rough grain of a piece of wood. The teacher sees the student bearing down really hard with the sander and hears the RPM's of the sander declining as measured by the frequency of the sound.…

But that's not at all how "supervised learning" works. You would do something like have a thousand sanded pieces of wood and columns of attributes of the sanding parameters that were used, and have a human label the wood pieces that meet the spec. Then you solve for the parameters that were likely to generate those acceptable results. ML is brute force compared with the heuristics that human learning can apply. And M…

[deleted]

Re: How to recognize AI snake oil [pdf]

#218

Earlier quoted context omitted.

Right-- but often an NPC is merely scenery or traffic, in the sense that their behavior does not compete with your interests. What I found interesting about the football example is that the random strategy of the NPC-oach both suggested a deeper intelligence AND proved to be an apparently effective opponent. Beyond reinforcing our tendency to project, as you say, a personal history on random behavior, it also highlig…

It's how in Rock Paper Scissors you can probably do better against a smart opponent by playing randomly than by trying to trick them. At least they can't get in your head, because there's nothing there. You won't do better than chance, but at least you can't do much worse.

So, use a random strategy and measure the entropy of your opponent's behavior. And modify accordingly, as needed.

A lot of the best ml right now is effectively about making better conditional probability distributions. You always get random output, but skewed according to the circumstances, and sharp according to confidence in the result.

Re: How to recognize AI snake oil [pdf]

#219

I don't have time to read the entire paper but I would like to share an anecdote. I worked at a company with a well staffed/funded machine learning team. They were in charge of recommendation systems - think along the lines of youtube up next videos. My team wanted better recommendations (really, less editorial intensive) so the ML team spent weeks crafting 12 or more variants of their recommendation system for our c…

Since we are sharing anecdotes, I can report it's been 20 years of buying stuff on the internet and the combined billions of ad tracking research dollars spent by Amazon and Google have not yet come up with a better algorithm than to bombard me with ads for the exact same thing I just bought .

Hey, personal experience and to be fair for them, Google does occasionally give me ads about a Haas CNC machine. I really want one. But I don't have the disposable $100k for one and I don't have the space... nor do I have 3 phase power. But I do want one and haven't bought one. So, good on them, right?

Re: How to recognize AI snake oil [pdf]

#220
post #187
post #156

Earlier quoted context omitted.

> Supervised learning in machine learning is nothing remotely like a human teaching anyone anything. I disagree, I think it's exactly the same. As an example, a human teaching a human how to use an orbital sander to smooth out the rough grain of a piece of wood. The teacher sees the student bearing down really hard with the sander and hears the RPM's of the sander declining as measured by the frequency of the sound.…

But that's not at all how "supervised learning" works. You would do something like have a thousand sanded pieces of wood and columns of attributes of the sanding parameters that were used, and have a human label the wood pieces that meet the spec. Then you solve for the parameters that were likely to generate those acceptable results. ML is brute force compared with the heuristics that human learning can apply. And M…

One of the columns of sanding parameters is the sound of the sander.
Post reply on HN