Live data from Hacker News

A Kaggle Grandmaster cheated in $25k AI contest with hidden code

theregister.co.uk

111–120 of 197 posts

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#111
post #87

It really worries me how many people are so quick to forgive him and tell him so. In my family if someone cheated they got called a cheater and suffered consequences. At least, they would have, if someone did something like that. But my parents didn’t raise mendacious villains. Look at this crap on Twitter: “Everyone makes mistakes. Thank you for the apology”. “Kagglers will still love to have you back” “It's great t…

I'm thinking Kaggle should look to Epic/Fortnite for how to handle cheaters.

A lifetime ban would not be out of the place here, considering it went on to win the competition with no admission until caught.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#112
post #109

Slightly off topic: The nature of the competition is a bit worrying to me. They're essentially letting an algorithm decide which dogs have the best chance of adoption and which to euthanize, aren't they?

THIS! Drop the "slightly"... I can't understand how people focus so much on the competition cheating, and so less on "wait, wtf are they doing here"... I mean, even if they are not deciding whether to euthanize or not based on this, you're still building a system that introduces "good looks" as a factor in a life-and-death decision regarding a living being.

It's not hard to jump from this a system that would use your facebook file to grant or deny medical health coverage or a similar life-and-death thing. Shift the Overton-window a little, push is a few notches further, and you're re-inventing phrenology with deep learning...

This is bone-chilling! I mean the fact that so many people overlook this...

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#113
post #59

I'm not going to defend Pleskov but organizers shouldn't have put out the competition with money attached that can simply be solved by scraping data. Good ML competition in fact should even invite cheats because the end goal is not ML for the sake of ML but rather cracking the prediction problem by whatever shortest path possible.

He scraped the test set's labels. How is that useful? It's not about "ML for the sake of ML", it's the equivalent of stealing the answers to a math test then writing them down. Why should that be rewarded?

The goal of this competition is to build a system (using ML or not) which is useful for predicting how quickly pets will be adopted. Any information used during the competition should be realistically available at inference time for future predictions... clearly, the expected answer cannot be available at the time you're trying to predict when a pet will be adopted.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#114

They would have able to win and get away with it if they incorporated the knowledge of the external dataset directly into the ML model, provided they had a reasonable estimate on the fraction of overlap between the external data and the test set. A weak version of this would be to just train on the external data in addition to the provided data. A stronger version would train regularly on the provided training data a…

This is a really good point!

Considering the guy was smart (he is kaggle grandmaster), I would really like to know what prevented him from training on the scraped data, and what motivated him to obfuscate the known sample lookup.

Maybe there's some technicality they made it impossible to tune the model on the additional scraped training data.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#116
post #109

Slightly off topic: The nature of the competition is a bit worrying to me. They're essentially letting an algorithm decide which dogs have the best chance of adoption and which to euthanize, aren't they?

Are you sure? Aren't they just trying to show the most appropriate pets to the most appropriate people in order to make adoptions quicker? Where do you read the euthanization part?

Unfortunately, most pet shelters get more pets than their facilities are able to handle. Unless they're a no-kill shelter (in order to be considered one they need to kill <10%), they will need to euthanize in order to make room for incoming animals. This is as opposed to promoting adoption, fostering, etc. An algorithm that can prematurely decide what animals will be adopted, or have the best chance to be, can create a perverse incentive to euthanize early the less optimal pets.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#117
post #109

Slightly off topic: The nature of the competition is a bit worrying to me. They're essentially letting an algorithm decide which dogs have the best chance of adoption and which to euthanize, aren't they?

Having read the original description of the challenge[1] I don't think you're correct. I think what they're doing is trying to identify the types of photos and descriptions that are successful so they can do more of them. Like: Do you want the dog bounding through a field or do you want them snuggling up on someone's lap in the photo? That sort of stuff.

I mean, you're right, you could just run the tool, find the ones that are unlikely to get adopted and euthanize them, but I don't think there's any reason to believe that's actually their intention.

[1]:https://www.kaggle.com/c/petfinder-adoption-prediction/overv...

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#118
It was briefly discussed about 7 days ago.

https://news.ycombinator.com/item?id=22045696

I posted there, same self-addressed question, that I cannot figure out the answer to...

It seems that intensives to cheat, and environment where 'means justify the ways' -- are overpowering.

For people who are naturally gifted, successful at young age -- why cheat?

Was this historically, always like this?

These insensitive to cheat, to gain unfair advantage, to treat life opportunities without any 'honor code' just seem to be so pervasive now, it seems.

There is a cheating scandal every other week involving most prestigious institutions, competitions, and so on.

These incentives to cheat, basically destroy from inside our commercial model, academics, judicial system, political system and probably military too.

This also creates a new type of powerful currency, and therefore the 'billionaires' in that currency have infinite power -- and that currency is 'dirt on somebody'.

Dirt on somebody who cheated before -- forever makes the cheaters into tools of injustice.

---

Public shaming is reactive, there we need something more proactive at various points. There are needs to be incentives for work verification, as an example.

I also think it is unfortunate but at least civil/commercial law in many countries is pretty much riddled with 'more expensive lawyers produce better results'. And it skews society into basically thinking 'anything goes, really. means justify the ways, and cheating something one can get away with'

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#119
post #53
post #29

Earlier quoted context omitted.

This only works when you don't win. You have to upload your code, including model training code, they won't accept a trained model binary.

If they were smart, they would have used the whole set for hyperparameter tuning. That would be essentially undetectable.

Not if the underlying model was bad, no tweaking of hyperparameters can change that. It's safe to say, he probably did consider this (and it probably didn't work well enough).

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#120

It was briefly discussed about 7 days ago. https://news.ycombinator.com/item?id=22045696 I posted there, same self-addressed question, that I cannot figure out the answer to... It seems that intensives to cheat, and environment where 'means justify the ways' -- are overpowering. For people who are naturally gifted, successful at young age -- why cheat? Was this historically, always like this? These insensitive to che…

>Was this historically, always like this?

Of course. 20$ bills left on the floor are bound to be picked by someone even if most won't.

..At least as long as the expexted punishment comes down to less than the value of said bills.

Post reply on HN