Live data from Hacker News

A Kaggle Grandmaster cheated in $25k AI contest with hidden code

theregister.co.uk

161–170 of 197 posts

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#161
post #156

Earlier quoted context omitted.

That's not a hack, that's blatant cheating, their solution litterally looked at the anwsers. It bypassed the ML model prediction, so that's not a ML solution, which to my understanding was the constraint of the competition. And in the end, that solution is useless for the adoption site, since the objective is to get adoption predictions (the animal has not been adopted yet). I'm confused too, by how can anyone think…

Their solution made use of the data available to them in a "prize" that was nothing more than a made-up competition for fun. Did the rules expressly say somewhere that one could not make use of available datasets?

Kaggle rules routinely restrict use of datasets to approved ones, almost certainly yes but you would have to check to be certain.

The prize was $10,000, not "fun".

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#162
post #153

Earlier quoted context omitted.

I'm confused by this reaction here - this was a _brilliant_ hack of the system and I think his work should be celebrated. Was it in keeping with the intention of the competition? Of course not. Were lives threatened by their creative solution to the competition? Also no. At the end of the day it was a fun, inventive approach to a made-up problem. So what's the problem?

The problem was that there was a prize at stake, and everyone else was playing by the rules.

Not trying to be argumentative here, but is what they did actually against the rules? I'm not at all familiar with this competition which might be why I'm not quite so worked up about this, so maybe I have misunderstood something. Did the competition really require that they only train on the provided data?

I compare this to my favorite sport, Formula 1 racing. In F1, teams of engineers with nearly unlimited budgets spend an absurd amount of effort doing everything in their power to bend the regulations (the "formula") to squeeze out some extra advantage.

For example, in this past season Ferrari was suddenly outperforming the pack (and their own recent performance) and it was clear something had changed on the car, they had power in places they didn't before. What finally came down is a clarification of the rules around fuel-rate metering, without directly calling out Ferrari. After the clarification, Ferrari power was back where it used to be. Nothing more was said of the matter by the FIA.

What we all _think_ happened is that Ferrari, knowing the fuel rate meters ran at 10kHz, discovered they could pulse their fuel pump so that the low-end of the flow rate cycle happened during that sampling interval. This means they could increase their overall fuel rate beyond what was technically allowed, due to how that technical requirement was being measured on the car (and reported back to the FIA).

Is it in keeping with the spirit of the rules? Of course not! Does it make for an interesting engineering puzzle on top of an already-exciting sport? Sure does!

Clearly I'm in the minority here, but I think this sort of problem-solving approach can be useful. If you're looking to compete against a field of entrants who are all looking for obvious and well-understood approaches to solving the problem at hand, I think sometimes the best solution to stand out is to look where the other teams aren't looking.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#163
post #59

I'm not going to defend Pleskov but organizers shouldn't have put out the competition with money attached that can simply be solved by scraping data. Good ML competition in fact should even invite cheats because the end goal is not ML for the sake of ML but rather cracking the prediction problem by whatever shortest path possible.

"Winning" by anything that can reasonably called cheating, as in this case, does not advance the general state of the art. Innovation is best served through appropriate rules and competition structure.

You see the same thing in grad school.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#164
I was pretty surprised by how common this type of behavior is on Kaggle. I work in machine learning and data science, but I don't use Kaggle much because it quickly became clear competitions boiled down to who could eek out the last hundredths of a percent in accuracy from models trained for weeks on multithousand dollar machines, and because the behavior described in the article was surprisingly common.

That said, the site is a fantastic resource for datasets. Lots of fantastic data uploaded both from old competitions and by the community.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#165
post #162

Earlier quoted context omitted.

The problem was that there was a prize at stake, and everyone else was playing by the rules.

Not trying to be argumentative here, but is what they did actually against the rules? I'm not at all familiar with this competition which might be why I'm not quite so worked up about this, so maybe I have misunderstood something. Did the competition really require that they only train on the provided data? I compare this to my favorite sport, Formula 1 racing. In F1, teams of engineers with nearly unlimited budgets…

The issue is that the entire reason the competition exists is because the company is sponsoring it and putting forward the prize money so that the top performing models can then be put into production, thereby solving some problem the company has. This type of cheating is dishonest and against the spirit of the competition, but it also defeats the entire purpose of the exercise. Simply keeping a lookup table of answers for the data isn't machine learning, and will not generalize into a production system. As stated in the article, without these hacks, he wouldn't have even placed in the top 100.

To use your F1 analogy, this isn't the equivalent of tweaking the cars in whatever way possible is within the rules. This is the equivalent of completely cutting across the grass and bypassing 90% of the track, which is indeed illegal and would get you penalized.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#166

How is this not criminal fraud for $10k? He deserves to go to prison. "Boo cheater" would be an appropriate response if he did it purely for ranking, not for money (or something easily sold for money).

How _is it_ criminal though? Please do tell which law his team broke here.

Put down the pitchfork and calm down. The guy is a raging a* and a cheater, but a criminal he is not. He lost his job, was publicly shamed and his reputation is tarnished basically forever. I think he got enough coming his way.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#167
post #107

Earlier quoted context omitted.

Your parent is saying kaggle takes the training code and runs it themselves. This makss secretly training on the known right answers impossible.

There is no such thing as "training code". You have a model that you have trained on some provided data - the training set. You give kaggle this model. Kaggle grades your model on some different data your model has never seen. The better your model classifies this data it hasn't seen the higher it scores and the more money you win. So again if you trained your model (code) on a training set that you have illegally ob…

> There is no such thing as "training code".

I'm by no means an expert in ML, but my understanding is there's some code that is run to train the modal. I meant that by "training code". My regrets if my terminology was unclear.

> They took the official training set. Said, we need more. And scraped websites to get a bigger, illegal training set. This is against the rules and is cheating. They got caught.

No, this is wrong

I suggest you look at one of the other comment where people have explained why this is wrong. They did a better job than me.

https://news.ycombinator.com/item?id=22124760

https://news.ycombinator.com/item?id=22126489

https://news.ycombinator.com/item?id=22124193

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#168
post #163

Earlier quoted context omitted.

"Winning" by anything that can reasonably called cheating, as in this case, does not advance the general state of the art. Innovation is best served through appropriate rules and competition structure.

You see the same thing in grad school.

Indeed you do, and I know someone who got screwed by a couple of plagiarizing cheaters. This probably contributed to his abandoning his studies and possibly to his suicide. I take a dim view of suggestions that cheating is reasonable, let alone that it is the smart option.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#169
post #17

For an HN audience, the "How Bestspotting cheated" post on Kaggle may be a better place to start: https://www.kaggle.com/bminixhofer/how-bestpetting-cheated

It's an odd way to cheat too, if they realised they had data from the validation set, couldn't they have over-trained a model with the validation set in the training data?

Yes, that would have been harder to detect if it is allowed. My understanding is that for a kernel competition (like this one), you can't use model weights that you've trained outside the kernel. Oddly, I can't find a rule explicitly prohibiting it.

Re: A Kaggle Grandmaster cheated in $25k AI contest with hidden code

#170
post #158

A large number of commenters fundamentally misunderstand what happened. They are saying "why did he upload his scraped training data? Why didn't he just train on it and upload the resulting model?" If you are making this argument, it means that you don't understand the contest. User rahimnathwani explains: > In this competition, the training code was run on Kaggle's system, so you'd still need to smuggle in the extra…

AFAIK it's impossible to smuggle it in a way that you can't be caught. Maybe the use of MD5 made it easier to see but if examined I don't think there's a totally hidden solution.
Post reply on HN