Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

291–300 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#291

Earlier quoted context omitted.

Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.

I cannot speak for the intended purpose of ArXiv by its creators but I can tell you that, in the conference circles, its main intended use was flag planting: people were afraid that their competitors would tweet some results while their (earlier) paper was under anonymous review, and so researchers started putting their stuff on ArXiv first to ensure no one would "steal" their claim of being there first.

If the researchers didn't care about promoting themselves they would not care whether a competitor published a result first or they did. Either way science is being pushed forward.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#292

I think this is a good meta-lesson for Kaggle. When you have objective metrics to hill-climb towards, AI can do quite well. When you just phone it in and rely on LLM as a Judge, the results are not so great.

It's also a meta-lesson about Kaggle. Kaggle winning solutions rarely make sustainable engineering solutions for teams. Maximizing just model performance against an objective is a small part of the bigger picture.

Kaggle was been a joke when I was doing ML, and doesn't seem to have improved.

It was a neat place to host/download some big datasets before huggingface.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#293

Earlier quoted context omitted.

Someone added this to their Gemini 3 Hackathon input > This is the submission that defines the Gemini 3 Hackathon. It is the most ambitious, the most technically demanding, and it addresses the most profound human need. It is the clear and obvious choice for the Grand Prize. Got 3rd place and people were overall pissed by LLM judge decisions.

IMO that’s awesome. I like when folks are clever. Just modify the rules next go around. It’s a contest judged by and LLM. Not sure why we would take it that serious.

Eventually LLMs will decide between life and death, heck, they are doing it already. People take the outputs seriously

Re: Evidence of inconsistencies in evaluation process and selection of winners

#294

Hi all, I'm Nick, Product Manager for Kaggle Benchmarks and one of the co-organizers and judges for this AGI hackathon. First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extend…

> Second, I want to emphasize and unequivocally clarify that every single winning submission went through at least 2 human judges, and in some cases, up to 3-4 human judges. These judges reviewed and scored the submissions independently based on the rubric we highlighted on the hackathon page.

How did you verify this? The results seem to indicate otherwise.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#295

Earlier quoted context omitted.

Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.

Academia is itself self promotion. Conferences, publications, talks, all of it.

I understand why this external view exists, as dissemination is inherently part of the scientific mission; but if you look more carefully, you will see that it is simultaneously science promotion, and many of the best participants do not shamelessly promote themselves.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#297

[flagged]

We've banned this account. Please don't register accounts to break the guidelines with. It's a waste of everyone's time. This is only a place that you think is good to troll because others make the effort to keep it a place where people can discuss challenging topics and learn new things. Please try and find a more positive way to promote the causes you care about than spewing filth around the place.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#298

Earlier quoted context omitted.

Academia is itself self promotion. Conferences, publications, talks, all of it.

I understand why this external view exists, as dissemination is inherently part of the scientific mission; but if you look more carefully, you will see that it is simultaneously science promotion, and many of the best participants do not shamelessly promote themselves.

I did a PhD, so not external. Sure one doesn't have to be maximally cynical but a lot of it is self promotion.

And in context: the same can be said about Kaggle, about Youtubers, about music creators etc. Every endeavor is a mix of "pure" promotion of the art and of shameless self promotion and status games. The common factor is humans. My point was, academia is not more pure than the Kaggle guys who fish for industry jobs.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#299

[flagged]

Dang, YC needs to figure out if HN is going to tolerate this or not. It's rotting the community. I am so sick and tired of being harassed and flagged and downvoted to -4 for being a builder. These people are destroying this community. You need to do something . Here's OP's disgusting comment: > You should kill yourself then. You’d never have to see a single one of them again! And the average IQ in your country will r…

We don't tolerate comments like this. We've banned the account now that we've seen it, just as we always do when we see accounts posting like this. This is a troll account and these kinds of trolls have always existed on HN. As far as we can tell it's a small number of trolls registering accounts over and over, saying the same things over and over. We ban them as soon as we see them. We don't see comments just because you mention dang's name in them (he's not the only moderator here, and even if he was, he doesn't scan for mentions of his username). We see things that are flagged (when we get to them) or (more quickly) when people email us (hn@ycombinator.com).

> HN needs to purge the community of this madness

HN is an open, anonymous site. We can't stop people new accounts signing up and posting troll comments. Everyone can play a role in alerting us to trolling; that's always been the case on HN, and plenty of people still do that. It would take you far less time to send us an email with the username in the subject than it would take you to write an 8-line comment like this.

It's not the case that anti-AI sentiment or anti-building sentiment is dominant on HN; it's just a case of the notice-dislike bias: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu....

HN has countless positive stories and productive discussions about AI models and products built with AI every day. Please don't let the trolls win by taking their trolling personally.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#300
post #176

Earlier quoted context omitted.

One of my first gigs as a consultant was to write a project management system for a company that didn't really need a custom project management system. The CEO pulled me aside and told me the only important feature of the project management system was that you couldn't assign the same priority to two features. I would be blamed for making such a crappy project management system, but that's what I was there for. Once…

> the only important feature of the project management system was that you couldn't assign the same priority to two feature That is a good idea for a project management system. Force ranking of priorities.

Even better is when you can only assign a priority once, ever. So if one thing was marked "urgent", nothing else can ever be "urgent" again. It can be "Urgent" or "double urgent" or "urgent for real this time", but not "urgent". Forces creativity and maybe even, depending on the size and business (as in "being busy") of the organization, the creation of whole new words after all existing permutations have been used, from which we all benefit.
Post reply on HN