Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

31–40 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#31
AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#32

[flagged]

Someone might want to downvote you because you just state something which is very controversal and you do not add any arguments to your 'empty' comment.

Its hard to even have a discussion because someone else needs to give you enough content like ask you first why do you even think that.

So how do you define AI? LLMs? GenAI stuff?

What is 95%? Does it mean that these 5% are unable to disrupt industries or does it mean for you that these 5% will change the world as we know it but stil 95% of other AI stuff is useless?

I personally think that AI/AGI progress is faster than i expected it, I think its very useful already today, I also think we still need to build a lot of obvious stuff (like proper AI Agentic Platforms), more hardware, cheaper hardware etc. but the way quite clear, but some peple might think the current state is the AGI future people talk about it but I think we will only see this in 4/5-15 years and then it will have disrupted a lot.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#33
post #28

> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" This is the most blatant Claude line, or as Claude would put it, the smoking gun.

But is it load-bearing?

I thought it was a belt and suspenders conclusion

Re: Evidence of inconsistencies in evaluation process and selection of winners

#35

"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?" That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.

We've had about a century now of science-fiction literature hyping up AI as a higher intelligence that is based solely on some ill-defined yet universal system of "logic" and is therefore not prone to human flaws such as pride, hate, envy, lust, etc. Now it has become extremely apparent that was always an unsubstantiated assumption but its too late because there are billions of people primed to never question the machine.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#36

It’s a shame that Arvix (and once thoughtful places like Kaggle) are used for self-promotion. I get people want to work at an AI lab but slopping it in public in this manner is counterproductive to the original intended purpose of these places.

Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#37

AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#39
post #33
post #28

Earlier quoted context omitted.

But is it load-bearing?

I thought it was a belt and suspenders conclusion

Honestly? It’s the shape that makes it clearly AI. That’s the quiet admission at the heart of the problem.
Post reply on HN