Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

241–250 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#241

Please note this Post was just renamed without my involvement from: Blatant AI slop just won a 25K USD Deepmind Kaggle Grand Prize into "Evidence of inconsistencies in evaluate process and selection of winners"

That's because HN doesn't allow editorialised titles, which means injecting your own opinion into the title. The author of the linked article's opinion is still allowed. If you wanted to give your own title you would have to write the article. Even though your title was objectively more correct and useful than the one it's been renamed to.

I am the author of the linked article, so you say my opinion should be allowed in the title?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#242
post #24

Earlier quoted context omitted.

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

Your perspective is a short arc. “Look what I can do now and look it made me way more productive.” I have no doubt it is true, you are on the up right now. However thinking of the long arc is important to, even though it has no consequence for you right now. AI is a force multiplayer and scarily dangerous in the wrong hands. We can already see by these discussions how uncertain things are. Just food for thought.

On the long arc is a better world: post-scarcity, accelerated advancement in important fields (such as medicine), and just general improvement in lifestyle for the greater global population. The primary scare is that the higher capabilities will be restricted to some small subset as most resources currently are today, but the increasing release of open models is helping to counteract that possibility.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#243
post #28

> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" This is the most blatant Claude line, or as Claude would put it, the smoking gun.

But is it load-bearing?

You're absolutely right to push back on that.

Bottom line⸻it's not load-bearing, it's structural.

And honestly⸻that's not nothing.

No loads. No bears. Just structure.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#244
post #46
post #19

What's up with all the AI generated responses on that page?

That whole thread had a strong stench of AI about it, across multiple participants.

The whole thing seems like a downward spiral of brain rot. I don't understand what the reward is for participating, it seems like a total waste of human effort.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#245

I thought Kaggle was a website where you download dubious CSV files of annualized bean consumption in Bolivia, or whatever. Was Kaggle ever a reputable source of original research, or a source of anything with any provenance at all? That would be news to me. The fact that 25 grand was involved this time is unique, I guess.

They also host the arc prize challenge with prices up to 850k https://www.kaggle.com/competitions/arc-prize-2026-arc-agi-3...

Re: Evidence of inconsistencies in evaluation process and selection of winners

#246
post #197

Earlier quoted context omitted.

In startups it's extremely common to have management still write code. Hell I'm CTO and I write a lot of code.

I know there can be CTOs who write a lot of code and be good, but I hope to God you're not my CTO because he is a nightmare in this regard right now. We aren't even a startup anymore we have 80 engineers and he pushed 200k lines this week

Whoa, 200k lines is a lot. Without knowing the details I wouldn't want to pass judgment, but that would be a flag to say the least. I don't know how he is reading all that code before pushing...

There's a danger in it too because (whether it should or not) CTO title carries a level of weight that might result in less scrutiny than an IC (which it should not IMHO). I hope he is encouraging objective reviews and not pushing stuff through on his title.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#247

Earlier quoted context omitted.

Decision trees are one of the oldest forms of AI! There are even algorithms to automatically form decision trees, though you don't need ML to have AI.

Those algorithms were only called machine learning and it's the opposite. Before the marketers got ahold of it we reserved AI for the actually intelligent, sentient, full strength vision of intelligent systems. We now talk about general AIs and A[Super]Is.

That's not true. AI was widely used to refer to the decision systems and state machines that produced NPC behaviour in video games, and I'm sure many other things than just science fiction

Re: Evidence of inconsistencies in evaluation process and selection of winners

#248

Earlier quoted context omitted.

That's because HN doesn't allow editorialised titles, which means injecting your own opinion into the title. The author of the linked article's opinion is still allowed. If you wanted to give your own title you would have to write the article. Even though your title was objectively more correct and useful than the one it's been renamed to.

They also removed the fact it was a competition hosted by a prestigious AI lab

your title was rewritten by AI, correct? :)

the slop doesn't bother me as much as the bowdlerized nature of said slop (and search results). sure, there was a lot of slop after the Lapsing of the Press Act (1695), but it was human slop with convictions and syphilis!

Re: Evidence of inconsistencies in evaluation process and selection of winners

#249

Earlier quoted context omitted.

That's fair and your good right. However know that my frustrations stems from spending two and a half days feeling like in the story of the emperors new clothes digging into this shit that was made king by a number of employees of one of the most prestigious AI labs.

Then write a blog post and make your point there, not in the title of the submission.

What do you think my two long posts in the discussion forum of kaggle are? I am the author and the evidence I present very much justified the original title.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#250
post #28

> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" This is the most blatant Claude line, or as Claude would put it, the smoking gun.

But is it load-bearing?

You're asking the right questions.
Post reply on HN