Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

21–30 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#21
post #5

It was probably scored by AI too. Same reason why slop-filled resumes apparently work better these days.

> Same reason why apparently slop-filled resumes work better these days. It'll also filter the kinds of employers that'll hire such candidates, so people that do this will likely land in terrible workplaces.

It’s a nice thought, but it’s probably not true as the AI becomes integrated into standard hiring tools

Re: Evidence of inconsistencies in evaluation process and selection of winners

#22

[flagged]

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

Apparently "being transformed" has "been transformed" because most things look a hell of a lot like the same thing these days.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#23

[flagged]

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

People's workday has been transformed by it. But I fail to see actual transformation beyond "more crap, faster."

AI hasn't done anything we couldn't already do. It's just doing it faster and with more mistakes.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#24

[flagged]

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

Your perspective is a short arc. “Look what I can do now and look it made me way more productive.” I have no doubt it is true, you are on the up right now.

However thinking of the long arc is important to, even though it has no consequence for you right now. AI is a force multiplayer and scarily dangerous in the wrong hands. We can already see by these discussions how uncertain things are.

Just food for thought.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#25
> "Finding 1: Scale Buys Evaluation, Not Control"

The attached paper's (https://arxiv.org/pdf/2604.16009) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition"

This is the most blatant Claude line, or as Claude would put it, the smoking gun.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#27

[flagged]

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

> As it can run in the background doing “chores” while I sleep.

Slightly off topic, but this reminds me of the night family in Rick and Morty.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#28

> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" This is the most blatant Claude line, or as Claude would put it, the smoking gun.

But is it load-bearing?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#29
I don't know about this exact competition but overall fair hackathons have been killed by AI.

It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners.

It used to be about human skill, now it's about ideas and of course insiders are the main winners.

Post reply on HN