[flagged]
Evidence of inconsistencies in evaluation process and selection of winners
41–50 of 338 posts
Re: Evidence of inconsistencies in evaluation process and selection of winners
#42I don't know about this exact competition but overall fair hackathons have been killed by AI. It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners. It used to be about human skill, now it's about ideas and of course insiders are the main winners.
Can you share any examples of that? I'd love to see them myself.
Re: Evidence of inconsistencies in evaluation process and selection of winners
#43> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" This is the most blatant Claude line, or as Claude would put it, the smoking gun.
But is it load-bearing?
Broadly, I keep thinking about this over last year or two: while LLMs have nearly eliminated the bar for slop and coding slop, the reviewers are still expected to perform their job diligently. The asymmetry here is extremely taxing for reviewers of all AI generated content. And this is one thing that AI can't help with (as with any statistical process that lacks world understanding and grasp of logical inference).
That's why I fully support Arxiv's tough stance on the AI use responsibility.
Re: Evidence of inconsistencies in evaluation process and selection of winners
#44Earlier quoted context omitted.
AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.
People's workday has been transformed by it. But I fail to see actual transformation beyond "more crap, faster." AI hasn't done anything we couldn't already do. It's just doing it faster and with more mistakes.
AI is capable of performing a lot of grunt work reliably. Still must be reviewed. But a big productivity gain over doing everything yourself.
Re: Evidence of inconsistencies in evaluation process and selection of winners
#45What's up with all the AI generated responses on that page?
Re: Evidence of inconsistencies in evaluation process and selection of winners
#46What's up with all the AI generated responses on that page?
Re: Evidence of inconsistencies in evaluation process and selection of winners
#47AI is not there yet, instead of working hard, everyone is choosing the easy way out.
AI slop wins prize, I wonder if Ai slop read it also. would not be surprised. however not to judge anyone, I think we are seeing slop everywhere, hope some things still require hard blocks for low quality.
its difficult to justify lack of attention and details
Re: Evidence of inconsistencies in evaluation process and selection of winners
#48[flagged]
Re: Evidence of inconsistencies in evaluation process and selection of winners
#49"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?" That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.
We've had about a century now of science-fiction literature hyping up AI as a higher intelligence that is based solely on some ill-defined yet universal system of "logic" and is therefore not prone to human flaws such as pride, hate, envy, lust, etc. Now it has become extremely apparent that was always an unsubstantiated assumption but its too late because there are billions of people primed to never question the mac…
People interact with AI, talking to it like a human. Of course they start to believe it’s rational like a human.
LLM does all of the entry level tasks better than the students. Partially because the answers are in the training set, and partially because it has gotten that good now. Hard not to start to believe it is “competent”.
I personally have had a real hard time getting traction talking about making sure the way we assess AI is not based on material it has trained on. YMMV as always, but I think the large training corpus contributes to the (unreasonably) high level of faith in the machine.
Re: Evidence of inconsistencies in evaluation process and selection of winners
#50It’s a shame that Arvix (and once thoughtful places like Kaggle) are used for self-promotion. I get people want to work at an AI lab but slopping it in public in this manner is counterproductive to the original intended purpose of these places.
Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.