Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

151–160 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#151
post #29

I don't know about this exact competition but overall fair hackathons have been killed by AI. It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners. It used to be about human skill, now it's about ideas and of course insiders are the main winners.

[flagged]

[deleted]

Re: Evidence of inconsistencies in evaluation process and selection of winners

#152
post #49

Earlier quoted context omitted.

We've had about a century now of science-fiction literature hyping up AI as a higher intelligence that is based solely on some ill-defined yet universal system of "logic" and is therefore not prone to human flaws such as pride, hate, envy, lust, etc. Now it has become extremely apparent that was always an unsubstantiated assumption but its too late because there are billions of people primed to never question the mac…

I really don’t think that’s what’s going on here. People interact with AI, talking to it like a human. Of course they start to believe it’s rational like a human. LLM does all of the entry level tasks better than the students. Partially because the answers are in the training set, and partially because it has gotten that good now. Hard not to start to believe it is “competent”. I personally have had a real hard time…

A lot of philosophy starts from the fundamental observation that, to solve a problem, you could either solve a problem, or state that the problem is ill-posed in some way (and dissolve it). Either answer the question, or question the question.

It's not a new problem in some sense. If you've dealt with really smart but really arrogant friends, they might jump ahead 10 steps and assume your rebuttal, posit theirs, assume your rebuttal to their posit, etc. etc. without... actually taking the time to listen carefully. On the national scale, this looks like forced trust in government authorities about what is "objectively best".

People need to get it through their skulls that, even if an AI, or any intelligence, could even solve the damn Riemann Hypothesis: if it's wrong, it's wrong. Of course, I think all of us know the objection - we see it on hackernews all the time. "You guys are just stupid contrarians who can't understand AI's deep reasoning". OK. The second inference? "therefore you are unable to govern yourselves properly - your concerns are all fallacies, misunderstandings, bad for you, etc.".

You might think that the second inference is extreme and nobody actually believes that, but as always, it's a gradient. Before AI, you might've had an extremely strong sense of self. Now? You look at OpenAI solving open math problems left and right on a public foundation model, and you think, "Maybe I should just trust it more. If I spend cycles thinking, it's probably going to outdo me anyways." The AI model silently makes 5 different assumptions and transformations? "Well, maybe it was rational in the space of tradeoffs to do that. The AI knows best, after all". You might be thinking of an architecture with 5 different key constraints based on lived experience, in which the AI keeps misunderstanding. "Oh, well, this genius-level mathematician/programmer AI isn't understanding my words - surely I must be mistaken, right? It's only humble to think that way".

I can't convince people otherwise though. After all, I can't "prove" that you should have a backbone when talking to AI. it could just as easily be "you're arrogant, this machine is in the top of all academic fields and is coming for all white collar jobs, who's to say you're right about anything?" All I can say is, there's a reason why Dostoevsky is one of my favorite authors.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#153

Earlier quoted context omitted.

IMO that’s awesome. I like when folks are clever. Just modify the rules next go around. It’s a contest judged by and LLM. Not sure why we would take it that serious.

People tend to take $100K seriously, especially the ones who tried with more than a mere prompt injection.

Sure but it’s one of those things imo that happens. Google should have been more rigorous with their judging. They may lose face for future hackathons or simply it becomes a gating item in the rules. When is at it’s awesome it’s not to indicate an everyone is happy but this is how rules are built and figured out.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#155
post #14

Sadly, the major ML/AI/NLP conferences are being inundated with AI slop papers. That will arguably have a bigger impact on the quality of research moving forward.

" impact on the quality of research moving forward. " It'll affect everything that depends on manipulating symbols! The enormous body of knowledge humanity has accumulated over the past 6.000 years or so is about to be flooded with slop!And That's the real threat genai poses to humans that i don't see anyone talking about..

Yes! People judge the "usefulness" of the AI generated material by how much they can get away with using it themselves, now. But what happens when more and more is built on top of these generated falsehoods. The errors and chaos bubbles up exponentially.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#156
post #63

Earlier quoted context omitted.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

With the exception of _one_ company that I worked at, pretty much every[0] company was a struggle between engineering and management. Engineering wants to get the software correct, and management wants to fire-hose features into the market. Most of the time (so more than half, at least), management tends to have a compulsion to mindlessly imitate what other companies/competitors are doing, usually without prioritizat…

Management is just responding to idiot end-users. I have been on plenty of sales calls where customers ask if features X,Y,Z are available, knowing that there’s a 99% chance they’ll never need them, but they ask anyway just because they’ve heard that someone else used a feature like that at some point in the past. If it’s not, they just assume the software is inferior.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#157

I think that a lot of software engineers are using LLMs and a lot of very popular tools are developed by, or are assisted by, LLMs. Is this not just going to be a thing going forward? This feels akin to traditional artists getting angry at digital art winning competitions when that was a new concept. We're simply in the early stages of a paradigm shift, no?

The issue we’re dealing with is that the tool is as likely to write confident sounding, well-written but completely wrong everything and if you don’t know the difference you might accidentally give it a gold medal. Like a chainsaw: yes the tools are useful and will be used in the future, but we may not want to use chainsaws to carve up the turkey.

If the judge doesn't know the difference, why are they judging a competition, and also, why is the competition esteemed?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#158

AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

Ego is also at play here. In inherently competitive fields like academic science using LLMs to get results more quickly is an alluring siren.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#159
post #92
post #63

Earlier quoted context omitted.

With the exception of _one_ company that I worked at, pretty much every[0] company was a struggle between engineering and management. Engineering wants to get the software correct, and management wants to fire-hose features into the market. Most of the time (so more than half, at least), management tends to have a compulsion to mindlessly imitate what other companies/competitors are doing, usually without prioritizat…

> It very frequently feels like management is making strategic decisions after snorting a long line of social-media-psychosis and TED talks. It is remarkable that investors have any faith in such founders/entrepreneurs at all. My guess is the causality is usually * the managers are pursuing things because their investors (/government ministries) hinted it was the future after snorting a line of "TED talks" and "socia…

One piece of managment advice I've gotten is to be very careful about what ideas you're throwing out there, because people can very easily get the wrong idea about the priority, and this can get worse as multiple layers get involved.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#160
post #57

Earlier quoted context omitted.

There is also the "You aren't paid to think, you are paid to do exactly what I tell you, nothing more or less!" school of management. I'm not sure how prevalent this attitude is now but it was very common in the 90s and 2000s. The AI and the bosses that want you to use it all speak from positions of authority and confidence. That's their right, granted to them by their position. You don't speak that way because as a…

So your position is that people actually want to do more work, but their managers are forcing them to work less? I don't buy it.

No, you missed the point. The GP was talking about causes for the situation where people are apparently prone to "outsourcing their thinking", resulting in low-effort AI slop being produced by entities that really should know better. The style of management I describe destroys the employee's confidence and agency, so they are much more likely to just submit and blindly accept whatever management/AI/etc. tells them than to do any sort of critical thinking.
Post reply on HN