Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

61–70 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#61
post #53
post #48

Earlier quoted context omitted.

2024-era take.

LLM's are really good with Django by the way! Must be partly due to the excellent documentation.

For several years one of the most widely used LLM coding benchmarks - SWE-bench Verified - consisted mainly of PRs from the Django project!

Re: Evidence of inconsistencies in evaluation process and selection of winners

#62

Earlier quoted context omitted.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

I just shame people that give slop. Slop PR? Fix the slop. Slop design? I’m not implementing slop, fix it. Innundated with slop PRs? Send half of them to my super and tell him to deal with it. We’ve fired people that wouldn’t get their shit together. Deadlines are being missed because we need to spend more time fixing slop? That’s a planning (management) problem, not mine. Management are the ones that forced everyone…

What if the slop PRs come from your super?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#63

AI is useful. But the amount of people that are simply offloading all of their thinking to AI and blindly accepting the answer is absurd. Kaggle is most likely using ai to assess the submissions and are not using any common sense by blindly accepting the results.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

With the exception of _one_ company that I worked at, pretty much every[0] company was a struggle between engineering and management. Engineering wants to get the software correct, and management wants to fire-hose features into the market. Most of the time (so more than half, at least), management tends to have a compulsion to mindlessly imitate what other companies/competitors are doing, usually without prioritization (so even if feature-parity is a good idea, usually management will want to prioritize whatever the newest feature is, and to put existing work on the back-burner). It very frequently feels like management is making strategic decisions after snorting a long line of social-media-psychosis and TED talks. It is remarkable that investors have any faith in such founders/entrepreneurs at all.

[0]: Various people I know do not even have the luxury of that one good company. Also, it -- unbelievably -- sounds much worse at other companies.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#64

"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?" That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.

We've had about a century now of science-fiction literature hyping up AI as a higher intelligence that is based solely on some ill-defined yet universal system of "logic" and is therefore not prone to human flaws such as pride, hate, envy, lust, etc. Now it has become extremely apparent that was always an unsubstantiated assumption but its too late because there are billions of people primed to never question the mac…

Aside from all the stories where AI does exactly what it was told to instead of what the creators meant? Apart form them, never underestimate British humour's ability to contradict narratives of competence*:

https://www.youtube.com/watch?v=BxWQo_vZgR8

https://www.youtube.com/watch?v=0mfvPHCVMp0

* artificial, political, workplace, nothing is beyond mockery

Re: Evidence of inconsistencies in evaluation process and selection of winners

#65
post #29

I don't know about this exact competition but overall fair hackathons have been killed by AI. It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners. It used to be about human skill, now it's about ideas and of course insiders are the main winners.

Hackathons were unfair long before AI. See https://news.ycombinator.com/item?id=48468766

The solution is to host and join hackathons without prizes. The point isn't to win, but to create and present something cool and have fun.

If anything, AI's assistance making a fast prototype means hackathons should be better.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#66
I think that a lot of software engineers are using LLMs and a lot of very popular tools are developed by, or are assisted by, LLMs. Is this not just going to be a thing going forward?

This feels akin to traditional artists getting angry at digital art winning competitions when that was a new concept.

We're simply in the early stages of a paradigm shift, no?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#67

Earlier quoted context omitted.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

Pragmatically speaking, a half-assed answer now, is often better than a perfect answer tomorrow.

Yeah, we live finite lives. Time is the one thing the vast majority of us aren’t getting more of. Of course speed is a priority. This isn’t a “capital” thing, it’s a fundamental part of the human experience.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#68

Earlier quoted context omitted.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

Pragmatically speaking, a half-assed answer now, is often better than a perfect answer tomorrow.

Depends what the cost of failure is.

If you're designing powerpoints or entertainment software; perhaps that's true. In the worst case you'll be embarrassed for producing AI slop or lose some revenue.

If your tool has the power to seriously harm or inconvenience people if built wrong, then it's just investor-fuelled myopia.

Post reply on HN