Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

301–310 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#301
post #199

Earlier quoted context omitted.

[flagged]

[flagged]

We appreciate your contributions here, but can you dial back this kind of sentiment in your comments?

> Go build and never speak to a human again, if that's what you want. Reject the humanity you despise

It's a really unfair stereotype to apply to those you're disagreeing with, and I see you do it repeatedly here; characterizing those you disagree with, or the HN community in general, with terms like anti-humanity and anti-human.

Yes I know there's a transhumanist element and an elitist element in Silicon Valley and among AI zealots. But it's not true of everyone or even many people who have big hopes for AI, and it's not a dominant sentiment on HN, or within YC. Most people I see being hopeful about AI are hopeful that it can make things better for humans – better jobs (more pleasant and better paying in real terms), better health/medicine and education that's more accessible to all, etc.

Sure, it's not a given that this will eventuate, and there are plenty of unknowns and plenty of ways in which things could get worse if the wrong decisions are made. Which is exactly why we need to have ongoing, vigorous discussion, which is what we're always trying to cultivate here on HN.

It's fine to express concerns about the real challenges and negative aspects of AI (and tech in general) and the potential for things to get worse. It's good to have that perspective represented.

But, continuing to name or characterize your debating opponents, or the HN/YC/tech community in general as "anti-human" is unfair often-inaccurate. It's a low-substance slur that serves to poison discussions. Enough, please.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#302
post #58

Earlier quoted context omitted.

While I agree with you in principle, I think the parent has a point here: where's the amazing product that couldn't have been done without AI? By now we should have seen some major new invention/company, incredibly fast revolutionary feature rollouts etc but I'm just seeing more of the same.

> " where's the amazing product that couldn't have been done without AI? " My phone can search my photos with text based on what's represented in the photo. That wasn't a thing for computers from 1950 through about 2012, and now it's just normal. Google Lens will tell you what a plant is, Merlin app will identify birds from birdsong. In 2000 we had Dragon Naturally Speaking and Kurzweil VoicePad, now we have voice re…

when your longest tenure is 2.5 years... /s

Well said. This is exactly was I was alluding to. Moving goal posts and making things practical that they didn't even notice (or weren't aware of at the time).

Re: Evidence of inconsistencies in evaluation process and selection of winners

#303

It’s a shame that Arvix (and once thoughtful places like Kaggle) are used for self-promotion. I get people want to work at an AI lab but slopping it in public in this manner is counterproductive to the original intended purpose of these places.

Hasn't this always been the case? Arxiv being used for self promotion and Kaggle being used to pivot into the industry. It is not a recent phenomenon.

[deleted]

Re: Evidence of inconsistencies in evaluation process and selection of winners

#304

Went through the comments here and there and one thing to note is that there was a question about who do you think should have won instead. This is a good question because it is possible that all submissions were like this or there were ones that looked just worse. It would be quite useful to know who came close as well in this case. If you knew which submissions were good you could have a process to revoke the prize…

I think the underlying problem here is that no single human brain has enough glycogen in reserve to thoughtfully process all the AI slop. It simply cannot be done by mortals.

I've noticed this over and over again with "professionals actually prefer LLM responses" studies. Typically the human generated responses seem better to me on a quick sample, but if I had to review 50 of them I'd probably start taking lazy shortcuts; using superficial language aptitude or factual comprehensiveness instead of critically reading.

It does seem like the human judges here might have given credit for e.g. a 20pg arXiv paper without actually reading it. I can blame them professionally but emotionally I have nothing but sympathy. I truly hate LLMs.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#305
post #301
post #199

Earlier quoted context omitted.

[flagged]

We appreciate your contributions here, but can you dial back this kind of sentiment in your comments? > Go build and never speak to a human again, if that's what you want. Reject the humanity you despise It's a really unfair stereotype to apply to those you're disagreeing with, and I see you do it repeatedly here; characterizing those you disagree with, or the HN community in general, with terms like anti-humanity an…

>It's a really unfair stereotype to apply to those you're disagreeing with, and I see you do it repeatedly here; characterizing those you disagree with, or the HN community in general, with terms like anti-humanity and anti-human.

I searched my comments, because that doesn't seem like something I would do. It seems I've only used the term "anti-human" twice since I've been here, two years ago in reference to utopianism (https://news.ycombinator.com/item?id=41561281) and three years ago in reference to LLMs and copyright (https://news.ycombinator.com/item?id=36195596) and I've never used the term "anti-humanity" at all.

>Yes I know there's a transhumanist element and an elitist element in Silicon Valley and among AI zealots. But it's not true of everyone or even many people who have big hopes for AI, and it's not a dominant sentiment on HN, or within YC.

I was answering a specific person about their specific attitude and their specific comments.

Fair game if you want to mod me over that comment, I'll admit I might have gone overboard but It doesn't seem as if you bothered to understand the context of what's actually being said at all here. I'm going to be charitable and assume there wasn't an AI moderator involved, I know you're working on that, if so it might need a bit more debugging.

>Most people I see being hopeful about AI are hopeful that it can make things better for humans – better jobs (more pleasant and better paying in real terms), better health/medicine that's more accessible to all, etc.

Again, I wasn't talking to or about most people.

>But, continuing to name or characterize your debating opponents, or the HN/YC/tech community in general as "anti-human" is unfair, inaccurate and only serves to poison discussions here.

I have. Literally. Never. Used. That. Term. About. A. Person. If you're going to criticize me, do it for things I've actually done. I show my ass all the time here, you don't need to make shit up.

But comments like https://news.ycombinator.com/item?id=48946901 and https://news.ycombinator.com/item?id=48928607 and https://news.ycombinator.com/item?id=48924340 and https://news.ycombinator.com/item?id=48337858 are no better. Reducing music to nothing more than a means to stimulate endorphines (https://news.ycombinator.com/item?id=48722569) and declaring AI better than 99% of human effort (https://news.ycombinator.com/item?id=48302356) and wishing all critics could be branded with a scarlet letter and eliminated from the timeline (https://news.ycombinator.com/item?id=48340190), the sheer glee with which they describe the future in which AI puts everyone out of work and the contempt towards anyone who values human effort does read as at least a bit "anti-human" to me.

There. I actually said the thing. Now you can be honest.

But fine. I shouldn't have added fuel to the fire. I'm clearly experiencing the same frustration echelon is, just from a different direction. HN has been incredibly frustrating of late.

@echelon - that comment of mine that you posted was absolutely tame and reasoned compared to some of the stuff you've been posting lately. And it was correct. And I didn't call anyone names. And you clearly agree with my sentiment where AI is concerned. And HN is never going to be what you (or I) want it to be. And I will apologize for my tone and I will dial it back but I think you need do do the same. You need to recognize that people can and do have legitimate criticisms about AI and that the reason you get flagged isn't because of "witchhunts" but because you go on a tilt reading the slightest bit of anti-AI sentiment. If I deserve to be flagged then you do too.

I clearly need to take a break from this place, touch some grass, something.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#306
post #176

Earlier quoted context omitted.

One of my first gigs as a consultant was to write a project management system for a company that didn't really need a custom project management system. The CEO pulled me aside and told me the only important feature of the project management system was that you couldn't assign the same priority to two features. I would be blamed for making such a crappy project management system, but that's what I was there for. Once…

> the only important feature of the project management system was that you couldn't assign the same priority to two feature That is a good idea for a project management system. Force ranking of priorities.

For a novel problem (sub)domain that is great, if no prior optimization occurred globally dominant bottlenecks / inefficiencies still exist.

But once the worst bottleneck is widened to the same width as the second-worst bottleneck, you are from then on optimizing both until you reach the width of the third-worst bottleneck, and from then on you need to elevate all 3 to improve the situation, and so on.

to make it more concrete with an example:

you can identify that the friction on a bicycle comes predominantly from the front wheel, so you optimize the front wheel bearing/lubrication/... until you discover the front wheel has the same friction as the rear wheel bearings, so if you want to improve you'd have to improve both front and rear wheel friction, which helps until they have improved beyond the friction on the pedal bearings, from then on you need to improve all 3, until you discover the chain links became the friction bottleneck, etc...

Re: Evidence of inconsistencies in evaluation process and selection of winners

#307

"I think you just need to accept the results of the competition. The winning submissions clearly provide value and had a lot of effort invested in them. I'm not really worried about a few inconsistencies or mistakes if the value is still there. Did you think another submission deserved to win over these?" That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.

We've had about a century now of science-fiction literature hyping up AI as a higher intelligence that is based solely on some ill-defined yet universal system of "logic" and is therefore not prone to human flaws such as pride, hate, envy, lust, etc. Now it has become extremely apparent that was always an unsubstantiated assumption but its too late because there are billions of people primed to never question the mac…

FWIW I think the more salient learned behavior is 40 years of using calculators / desktops / laptops / smartphones which were basically 99.9999% reliable at retrieving text and doing computations. It is very hard to undo the learning of "the computer is a machine designed to be accurate."

Re: Evidence of inconsistencies in evaluation process and selection of winners

#308

Earlier quoted context omitted.

To be the devil's advocate, engineers themselves can be hilariously incompetent too. Engineers have a tendency of assuming that budget is infinite and target audience is other engineers from same specialization. Open-source projects often have this problem where you can have dozens of thousands of man-hours poured into a project without a single end-user opinion taken into account. At some point my manager, who himse…

> Open-source projects often have this problem where you can have dozens of thousands of man-hours poured into a project without a single end-user opinion taken into account. How is this "a problem"? The reason there are dozens of thousands of hours on an open source project is because the end-users are working on it. Some projects exist solely for someone to work on it (that is, the "working on it" is the "use case"…

Any open source project that becomes any size 'uses' money, maybe not directly, but at least from corporate handouts like free hosting.

And once a project starts getting a fair number of contributers political problems arise, feelings get hurt, and forks happen. Quite often the forks take the contributers leaving the original project a shell and a warning to others.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#309

Earlier quoted context omitted.

Something like the time value of money. But on the other hand, a bad answer can have negative value. Although "wrong and early" is better than "wrong and late".

If “wrong” breaks things, then late is better than early.

Quite often the only way to know if you're wrong or right is to start building it.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#310

Earlier quoted context omitted.

I just shame people that give slop. Slop PR? Fix the slop. Slop design? I’m not implementing slop, fix it. Innundated with slop PRs? Send half of them to my super and tell him to deal with it. We’ve fired people that wouldn’t get their shit together. Deadlines are being missed because we need to spend more time fixing slop? That’s a planning (management) problem, not mine. Management are the ones that forced everyone…

What if the slop PRs come from your super?

I have a good enough relationship with mine that I feel comfortable telling him his code is garbage - AI slop or Meatsack slop doesn't matter.
Post reply on HN