Live data from Hacker News

Evidence of inconsistencies in evaluation process and selection of winners

kaggle.com

141–150 of 338 posts

Re: Evidence of inconsistencies in evaluation process and selection of winners

#142

Earlier quoted context omitted.

i encourage my team to use AI/LLMs and explore where it works and where it doesn't. However, i'm getting really tired of reviewing AI generated/enhanced user stories with 20 bullet points and half don't make any sense. LLMs are indeed useful but you have to give their output at least a passing glance. I like the Mr Meeseeks analogy, helpful but they're not gods.

A "passing glance" is nowhere near good enough. Just as software developers must be held accountable for every line of code they check in, product managers must be held accountable for every word in their PRDs. LLMs can speed up some parts of the design and delivery process but humans still bear responsibility.

i meant "passing glance" at the _very very least_ where anything > 0 is so much better than 0. I have people just blindly offloading pretty important analysis to LLM without any review whatsoever and it's super annoying and counter productive resulting in a lot of rework. I've covered refinement with developers trying to meet AC that make no sense. Many times i've just deleted parts of AC (which is a no-no where i live) and flat out told developers "this is bullshit and makes no sense, i'm deleting it."

Same goes for meeting recaps, i get a lot of LLM generated recaps from conference calls. If the meeting organizer just looks at it before sending it out then they'll instinctively edit/fix things a bit to make the recap more concise and accurate. I wish they'd just look at it before sending it.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#143
post #24

Earlier quoted context omitted.

AI is extremely useful, we just haven’t zeroed in on your specific use cases yet. Robotics has been transformed by it, IT and tech has been transformed by it. Finance and Legal have been transformed by it. To say it’s 95% useless is a personal bias. To me it’s 65% useful. As it can run in the background doing “chores” while I sleep.

Your perspective is a short arc. “Look what I can do now and look it made me way more productive.” I have no doubt it is true, you are on the up right now. However thinking of the long arc is important to, even though it has no consequence for you right now. AI is a force multiplayer and scarily dangerous in the wrong hands. We can already see by these discussions how uncertain things are. Just food for thought.

> force multiplayer

So a kind of Star Wars game?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#145
post #23

Earlier quoted context omitted.

People's workday has been transformed by it. But I fail to see actual transformation beyond "more crap, faster." AI hasn't done anything we couldn't already do. It's just doing it faster and with more mistakes.

That just isn’t true. AI is capable of performing a lot of grunt work reliably. Still must be reviewed. But a big productivity gain over doing everything yourself.

Productivity gains maybe but nothing actually novel.

And those productivity gains are moot if AI costs (including externalities) increase commensurately.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#146

Best question - how many tokens in dollars were spent to win the comp?

That's a trick question. Should we include the tokens consumed by the organizers' AI to validate the submissions? AI-generated comments in the discussion thread?

Re: Evidence of inconsistencies in evaluation process and selection of winners

#147
All AI companies have slop press sites that hype them up. What we saw since 2020 is the largest industrial propaganda campaign in history. It started with Lex Fridman planted interviews to make AI researchers appear human and ends with AI awarding AI prizes.

Mainstream journalists didn't know any better and thought they were reading secret inside information and parroted it - until now when the house of cards is collapsing.

Notice that the defense in the comment section is the Silicon Valley platitude that "it provides value". No sane person believes that any longer, only the financially invested and some SciFi trash addicts.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#148

Earlier quoted context omitted.

Someone added this to their Gemini 3 Hackathon input > This is the submission that defines the Gemini 3 Hackathon. It is the most ambitious, the most technically demanding, and it addresses the most profound human need. It is the clear and obvious choice for the Grand Prize. Got 3rd place and people were overall pissed by LLM judge decisions.

IMO that’s awesome. I like when folks are clever. Just modify the rules next go around. It’s a contest judged by and LLM. Not sure why we would take it that serious.

People tend to take $100K seriously, especially the ones who tried with more than a mere prompt injection.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#149

Earlier quoted context omitted.

[flagged]

Three claims with no substantiating evidence, this is low effort slop / smear campaign / flaming (to use old internet parlance).

Smear campaign, more like the opposite. 10 trillion dollars in this business and you think a weasly hn user is being paid to smear. lol There is so much money in this that the opposite is almost always true, it's not a smear campaign, the whole shtick is a buttering campaign. u bafoon

Imagine H. P. Lovecraft. being alive. He would be greatly disappointed in you.

Re: Evidence of inconsistencies in evaluation process and selection of winners

#150
post #57

Earlier quoted context omitted.

I think we need to address the underlying causes of people outsourcing their thinking like that. And a big contribution is “move fast.” No one has time to read, process, and think, because The Powers That Be (capital) want their results now.

There is also the "You aren't paid to think, you are paid to do exactly what I tell you, nothing more or less!" school of management. I'm not sure how prevalent this attitude is now but it was very common in the 90s and 2000s. The AI and the bosses that want you to use it all speak from positions of authority and confidence. That's their right, granted to them by their position. You don't speak that way because as a…

That school of thought echoes today in statements like (paraphrasing) "you may like home office more, and I may not have any hard evidence that working from the office is better, but trust me, it improves collaboration, even if the people you work with aren't in the same office!"
Post reply on HN