Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

81–90 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#81

> But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. I cannot believe this is just put out there unexamined of any level of "maybe we shouldn't help this happen". This is complete moral abdication. And to be clear, being "good" is no defense. Being good often…

To be clear...prior to this recent explosive interest in LLMs, this was already true. Snowden was over 10 years ago. We can't start clutching our pearls now as if programmatic mass surveillance hasn't been running on all cylinders for over 20 years. Don't get me wrong, we should absolutely care about this, everyone should. I'm just saying any vague gestures at imminent privacy-doom thanks to LLMs is liable to be doin…

Who, exactly, is the "we" who you see "pearl clutching" instead of "yes and-ing"?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#83
post #13

#272, I got a B+! Neat. It would be very interesting to see this applied year after year to see if people get better or worse over time in the accuracy of their judgments. It would also be interesting to correlate accuracy to scores, but I kind of doubt that can be done. Between just expressing popular sentiment and the first to the post people getting more votes for the same comment than people who come later it pro…

#250, but then I wasn't trying to make predictions for a future AI. Or anyone else, really. Got a high score mostly for status quo bias, e.g. visual languages going nowhere and FPGAs remain niche.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#84

One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…

"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#86
Looking at the results and the prompt, I would tweak the prompt to

* ignore comments that do not speculate on something that was unknown or had not achieved consensus as of the date of yyyy-mm-dd

* at the same time, exclude speculations for which there still isn’t a definitive answer or consensus today

* ignore comments that speculate on minor details or are stating a preference/opinion on a subjective matter

* it is ok to generate an empty list of users for a thread if there are no comments meeting the speculation requirements laid out above

* etc

Re: Auto-grading decade-old Hacker News discussions with hindsight

#87

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

The problem seems underspecified; what does it mean for a comment to be accurate? It would seem that comments like "the sun will rise tomorrow" would rank highest, but they aren't surprising.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#88
post #31

Earlier quoted context omitted.

> for some reason, whole conversations get reset to a single timestamp. What do you mean?

Submissions put in the second-chance pool briefly appear (sometimes "again") on the frontpage, and the conversation timestamps are reset so it appears like they were written after the second-chance submission, not before.

I never noticed that. What a weird lie!

I suppose they want to make the comments seem "fresh" but it's a deliberate misrepresentation. You could probably even contrive a situation where it could be damaging, e.g. somebody says something before some relevant incident, but the website claims they said it afterwards.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#89
post #84

One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…

"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.

Perhaps a new category, 'highest risk guess but right the most often'. Those is the high impact predictions.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#90
post #84

Earlier quoted context omitted.

"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.

Perhaps a new category, 'highest risk guess but right the most often'. Those is the high impact predictions.

Prediction markets have pretty much obviated the need for these things. Rather than rely on "was that really a hot take?" you have a market system that rewards those with accurate hot takes. The massive fees and lock-up period discourage low-return bets.
Post reply on HN