> But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. I cannot believe this is just put out there unexamined of any level of "maybe we shouldn't help this happen". This is complete moral abdication. And to be clear, being "good" is no defense. Being good often…
To be clear...prior to this recent explosive interest in LLMs, this was already true. Snowden was over 10 years ago. We can't start clutching our pearls now as if programmatic mass surveillance hasn't been running on all cylinders for over 20 years. Don't get me wrong, we should absolutely care about this, everyone should. I'm just saying any vague gestures at imminent privacy-doom thanks to LLMs is liable to be doin…
Auto-grading decade-old Hacker News discussions with hindsight
81–90 of 285 posts
Re: Auto-grading decade-old Hacker News discussions with hindsight
#82Re: Auto-grading decade-old Hacker News discussions with hindsight
#83#272, I got a B+! Neat. It would be very interesting to see this applied year after year to see if people get better or worse over time in the accuracy of their judgments. It would also be interesting to correlate accuracy to scores, but I kind of doubt that can be done. Between just expressing popular sentiment and the first to the post people getting more votes for the same comment than people who come later it pro…
Re: Auto-grading decade-old Hacker News discussions with hindsight
#84One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…
Re: Auto-grading decade-old Hacker News discussions with hindsight
#85Re: Auto-grading decade-old Hacker News discussions with hindsight
#86* ignore comments that do not speculate on something that was unknown or had not achieved consensus as of the date of yyyy-mm-dd
* at the same time, exclude speculations for which there still isn’t a definitive answer or consensus today
* ignore comments that speculate on minor details or are stating a preference/opinion on a subjective matter
* it is ok to generate an empty list of users for a thread if there are no comments meeting the speculation requirements laid out above
* etc
Re: Auto-grading decade-old Hacker News discussions with hindsight
#87This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…
Re: Auto-grading decade-old Hacker News discussions with hindsight
#88Earlier quoted context omitted.
> for some reason, whole conversations get reset to a single timestamp. What do you mean?
Submissions put in the second-chance pool briefly appear (sometimes "again") on the frontpage, and the conversation timestamps are reset so it appears like they were written after the second-chance submission, not before.
I suppose they want to make the comments seem "fresh" but it's a deliberate misrepresentation. You could probably even contrive a situation where it could be damaging, e.g. somebody says something before some relevant incident, but the website claims they said it afterwards.
Re: Auto-grading decade-old Hacker News discussions with hindsight
#89One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…
"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.
Re: Auto-grading decade-old Hacker News discussions with hindsight
#90Earlier quoted context omitted.
"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.
Perhaps a new category, 'highest risk guess but right the most often'. Those is the high impact predictions.