Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

71–80 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#71

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

I long had a similar idea for stocks. Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy. I would be very surprised if this had not been done a thousand times by various trading firms and enterprising individuals.

Of course in the above example of stocks there are clear predictions (HNWS will go up) and an oracle who resolves it (stock market). This seems to be a way harder problem for generic free form comments. Who resolves what prediction a particular comment has made and whether it actually happened?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#72
A majority don't seem to be predictions about the future, and it seems to mostly like comments that give extended air to what was then and now the consensus viewpoint, e.g. the top comment from pcwalton the highest scored user: https://news.ycombinator.com/item?id=10657401

> (Copying my comment here from Reddit /r/rust:) Just to repeat, because this was somewhat buried in the article: Servo is now a multiprocess browser, using the gaol crate for sandboxing. This adds (a) an extra layer of defense against remote code execution vulnerabilities beyond that which the Rust safety features provide; (b) a safety net in case Servo code is tricked into performing insecure actions. There are still plenty of bugs to shake out, but this is a major milestone in the project.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#73
post #70

Earlier quoted context omitted.

The RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose…

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

it depends on if you vote based on the quality of contribution to the discussion or based on how much you agree/disagree.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#74
One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst.

If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse.

Something like "Lithium-ion battery pack prices fall to $108/kWh" is classic cost-curve progress. Boring, steady, and historically extremely reliable over long horizons. Probably one of the most likely headlines today to age correctly, even if it gets little attention.

On the flip side, stuff like "New benchmark shows top LLMs struggle in real mental health care" feels like high-risk framing. Benchmarks rotate constantly, and “struggle” headlines almost always age badly as models jump whole generations.

I bet theres many "boring but right" takes we overlook today and I wondr if there's a practical way to surface them before hindsight does

Re: Auto-grading decade-old Hacker News discussions with hindsight

#75
I understand the exercise, but I think it should have a disclaimer, some of the LLM reviews are showing a bias and when I read the comments they turned out not to be as bad as the LLM made them. As this hits the front page, some people will only read the title and not the accompanying blog post, losing all of the nuance.

That said, I understand the concept and love what you did here. By this being exposed to the best disinfectant, I hope it will raise awareness and show how people and corporations should be careful about its usage. Now this tech is accessible to anyone, not only big techs, in a couple of hours.

It also shows how we should take with a grain of salt the result of any analysis of such scale by a LLM. Our private channels now and messages on software like Teams and Slack can be analyzed to hell by our AI overlords. I'm probably going to remove a lot of things from cloud drives just in case. Perhaps online discourse will deteriorate to more inane / LinkedIn style content.

Also, I like that your prompt itself has some purposefully leaked bias, which shows other risks—¹for instance, "fsflover: F", which may align the LLM to grade worse the handles that are related to free software and open source).

As a meta concept of this, I wonder how I'll be graded by our AI overlords in the future now that I have posted something dismissive of it.

¹Alt+0151

Re: Auto-grading decade-old Hacker News discussions with hindsight

#76
post #38
post #26

Earlier quoted context omitted.

It's true that meta is the crack of internet forums, so we, er, crack down on it quite a bit. That's a longstanding view: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... Alternate metaphor: evil catnip - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... But yesterday's thread and this one are clearly exceptions—far above the median. https://news.ycombinator.com/item?id=46212180 was parti…

I love it when you share some insight about HN or internet communication for which you have relevant searches at the ready to explanations of the concept. A personal favourite is “the contrarian dynamic”. Do you have a list of those at the ready or do you just remember them? If you feel like sharing, what’s your process and is there a list of those you’d make public? I imagine having one would be useful, e.g. for onb…

I just remember them. Or forget them!

The process is simply that moderation is super repetitive, so eventually certain pathways get engraved in one's memory. A lot of the time, though, I can't quite remember one of these patterns and I'm unable to dig up my past comments about it. That's annoying, in that particular way when your brain can feel something's there but is unable to retrieve it.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#77
I believe that the GPA calculation is off, maybe just for F's.

I scrolled to the bottom of the hall of fame/shame and saw that entry #1505 and 3 F's and a D, with an average grade of D+ (1.46).

No grade better than a D shouldn't average to a D+, I'd expect it to be closer to a 0.25.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#78
post #70

Earlier quoted context omitted.

The RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose…

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

Reddit's current structure very much produces an echo chamber with only one main prevailing view. If everyone used an extension like this I would expect it to increase overall diversity of opinion on the site, as things that conflict with the main echo chamber view could still thrive in their own communities rather than getting downvoted with the actual spam.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#79

One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…

Instead of "LLM's will put developers out of jobs" the boring reality is going to be "LLM's are a useful tool with limited use".

Re: Auto-grading decade-old Hacker News discussions with hindsight

#80
post #70

Earlier quoted context omitted.

The RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose…

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

More than having exact same system but with any random reader voting ? I'd say as long as you don't do "I disagree therefore I downvote" it would probably be more accurate than having essentially same voting system driven by randoms like reddit/HN already does
Post reply on HN