Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

211–220 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#211
post #70

Earlier quoted context omitted.

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

it depends on if you vote based on the quality of contribution to the discussion or based on how much you agree/disagree.

I don't think you can change user behavior like this.

You can give them a "venting sink" though. Instead of having a downvote button that just downvotes, have it pop up a little menu asking for a downvote reason, with "spam" and "disagree" as options. You could then weigh downvotes by which option was selected, along with an algorithm to discover "user honesty" based on whether their downvotes correlate with others or just with the people on their end of the political spectrum, a la Birdwatch.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#212

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

I long had a similar idea for stocks. Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy. I would be very surprised if this had not been done a thousand times by various trading firms and enterprising individuals. Of course in the above example of stocks there are clear predictions (HNWS will go up) and an oracle who resolves it (stock market). This seems to be a way harder problem fo…

> Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy.

Didn't somebody make an ETF once that went against the prediction of some famous CNBC stock picker, showing that it would have given you alpha in the past.

> seems to be a way harder problem for generic free form comments.

That's what prediction markets are for. People for whom truth and accuracy matters (often concentrated around the rationalist community) will often very explicitly make annual lists of concrete and quantifiable predictions, and then self-grade on them later.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#213

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

I like the idea and certainly would try it. Although I feel in a way this would be an anti-thesis to HN. HN tries to foster curiosity, but if you're (only) ranked by the accuracy of your predictions, this could give the incentive to always fall back to a save and boring position.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#214
post #70

Earlier quoted context omitted.

The RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose…

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#216

It's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 Introducing…

Comment dates on hn frontend are sometimes altered when submissions are merged, do you handle this case properly?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#217

It's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 Introducing…

I like the "past" functionality here, maybe wished there was one for week/month I could scroll back as well.

Miss it for reddit as well. Top day/week/month/alltime makes it hard to find top a month in 2018.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#218
post #187

Earlier quoted context omitted.

Claiming that AI in anything resembling its current form is older than 5 years is like claiming the history of the combustion engine started when an ape picked up a burning stick.

Your analogy fails because picking up a burning stick isn’t a combustion engine, whereas decades of neural-net and sequence-model work directly enabled modern LLMs. LLMs aren’t “five years old”; the scaling-transformer regime is. The components are old, the emergent-capability configuration is new. Treating the age of the lineage as evidence of future growth is equivocation across paradigms. Technologies plateau when…

I think this is the first time I have ever posted one of these but thank you for making the argument so well.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#219

> https://karpathy.ai/hncapsule/2015-12-24/index.html#article-... I wonder why ChatGPT refused to analyze it? The HN article was "Brazil declares emergency after 2,400 babies are born with brain damage" but the page says "No analysis available".

My guess is that it’s because there’s a lot of very negative comments about Brazil in that article. Trying to grade people for their opinions on a topic like that gets into dangerous territory.
Post reply on HN