Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

221–230 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#221

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

The problem seems underspecified; what does it mean for a comment to be accurate? It would seem that comments like "the sun will rise tomorrow" would rank highest, but they aren't surprising.

just because an idea is qualitative doesn't mean its invalid

Re: Auto-grading decade-old Hacker News discussions with hindsight

#222
post #195

Earlier quoted context omitted.

You can't anonymize comments from well-known users, to an LLM: https://gwern.net/doc/statistics/stylometry/truesight/index

That's an overly strong claim, an LLM could also be used to normalise style

How would you possibly grade comments if you change them?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#223

Earlier quoted context omitted.

LaunchHN: Announcing Twoday, our new YC backed startup coming out of stealth mode. We’re launching a breakthrough platform that leverages frontier scale artificial intelligence to model, predict, and dynamically orchestrate solar luminance cycles, unlocking the world’s first synthetic second sunrise by Q2 2026. By combining physics informed multimodal models with real time atmospheric optimisation, we’re redefining w…

You joke, but, alas, there is a _real_ company kinda trying to do this. Reflect Orbital[1] wants to set up space mirrors, so you can have daytime at night for your solar panels! (Various issues, like around light pollution and the fact that looking up at the proposed satellites with binoculars could cause eye damage... don't seem to be on their roadmap.) This is one idea that's going to age badly whether or not they…

Reflecting sunlight from orbit is an idea that had been talked about for a couple of decades even before Znamya-2[1] launched in 1992. The materials science needed to unfurl large surfaces in space seems to be very difficult, whether mirrors or sails.

[1] https://en.wikipedia.org/wiki/Znamya_(satellite)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#224

Earlier quoted context omitted.

That's an overly strong claim, an LLM could also be used to normalise style

How would you possibly grade comments if you change them?

Extract the concrete predictions, evaluate them as true/false/indeterminate, and grade the user on the number of true vs false?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#225

Earlier quoted context omitted.

That's an overly strong claim, an LLM could also be used to normalise style

How would you possibly grade comments if you change them?

You don’t need comments, just facts in them to see if they’re accurate.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#227

Earlier quoted context omitted.

Yeah, I see this constantly any time Europe is mentioned in a submission. Early European morning/day, regular discussions, but as the European afternoon/evening comes around, you start noticing a lot anti-union sentiment, discussions start to shift into over-regulation, and the typical boring anti-Europe/EU talking points.

“Regular” to who? Pro EU sentiment almost only comes from the EU, which is what you’re observing. Pro-US sentiment is relatively mixed (as is anti-US sentiment) in distribution.

> Pro EU sentiment almost only comes from the EU

Says who? But also, it doesn’t suggest what you imply. I could as easily conclude: “Oh wow, the people who actually experience the system like it that much? Awesome!”

Re: Auto-grading decade-old Hacker News discussions with hindsight

#228

It doesn't look like the code anonymizes usernames when sending the thread for grading. This likely induces bias in the grades based on past/current prevailing opinions of certain users. It would be interesting to see the whole thing done again but this time randomly re-assigning usernames, to assess bias, and also with procedurally generated pseudonyms, to see whether the bias can be removed that way. I'd expect de-…

What a human-like critizicism of human-like behavior.

I [as a human] also do the same thing when observing others in IRL and forum interactions. Reputation matters™

----

A further question is whether a bespoke username could influence the bias of a particular comment (e.g. A username of something like HatesPython might influence the interpretation of that commenter's particular perception of the Python coding language, which might actually be expressing positivity — the username's irony lost to the AI?).

Re: Auto-grading decade-old Hacker News discussions with hindsight

#230

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

I long had a similar idea for stocks. Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy. I would be very surprised if this had not been done a thousand times by various trading firms and enterprising individuals. Of course in the above example of stocks there are clear predictions (HNWS will go up) and an oracle who resolves it (stock market). This seems to be a way harder problem fo…

Out of curiosity, I built this. I extended karpathy's code and widened the date range to see what stocks these users would pick given their sentiments.

What came back were the usual suspects: GLP-1 companies and AI.

Back to the "boring but right" thesis. Not much alpha to be found

Post reply on HN