Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

241–250 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#241
post #214
post #70

Earlier quoted context omitted.

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.

That’s what everyone living in an echo chamber (and especially one of their own creation) thinks they’re in.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#242

It's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 Introducing…

I'd love to see sentiment analysis done based on time of day. I'm sure it's largely time zone differences, but I see a large variance in the types of opinions posted to hn in the morning versus the evening and I'd be curious to see it quantified.

e.g. how many are cali tech bros vs nyc fintec vs 10am moscow shillbot time

Re: Auto-grading decade-old Hacker News discussions with hindsight

#243
post #214
post #70

Earlier quoted context omitted.

I am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.

Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.

flat earth creationists would describe their colleagues the same way.

a group of them certainly is an echo chamber; why isn't your view?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#244

Earlier quoted context omitted.

I long had a similar idea for stocks. Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy. I would be very surprised if this had not been done a thousand times by various trading firms and enterprising individuals. Of course in the above example of stocks there are clear predictions (HNWS will go up) and an oracle who resolves it (stock market). This seems to be a way harder problem fo…

> Analyze posts of people giving stock tips on WSB, Twitter, etc and rank by accuracy. Didn't somebody make an ETF once that went against the prediction of some famous CNBC stock picker, showing that it would have given you alpha in the past. > seems to be a way harder problem for generic free form comments. That's what prediction markets are for. People for whom truth and accuracy matters (often concentrated around…

Cramer is the stock picker guy. There is a well known "Cramer Effect" or "Cramer Bounce" where the stock peaks then drops hard.

Makes for great pump n dump if you're day trading and willing to ride

https://www.investopedia.com/terms/c/cramerbounce.asp

long-term his choices don't do well, so the Inverse Cramer basically says "do the opposite of this goober" and has solid returns (sorta; depends a lot on methodology, and the sole hedgefund playing that strategy shutdown)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#245

I've spent a weekend making something similar for my gmail account (which google keeps nagging me about being 90% full). It's fascinating to be able to classify 65k+ of emails (surprise: more than half are garbage), as well as summarize and trace the nature of communication between specific senders/recipients. It took about 50 hours on a dual RTX 3090 running Qwen 3. My original goal was to prune the account deleting…

so then what do you do with the useful stuff?

Re: Auto-grading decade-old Hacker News discussions with hindsight

#246

Earlier quoted context omitted.

Never call a man happy until he is dead. Also I don’t think your argument generalizes well - there are plenty of private research investment bubbles that have popped and not reached their original peaks (e.g. VR).

It wasn't a generalized argument, though, it was a specific one, about AI.

Here is one sentence from the referenced prediction:

> I don't think there will be any more AI winters.

This isn't enough to qualify as a testable prediction, in the eyes of people that care about such things, because there is no good way to formulate a resolution criteria for a claim that extends indefinitely into the future. See [1] for a great introduction.

[1]: https://www.astralcodexten.com/p/prediction-market-faq

Re: Auto-grading decade-old Hacker News discussions with hindsight

#247
post #160

One of the few use cases for LLMs that I have high hopes for and feel is still under appreciated is grading qualitative things. LLMs are the first tech (afaik) that can do top-down analysis of phenomena in a manner similar to humans, which means a lot of important human use cases that are judgement-oriented can become more standardized, faster, and more readily available. For instance, one of the unfortunate aspects…

This is wrong, just look at this comment here:

https://news.ycombinator.com/item?id=46222523

LLM can't grade reliably human text. It doesn't understand it.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#248

I've spent a weekend making something similar for my gmail account (which google keeps nagging me about being 90% full). It's fascinating to be able to classify 65k+ of emails (surprise: more than half are garbage), as well as summarize and trace the nature of communication between specific senders/recipients. It took about 50 hours on a dual RTX 3090 running Qwen 3. My original goal was to prune the account deleting…

I would love to do something like this, and weirdly I even have a dual 3090 home setup.

Any chance you can outline the steps/prompts/tools you used to run this?

I've been building a 2nd brain type project, that plugs into all my work places and a custom classifier has been on that list that would enhance that.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#249
post #241
post #214

Earlier quoted context omitted.

Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.

That’s what everyone living in an echo chamber (and especially one of their own creation) thinks they’re in.

I don't think I'm in any is my problem (HN is better than most, doesn't mean it's good in absolute terms...)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#250

Earlier quoted context omitted.

We do not currently have the political apparatus in place to stop the dystopian nightmares depicted in movies and media. They were supposed to be cautionary tales. Maybe they still can be, but there are basically zero guardrails in non-progressive forms of government to prevent massive accumulations of power being wielded in ways most of the population disapproves of.

Thats the whole point of democracy, to prevent the ruling parties from doing wildly unpopular things. Unlike a dictatorship, where they can do anything (including good things, that otherwise wouldn't happen in a democracy). I know that "X is destroying democracy, vote for Y" has been a prevalent narrative lately, but is there any evidence that it's true? I get that it's death by a thousand cuts, or "one step at a tim…

> I know that "X is destroying democracy, vote for Y" has been a prevalent narrative lately, but is there any evidence that it's true? I get that it's death by a thousand cuts, or "one step at a time" as they say.

I suggest reading [1], [2], and [3]. From there, you'll probably have lots of background to pose your own research questions. According to [4], until you write about something, your thinking will be incomplete, and I tend to agree nearly all of the time.

[1]: https://en.wikipedia.org/wiki/Democratic_backsliding

[2]: https://hub.jhu.edu/2024/08/12/anne-applebaum-autocracy-inc/

[3]: https://carnegieendowment.org/research/2025/08/us-democratic...

[4]: "Neuroscientists, psychologists and other experts on thinking have very different ideas about how our brains work, but, as Levy writes: “no matter how internal processes are implemented, (you) need to understand the extent to which the mind is reliant upon external scaffolding.” (2011, 270) If there is one thing the experts agree on, then it is this: You have to externalise your ideas, you have to write. Richard Feynman stresses it as much as Benjamin Franklin. If we write, it is more likely that we understand what we read, remember what we learn and that our thoughts make sense." - Sönke Ahrens. How to Take Smart Notes_ - Sonke Ahrens (p. 30)

Post reply on HN