Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

101–110 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#101
post #97

Earlier quoted context omitted.

I'm giving them higher marks than the people who say it won't. LLMs have seen huge improvements over the last 3 years. Are you going to make the bet that they will continue to make similarly huge improvements, taking them well past human ability, or do you think they'll plateau? The former is the boring, linear prediction.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

At every intellectual task.

They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person).

Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#102

Earlier quoted context omitted.

> because the website is a good "web citizen." It has urls that maintain their state over a decade. It's a shame that maintaining the web is so hard that only a few websites are "good citizens". I wish the web was a -bit- way more like git. It should be easier to crawl the web and serve it. Say, you browse and get things cached and shared, but only your "local bookmarks" persist. I guess it's like pinning in IPFS.

Yes, I wish we could serve static content more like bittorent, where your uri has an associate hash, and any intermediate router or cache could be an equivalent source of truth, with the final server only needing to play a role if nothing else has it. It is not possible right now to make hosting democratized/distributed/robust because there's no way for people to donate their own resources in a seamless way to keepin…

This is IPFS

Re: Auto-grading decade-old Hacker News discussions with hindsight

#103
post #97

Earlier quoted context omitted.

I'm giving them higher marks than the people who say it won't. LLMs have seen huge improvements over the last 3 years. Are you going to make the bet that they will continue to make similarly huge improvements, taking them well past human ability, or do you think they'll plateau? The former is the boring, linear prediction.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

“At what?” is really the key question here.

A lot of the press likes to paint “AI” as a uniform field that continues to improve together. But really it’s a bunch of related subfields. Once in a blue moon a technique from one subfield crosses over into another.

“AI” can play chess at superhuman skill. “AI” can also drive a car. That doesn’t mean Waymo gets safer when we increase Stockfish’s elo by 10 points.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#104
post #97

Earlier quoted context omitted.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

At every intellectual task. They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person). Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

What is an intellectual task? Once again, there's tons of stuff LLMs won't be trained on in the next 3 years. So it would be trivial to just find one of those things and say voila! LLMs aren't better than me at that.

I'll make one prediction that I think will hold up. No LLM-based system will be able to take a generic ask like "hack the nytimes website and retrieve emails and password hashes of all user accounts" and do better than the best hackers and penetration testers in the world, despite having plenty of training data to go off of. It requires out-of-band thinking that they just don't possess.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#105

Nice! Something must be in the air – last week I built a very similar project using the historical archive of all-in podcast episodes: https://allin-predictions.pages.dev/

I'll use this as evidence supporting my continued demand for a Friedberg only spinoff.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#106

I am surprised the author thought the project passed quality control. The LLM reviews seem mostly false. Looking at the comment reviews on the actual website, the LLM seems to have mostly judged whether it agreed with the takes, not whether they came true, and it seems to have an incredibly poor grasp of it's actual task of accessing whether the comments were predictive or not. The LLM's comment reviews are of often…

Are you sure? The third section of each review lists the “Most prescient” and “Most wrong” comments. That sounds exactly like what you're looking for. For example, on the "Kickstarter is Debt" article, here is the LLM's analysis of the most prescient comment. The analysis seems accurate and helpful to me.

https://karpathy.ai/hncapsule/2015-12-03/index.html#article-...

  phire

  > “Oculus might end up being the most successful product/company to be kickstarted… > Product wise, Pebble is the most successful so far… Right now they are up to major version 4 of their product. Long term, I don't think they will be more successful than Oculus.”

  With hindsight:

  Oculus became the backbone of Meta’s VR push, spawning the Rift/Quest series and a multi‑billion‑dollar strategic bet.
  Pebble, despite early success, was shut down and absorbed by Fitbit barely a year after this thread.

  That’s an excellent call on the relative trajectories of the two flagship Kickstarter hardware companies.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#107

I am surprised the author thought the project passed quality control. The LLM reviews seem mostly false. Looking at the comment reviews on the actual website, the LLM seems to have mostly judged whether it agreed with the takes, not whether they came true, and it seems to have an incredibly poor grasp of it's actual task of accessing whether the comments were predictive or not. The LLM's comment reviews are of often…

Examples: tptacek gets an 'A' for his comment on DF which the LLM claiming that the user "captured DF's unforgiving nature, where 'can't do x or it crashes is just another feature to learn' which remained true until it was fixed on ..."

Link to LLM review: https://karpathy.ai/hncapsule/2015-12-02/index.html#article-....

So the LLM is praising a comment as describing DF as unforgiving (a characterization of the present then, not a statement about the future). And worse, it seems like tptacek may in fact be implying the opposite of the future (e.g., x will continue to crash when it was eventually fixed.)

Here is the original comment: " tptacek on Dec 2, 2015 | root | parent | next [–]

If you're not the kind of person who can take flaws like crashes or game-stopping frame-rate issues and work them into your gameplay, DF is not the game for you. It isn't a friendly game. It can take hours just to figure out how to do core game tasks. "Don't do this thing that crashes the game" is just another task to learn."

Note: I am paraphrasing the LLM review, as the website is also poorly designed, with one unable to select the text of the LLM review!

N.b., this choice of comment review is not overly cherry picked. I just scanned the "best commentators" and tptacek was number two, with this particular egregiously unrelated-to-prediction LLM summary given as justifying his #2 rating.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#108

One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…

This suggests that the best way to grade predictions is some sort of weighting of how unlikely they were at the time. Like, if you were to open a prediction market for statement X, some sort of grade of the delta between your confidence of the event and the “expected” value, summed over all your predictions.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#109

It's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 Introducing…

I'd love to see sentiment analysis done based on time of day. I'm sure it's largely time zone differences, but I see a large variance in the types of opinions posted to hn in the morning versus the evening and I'd be curious to see it quantified.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#110
post #14

I have never felt less confident in the future than I do in 2025... and it's such a stark contrast. I guess if you split things down the middle, AI probably continues to change the world in dramatic ways but not in the all or nothing way people expect. A non trivial amount of people get laid off, likely due to a finanical crisis which is used as an excuse for companies scale up use of AI. Good chance the financial cr…

We do not currently have the political apparatus in place to stop the dystopian nightmares depicted in movies and media. They were supposed to be cautionary tales. Maybe they still can be, but there are basically zero guardrails in non-progressive forms of government to prevent massive accumulations of power being wielded in ways most of the population disapproves of.

Thats the whole point of democracy, to prevent the ruling parties from doing wildly unpopular things. Unlike a dictatorship, where they can do anything (including good things, that otherwise wouldn't happen in a democracy).

I know that "X is destroying democracy, vote for Y" has been a prevalent narrative lately, but is there any evidence that it's true? I get that it's death by a thousand cuts, or "one step at a time" as they say.

Post reply on HN