Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

111–120 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#111
post #97

Earlier quoted context omitted.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

At every intellectual task. They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person). Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

To be clear, you are suggesting “huge improvements” in “every intellectual task”?

This is unlikely for the trivial reason that some tasks are roughly saturated. Modest improvements in chess playing ability are likely. Huge improvements probably not. Even more so for arithmetic. We pretty much have that handled.

But the more substantive issue is that intellectual tasks are not all interconnected. Getting significantly better at drawing hands doesn’t usually translate to executive planning or information retrieval.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#112

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

The RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose…

That assumes your upvotes in the past were a good proxy for being correct today. You could have both been wrong.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#113

Looking at the results and the prompt, I would tweak the prompt to * ignore comments that do not speculate on something that was unknown or had not achieved consensus as of the date of yyyy-mm-dd * at the same time, exclude speculations for which there still isn’t a definitive answer or consensus today * ignore comments that speculate on minor details or are stating a preference/opinion on a subjective matter * it is…

You would also need to exclude “predictions” for things which already happened at the time they were predicted.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#114
Predictions are only valuable when they're actually made ahead of the knowledge becoming available. A man will walk on mars by 2030 is falsifiable, a man will walk on mars is not. A lot of these entries have very low to no predictive value or were already known at the time, but just related. Would be nice if future 'judges' put in more work to ensure quality judgments.

I would grade this article B-, but then again, nobody wrote it... ;)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#115
post #108

One thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Some…

This suggests that the best way to grade predictions is some sort of weighting of how unlikely they were at the time. Like, if you were to open a prediction market for statement X, some sort of grade of the delta between your confidence of the event and the “expected” value, summed over all your predictions.

Exactly, that's the element that is missing. If there are 50 comments against and one pro and that pro has it in the longer term then that is worth noticing, not when there are 50 comments pro and you were one of the 'pros'.

Going against the grain and turning out right is far more valuable than being right consistently when the crowd is with you already.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#116
post #104

Earlier quoted context omitted.

At every intellectual task. They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person). Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

What is an intellectual task? Once again, there's tons of stuff LLMs won't be trained on in the next 3 years. So it would be trivial to just find one of those things and say voila! LLMs aren't better than me at that. I'll make one prediction that I think will hold up. No LLM-based system will be able to take a generic ask like "hack the nytimes website and retrieve emails and password hashes of all user accounts" and…

I'll take a stab at this: LLMs currently seem to be rather good at details, but they seem to struggle greatly with the overall picture, in every subject.

- If I want Claude Code to write some specific code, it often handles the task admirably, but if I'm not sure what should be written, consulting Claude takes a lot of time and doesn't yield much insight, where as 2 minutes with a human is 100x more valuable.

- I asked ChatGPT about some political event. It mirrored the mainstream press. After I reminded it of some obvious facts that revealed a mainstream bias, it agreed with me that its initial answer was wrong.

These experiences and others serve to remind me that current LLMs are mostly just advanced search engines. They work especially well on code because there is a lot of reasonably good code (and tutorials) out there to train on. LLMs are a lot less effective on intellectual tasks that humans haven't already written and published about.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#117

> I was reminded again of my tweets that said "Be good, future LLMs are watching". You can take that in many directions, but here I want to focus on the idea that future LLMs are watching. Everything we do today might be scrutinized in great detail in the future because doing so will be "free". A lot of the ways people behave currently I think make an implicit "security by obscurity" assumption. But if intelligence r…

Given the quality of the judgment I'm not worried, there is no value here.

To properly execute this idea rather than to just toss it off without putting in the work to make it valuable is exactly what irritates me about a lot of AI work. You can be 900 times as productive at producing mental popcorn, but if there was value to be had here we're not getting it, just a whiff of it. Sure, fun project. But I don't feel particularly judged here. The funniest bit is the judgment on things that clearly could not yet have come to pass (for instance because there is an exact date mentioned that we have not yet reached). QA could be better.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#118
post #94
post #59

Now: compared to what? Is there a better source than HN? How's it compare to Reddit or lobsters? Compared to what happens next? Does tptacek's commentary become market signal equivalent to the Fed Chair or the BLS labor and inflation reports?

What makes you think it already isn't?

You've made me billions by now! Thank you...

Re: Auto-grading decade-old Hacker News discussions with hindsight

#119

This is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of th…

That’s what Elon’s vision was before he ended up buying Twitter. Keep a digital track record for journalists. He wanted to call it Pravda.

(And we do have that in real life. Just as, among friends, we do keep track of who is in whose debt, we also keep a mental map of whose voice we listen to. Old school journalism still had that, where people would be reading someone’s column over the course of decades. On the internet, we don’t have that, or we have it rarely.)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#120

> But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. I cannot believe this is just put out there unexamined of any level of "maybe we shouldn't help this happen". This is complete moral abdication. And to be clear, being "good" is no defense. Being good often…

I've had the same though as Karpathy over the past couple of months/years. I don't think it's good, exciting, or something to celebrate, but I also have no idea how to prevent it. I would read his "Best to be good." as a warning or reminder that everything you do or say online will be collected and analyzed by an "intelligence". You can't count on hiding amongst the mass of online noise. Imagine if someone were to co…

I think we need to stop focusing only on the AI aspect of this. Yes, it's an important component to the sort of mass surveillance system you're describing, but it's not the only component. The internet, advertising, privacy, all of these are integral to this outcome.

While I don't have a general solution, I do believe that the solution will need to be multi-faceted and address multiple aspects of the technologies enabling this. My first step would be for society to re-evaluate and shift its views towards information, both locally and internationally.

For example, if you proposed to get rid of all physical borders between countries, everyone would likely be aghast. Obviously there are too many disagreements and conflicting value sets between countries for this to happen. Yet in the west we think nothing have having no digital information borders, despite the fact that the lack of them in part enables this data collection and other issues such as election interference. Yes, erecting firewalls is extremely unpalatable to people in the west, but is almost certainly part of the solution on the national level. Countries like China long ago realized this, though they also use firewalls as a means of control, not just protection (it doesn't have to be this way).

But within countries we also need to shift away from a default position of "I have the right to say whatever I want so therefore I should" and into one of "I'm not putting anything online unless I'm willing to have my employer, parents, literally everyone, read it." Also, we need to systematically attack and dismantle the advertising industry. That industry is one of the single biggest driving factors behind the extreme systematic collection and correlation of data on people. Advertising needs to switch to a "you come to me" approach not a "I'm coming to you" approach.

Post reply on HN