Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

271–280 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#271
It would be great to run this on a collection of interesting threads over different periods and not just one snapshot. For example, the thread from the day Trump got elected in 2016, the thread from the day of brexit and so on. Those are the times when people make many passionate predictions about how the future will play out, be good to see them retroactively scored.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#272
post #214

Earlier quoted context omitted.

Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.

flat earth creationists would describe their colleagues the same way. a group of them certainly is an echo chamber; why isn't your view?

"flat earth creationists would describe their colleagues the same way."

Actually they mostly don't. Lots of infighting over the real true answer .. (infinite flat earth, finite but with impassable ice walls, ..)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#273

It would be great to run this on a collection of interesting threads over different periods and not just one snapshot. For example, the thread from the day Trump got elected in 2016, the thread from the day of brexit and so on. Those are the times when people make many passionate predictions about how the future will play out, be good to see them retroactively scored.

Assuming this keeps running, I suppose we just have to wait about a year.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#274

Earlier quoted context omitted.

“Regular” to who? Pro EU sentiment almost only comes from the EU, which is what you’re observing. Pro-US sentiment is relatively mixed (as is anti-US sentiment) in distribution.

> Pro EU sentiment almost only comes from the EU Says who? But also, it doesn’t suggest what you imply. I could as easily conclude: “Oh wow, the people who actually experience the system like it that much? Awesome!”

Or one could conclude that the bots were posting at a time of day intending you as the reading target. As long as they post things that you are inclined to agree with, you'll feel positive reinforcement about an issue regardless of the actual popularity or even viability.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#275

Earlier quoted context omitted.

The AOL scandal pretty much proved that anonymity is a mirage. You may think you are anonymous but it just takes combining a few unrelated databases to de-anonymize you. HN users think they are anonymous but they're not, they drop factoids all over the place about who they are. 33 bits... it is one of my recurring favorite themes and anybody in the business of managing other people's data should be well aware of the…

I think you're being too conspiracy theorist here by making everything black and white. Besides, the main problem of how difficult it is to deanonymize, not if possible . Privacy and security both have to perfect defense. For example, there's no passwords that are unhackable. There are only passwords that cannot be hacked with our current technology, budgets, and lifetime. But you could brute force my HN password, it…

I have enough industry insights to prove that your data is floating out there, unprotected, in plain text and that those that are not bound by the law are making very good use of it. Every breach leaks more bits about you.

This is the main driver behind the targeted scams that ordinary people now have to deal with. It is why people get voice calls from loved ones in distress, why they get 'tech support' calls that aim to take over their devices and why lots of people have lost lots of money.

If you think I am too conspiracy theorist by making everything black and white that is maybe simply because we live different lives and have different experience.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#276
post #263
post #50

This is a perfect example of the power and problems with LLMs. I took the narcissistic approach of searching for myself. Here's a grade of one of my comments[1]: >slg: B- (accurate characterization of PH’s “networking & facade” feel, but implicitly underestimates how long that model can persist) And here's the actual comment I made[2]: >And maybe it is the cynical contrarian in me, but I think the "real world" aspect…

I'm not so sure; that may not have been what you meant, but that doesn't mean it's not what others read into it. The broader context is HN is a startup forum and one of the most common discussion patterns is 'I don't like it' that is often a stand-in for 'I don't think it's viable as-is'. Startups are default dead, after all. With that context, if someone were to read your comment and be asked 'does this person think…

And this is a perfect example of how some people respond to LLMs, bending over backwards to justify the output like we are some kids around a Ouija board.

The LLM isn't misinterpreting the text, it's just representing people who misinterpreted the text isn't the defense you seem to think it is.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#277

I've spent a weekend making something similar for my gmail account (which google keeps nagging me about being 90% full). It's fascinating to be able to classify 65k+ of emails (surprise: more than half are garbage), as well as summarize and trace the nature of communication between specific senders/recipients. It took about 50 hours on a dual RTX 3090 running Qwen 3. My original goal was to prune the account deleting…

so then what do you do with the useful stuff?

Local archive + client for search

Re: Auto-grading decade-old Hacker News discussions with hindsight

#278
post #257

Earlier quoted context omitted.

Dang, posting links to searches for your own comments is so meta, no matter the topic, but even more meta when about meta crack. I love how the first hit of meta crack is this, your own message about meta crack.

I'm higher than my supplier!

You never meta crack you wouldn't hit.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#279
post #180

Earlier quoted context omitted.

Tiger: humans will never beat tigers because tigers are purpose built killing machines and they are just generalist --40,000BC

You don't think humans hunted tigers in 40,000BC?

I don't think it would make much sense to hunt large predators prior to the invention of agriculture, even though early humans were probably plenty smart enough to build traps capable of holding animals like tigers. But after that (less than 40k years ago, more than 10k years ago), I'd bet it was a common-ish thing for humans to try to hunt predators that preyed upon their livestock.

Tigers are terrifying, though. I think it takes extreme or perverse circumstances to make hunting a tiger make any sense at all. And even then, traps and poisons make more sense than stalking a tiger to kill it!

Re: Auto-grading decade-old Hacker News discussions with hindsight

#280
post #276
post #263

Earlier quoted context omitted.

I'm not so sure; that may not have been what you meant, but that doesn't mean it's not what others read into it. The broader context is HN is a startup forum and one of the most common discussion patterns is 'I don't like it' that is often a stand-in for 'I don't think it's viable as-is'. Startups are default dead, after all. With that context, if someone were to read your comment and be asked 'does this person think…

And this is a perfect example of how some people respond to LLMs, bending over backwards to justify the output like we are some kids around a Ouija board. The LLM isn't misinterpreting the text, it's just representing people who misinterpreted the text isn't the defense you seem to think it is.

And your response here is a perfect example of confidently jumping to conclusions on what someone's intent is... which is exactly what you're saying the LLM did to you.

I scoped my comment specifically around what a reasonable human answer would be if one were asked the particular question it was asked with the available information it had. That's all.

Btw I agree with your comment that it hallucinated/assumed your intent! Sorry I did not specify that. This was a bit of a 'play stupid games win stupid prizes' prompt by the OP. If one asks an imprecise question one should not expect a precise answer. The negative externality here is reader's takeaways are based on false precision. So is it the fault of the question asker, the readers, the tool, or some mix? The tool is the easiest to change, so probably deserves the most blame.

I think we'd both agree LLMs are notoriously overly-helpful and provide low confidence responses to things they should just not comment on. That to me is the underlying issue - at the very least they should respond like humans do not only in content but in confidence. It should have said it wasn't confident about its response to your post, and OP should have thus thrown its response out.

Rarely do we have perfect info, in regular communications we're always making assumptions which affect our confidence in our answers. The question is what's the confidence threshold we should use? This is the question to ask before the question of 'is it actually right?', which is also an important question to ask, but one I think they're a lot better at than the former.

Fwiw you can tell most LLMs to update its memory to always give you a confidence score 0.0-1.0. This helps tremendously, it's pretty darn accurate, it's something you can program thresholds around, and I think it should be built in to every LLM response.

The way I see it, LLMs have lots and lots of negative externalities that we shouldn't bring into this world (I'm particularly sensitive to the effects on creative industries), and I detest how they're being used so haphazardly, but they do have some uses we also shouldn't discount and figure out how to improve on. The question is where are we today in that process?

The framework I use to think about how LLMs are evolving is that of transitioning mediums. Like movies started as a copy/paste of stage plays before they settled into the medium and understand how to work along the grain of its strengths & weaknesses to create new conventions. Speech & text are now transitioning into LLMs. What is the grain we need to go along?

My best answer is the convention LLMs need to settle into is explicit confidence, and each question asked of them should first be a question of what the acceptable confidence threshold is for such a question. I think every question and domain will have different answers for that, and we should debate and discuss that alongside any particular answer.

Post reply on HN