Auto-grading decade-old Hacker News discussions with hindsight
271–280 of 285 posts
Re: Auto-grading decade-old Hacker News discussions with hindsight
#272Earlier quoted context omitted.
Echo chamber of rational, thoughtful and truthful speakers is what I’m looking for in Internet forums.
flat earth creationists would describe their colleagues the same way. a group of them certainly is an echo chamber; why isn't your view?
Actually they mostly don't. Lots of infighting over the real true answer .. (infinite flat earth, finite but with impassable ice walls, ..)
Re: Auto-grading decade-old Hacker News discussions with hindsight
#273It would be great to run this on a collection of interesting threads over different periods and not just one snapshot. For example, the thread from the day Trump got elected in 2016, the thread from the day of brexit and so on. Those are the times when people make many passionate predictions about how the future will play out, be good to see them retroactively scored.
Re: Auto-grading decade-old Hacker News discussions with hindsight
#274Earlier quoted context omitted.
“Regular” to who? Pro EU sentiment almost only comes from the EU, which is what you’re observing. Pro-US sentiment is relatively mixed (as is anti-US sentiment) in distribution.
> Pro EU sentiment almost only comes from the EU Says who? But also, it doesn’t suggest what you imply. I could as easily conclude: “Oh wow, the people who actually experience the system like it that much? Awesome!”
Re: Auto-grading decade-old Hacker News discussions with hindsight
#275Earlier quoted context omitted.
The AOL scandal pretty much proved that anonymity is a mirage. You may think you are anonymous but it just takes combining a few unrelated databases to de-anonymize you. HN users think they are anonymous but they're not, they drop factoids all over the place about who they are. 33 bits... it is one of my recurring favorite themes and anybody in the business of managing other people's data should be well aware of the…
I think you're being too conspiracy theorist here by making everything black and white. Besides, the main problem of how difficult it is to deanonymize, not if possible . Privacy and security both have to perfect defense. For example, there's no passwords that are unhackable. There are only passwords that cannot be hacked with our current technology, budgets, and lifetime. But you could brute force my HN password, it…
This is the main driver behind the targeted scams that ordinary people now have to deal with. It is why people get voice calls from loved ones in distress, why they get 'tech support' calls that aim to take over their devices and why lots of people have lost lots of money.
If you think I am too conspiracy theorist by making everything black and white that is maybe simply because we live different lives and have different experience.
Re: Auto-grading decade-old Hacker News discussions with hindsight
#276This is a perfect example of the power and problems with LLMs. I took the narcissistic approach of searching for myself. Here's a grade of one of my comments[1]: >slg: B- (accurate characterization of PH’s “networking & facade” feel, but implicitly underestimates how long that model can persist) And here's the actual comment I made[2]: >And maybe it is the cynical contrarian in me, but I think the "real world" aspect…
I'm not so sure; that may not have been what you meant, but that doesn't mean it's not what others read into it. The broader context is HN is a startup forum and one of the most common discussion patterns is 'I don't like it' that is often a stand-in for 'I don't think it's viable as-is'. Startups are default dead, after all. With that context, if someone were to read your comment and be asked 'does this person think…
The LLM isn't misinterpreting the text, it's just representing people who misinterpreted the text isn't the defense you seem to think it is.
Re: Auto-grading decade-old Hacker News discussions with hindsight
#277I've spent a weekend making something similar for my gmail account (which google keeps nagging me about being 90% full). It's fascinating to be able to classify 65k+ of emails (surprise: more than half are garbage), as well as summarize and trace the nature of communication between specific senders/recipients. It took about 50 hours on a dual RTX 3090 running Qwen 3. My original goal was to prune the account deleting…
so then what do you do with the useful stuff?
Re: Auto-grading decade-old Hacker News discussions with hindsight
#278Earlier quoted context omitted.
Dang, posting links to searches for your own comments is so meta, no matter the topic, but even more meta when about meta crack. I love how the first hit of meta crack is this, your own message about meta crack.
I'm higher than my supplier!
Re: Auto-grading decade-old Hacker News discussions with hindsight
#279Earlier quoted context omitted.
Tiger: humans will never beat tigers because tigers are purpose built killing machines and they are just generalist --40,000BC
You don't think humans hunted tigers in 40,000BC?
Tigers are terrifying, though. I think it takes extreme or perverse circumstances to make hunting a tiger make any sense at all. And even then, traps and poisons make more sense than stalking a tiger to kill it!
Re: Auto-grading decade-old Hacker News discussions with hindsight
#280Earlier quoted context omitted.
I'm not so sure; that may not have been what you meant, but that doesn't mean it's not what others read into it. The broader context is HN is a startup forum and one of the most common discussion patterns is 'I don't like it' that is often a stand-in for 'I don't think it's viable as-is'. Startups are default dead, after all. With that context, if someone were to read your comment and be asked 'does this person think…
And this is a perfect example of how some people respond to LLMs, bending over backwards to justify the output like we are some kids around a Ouija board. The LLM isn't misinterpreting the text, it's just representing people who misinterpreted the text isn't the defense you seem to think it is.
I scoped my comment specifically around what a reasonable human answer would be if one were asked the particular question it was asked with the available information it had. That's all.
Btw I agree with your comment that it hallucinated/assumed your intent! Sorry I did not specify that. This was a bit of a 'play stupid games win stupid prizes' prompt by the OP. If one asks an imprecise question one should not expect a precise answer. The negative externality here is reader's takeaways are based on false precision. So is it the fault of the question asker, the readers, the tool, or some mix? The tool is the easiest to change, so probably deserves the most blame.
I think we'd both agree LLMs are notoriously overly-helpful and provide low confidence responses to things they should just not comment on. That to me is the underlying issue - at the very least they should respond like humans do not only in content but in confidence. It should have said it wasn't confident about its response to your post, and OP should have thus thrown its response out.
Rarely do we have perfect info, in regular communications we're always making assumptions which affect our confidence in our answers. The question is what's the confidence threshold we should use? This is the question to ask before the question of 'is it actually right?', which is also an important question to ask, but one I think they're a lot better at than the former.
Fwiw you can tell most LLMs to update its memory to always give you a confidence score 0.0-1.0. This helps tremendously, it's pretty darn accurate, it's something you can program thresholds around, and I think it should be built in to every LLM response.
The way I see it, LLMs have lots and lots of negative externalities that we shouldn't bring into this world (I'm particularly sensitive to the effects on creative industries), and I detest how they're being used so haphazardly, but they do have some uses we also shouldn't discount and figure out how to improve on. The question is where are we today in that process?
The framework I use to think about how LLMs are evolving is that of transitioning mediums. Like movies started as a copy/paste of stage plays before they settled into the medium and understand how to work along the grain of its strengths & weaknesses to create new conventions. Speech & text are now transitioning into LLMs. What is the grain we need to go along?
My best answer is the convention LLMs need to settle into is explicit confidence, and each question asked of them should first be a question of what the acceptable confidence threshold is for such a question. I think every question and domain will have different answers for that, and we should debate and discuss that alongside any particular answer.