Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

121–130 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#121

> But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. I cannot believe this is just put out there unexamined of any level of "maybe we shouldn't help this happen". This is complete moral abdication. And to be clear, being "good" is no defense. Being good often…

It's nice that the LLM-enabled panopticon still cannot find this very recent related media, [0] but my silly mind can. It is actually an interesting commentary from a non-tech point of view. This is how the rest of the world feels:

Anyway, back to work trying to make my millions using Opus and such.

[0] https://old.reddit.com/r/funny/comments/1pj5bg9/al_companies...

Re: Auto-grading decade-old Hacker News discussions with hindsight

#122

I am surprised the author thought the project passed quality control. The LLM reviews seem mostly false. Looking at the comment reviews on the actual website, the LLM seems to have mostly judged whether it agreed with the takes, not whether they came true, and it seems to have an incredibly poor grasp of it's actual task of accessing whether the comments were predictive or not. The LLM's comment reviews are of often…

I haven’t looked at the output yet, but came here to say,LLM grading is crap. They miss things, they ignore instructions, bring in their own views, have no calibration and in general are extremely poorly suited to this task. “Good” LLM as a judge type products (and none are great) use LLMs to make binary decisions - “do these atomic facts match yes / no” type stuff - and aggregate them to get a score.

I understand this is just a fun exercise so it’s basically what LLMs are good at - generating plausible sounding stuff without regard for correctness. I would not extrapolate this to their utility on real evaluation tasks.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#123
post #88

Earlier quoted context omitted.

Submissions put in the second-chance pool briefly appear (sometimes "again") on the frontpage, and the conversation timestamps are reset so it appears like they were written after the second-chance submission, not before.

I never noticed that. What a weird lie! I suppose they want to make the comments seem "fresh" but it's a deliberate misrepresentation. You could probably even contrive a situation where it could be damaging, e.g. somebody says something before some relevant incident, but the website claims they said it afterwards.

I think the reason is much simpler than that. Resetting the timestamp lets them easily resurface things on the frontpage, because the current time - posting time delta becomes a lot smaller, so it's again ranked higher. And avoiding adding a special case, lets the rest of the codebase work exactly like it was before, basically just need to add a "set submission time to now" function and you get the rest for free.

But, I'm just guessing here based on my own refactoring experience through the years, may be a completely different reason, or even by mistake? Who knows? :)

Re: Auto-grading decade-old Hacker News discussions with hindsight

#125
post #84

Earlier quoted context omitted.

"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.

I'm giving them higher marks than the people who say it won't. LLMs have seen huge improvements over the last 3 years. Are you going to make the bet that they will continue to make similarly huge improvements, taking them well past human ability, or do you think they'll plateau? The former is the boring, linear prediction.

>The former is the boring, linear prediction.

right, because if there is one thing that history shows us again and again is that things that have a period of huge improvements never plateau but instead continue improving to infinity.

Improvement to infinity, that is the sober and wise bet!

Re: Auto-grading decade-old Hacker News discussions with hindsight

#126

Earlier quoted context omitted.

At every intellectual task. They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person). Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

To be clear, you are suggesting “huge improvements” in “every intellectual task”? This is unlikely for the trivial reason that some tasks are roughly saturated. Modest improvements in chess playing ability are likely. Huge improvements probably not. Even more so for arithmetic. We pretty much have that handled. But the more substantive issue is that intellectual tasks are not all interconnected. Getting significantly…

There’s plenty of room to grow for LLMs in terms of chess playing ability considering chess engines have them beat by around 1500 ELO

Re: Auto-grading decade-old Hacker News discussions with hindsight

#127

It's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 Introducing…

I'd love to see sentiment analysis done based on time of day. I'm sure it's largely time zone differences, but I see a large variance in the types of opinions posted to hn in the morning versus the evening and I'd be curious to see it quantified.

Yeah, I see this constantly any time Europe is mentioned in a submission. Early European morning/day, regular discussions, but as the European afternoon/evening comes around, you start noticing a lot anti-union sentiment, discussions start to shift into over-regulation, and the typical boring anti-Europe/EU talking points.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#128
post #83
post #13

#272, I got a B+! Neat. It would be very interesting to see this applied year after year to see if people get better or worse over time in the accuracy of their judgments. It would also be interesting to correlate accuracy to scores, but I kind of doubt that can be done. Between just expressing popular sentiment and the first to the post people getting more votes for the same comment than people who come later it pro…

#250, but then I wasn't trying to make predictions for a future AI. Or anyone else, really. Got a high score mostly for status quo bias, e.g. visual languages going nowhere and FPGAs remain niche.

Yeah, it be much more interesting to see the people who made (at the time) outrageous claims, but they came to be true, rather than a list of people who could state that the status quo most likely would stay as it is.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#130
post #129

Gotta auto grade every HN comment for how good it is at predicting stock market movement then check what the "most frequently correct" user is saying about the next 6 months.

As the saying goes, "past performance is not indicative of future results"
Post reply on HN