Live data from Hacker News

Auto-grading decade-old Hacker News discussions with hindsight

karpathy.bearblog.dev

161–170 of 285 posts

Re: Auto-grading decade-old Hacker News discussions with hindsight

#161
post #97

Earlier quoted context omitted.

I'm giving them higher marks than the people who say it won't. LLMs have seen huge improvements over the last 3 years. Are you going to make the bet that they will continue to make similarly huge improvements, taking them well past human ability, or do you think they'll plateau? The former is the boring, linear prediction.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

I imagine "better" in this case depends on how one scores "I don't know" or confident-sounding falsehoods.

Failures aren't just a ratio, they're a multi-dimensional shape.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#162

Earlier quoted context omitted.

I'd love to see sentiment analysis done based on time of day. I'm sure it's largely time zone differences, but I see a large variance in the types of opinions posted to hn in the morning versus the evening and I'd be curious to see it quantified.

Yeah, I see this constantly any time Europe is mentioned in a submission. Early European morning/day, regular discussions, but as the European afternoon/evening comes around, you start noticing a lot anti-union sentiment, discussions start to shift into over-regulation, and the typical boring anti-Europe/EU talking points.

“Regular” to who? Pro EU sentiment almost only comes from the EU, which is what you’re observing. Pro-US sentiment is relatively mixed (as is anti-US sentiment) in distribution.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#163

Earlier quoted context omitted.

Yes, I wish we could serve static content more like bittorent, where your uri has an associate hash, and any intermediate router or cache could be an equivalent source of truth, with the final server only needing to play a role if nothing else has it. It is not possible right now to make hosting democratized/distributed/robust because there's no way for people to donate their own resources in a seamless way to keepin…

This is IPFS

[deleted]

Re: Auto-grading decade-old Hacker News discussions with hindsight

#164
post #147

Earlier quoted context omitted.

Are you sure? The third section of each review lists the “Most prescient” and “Most wrong” comments. That sounds exactly like what you're looking for. For example, on the "Kickstarter is Debt" article, here is the LLM's analysis of the most prescient comment. The analysis seems accurate and helpful to me. https://karpathy.ai/hncapsule/2015-12-03/index.html#article-... phire > “Oculus might end up being the most succe…

Until someone publishes a systematic quality assessment, we're grasping at anecdotes. It is unfortunate that the questions of "how well did the LLM do?" and "how does 'grading' work in this app?" seem to have gone out the window when HN readers see something shiny.

Yes. And the article is a perfect example of the dangerous sort of automation bias that people will increasingly slide into when it comes to LLMs. I realize Karpathy is sort of incentivized toward this bias given his career, but he doesn't even spend a single sentence even so much as suggesting that the results would need further inspection, or that they might be inaccurate.

The LLM is consulted like a perfect oracle, flawless in its ability to perform a task, and it's left at that. Its results are presented totally uncritically.

For this project, of course, the stakes are nil. But how long until this unfounded trust in LLMs works its way into high stakes problems? The reign of deterministic machines for the past few centuries has ingrained a trust in the reliability of machines in us that should be suspended when dealing with an inherently stochastic device like an LLM.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#166
post #97

Earlier quoted context omitted.

> Are you going to make the bet that they will continue to make similarly huge improvements Sure yeah why not > taking them well past human ability, At what? They're already better than me at reciting historical facts. You'd need some actual prediction here for me to give you "prescience".

At every intellectual task. They're already better than you at reciting historical facts. I'd guess they're probably better at composing poems (they're not great but far better than the average person). Or you agree with me? I'm not looking for prescience marks, I'm just less convinced that people really make the more boring and obvious predictions.

> They're already better than you at reciting historical facts.

so is a textbook, but no-one argues that's intelligent

Re: Auto-grading decade-old Hacker News discussions with hindsight

#167

Earlier quoted context omitted.

I'm giving them higher marks than the people who say it won't. LLMs have seen huge improvements over the last 3 years. Are you going to make the bet that they will continue to make similarly huge improvements, taking them well past human ability, or do you think they'll plateau? The former is the boring, linear prediction.

LaunchHN: Announcing Twoday, our new YC backed startup coming out of stealth mode. We’re launching a breakthrough platform that leverages frontier scale artificial intelligence to model, predict, and dynamically orchestrate solar luminance cycles, unlocking the world’s first synthetic second sunrise by Q2 2026. By combining physics informed multimodal models with real time atmospheric optimisation, we’re redefining w…

You joke, but, alas, there is a _real_ company kinda trying to do this. Reflect Orbital[1] wants to set up space mirrors, so you can have daytime at night for your solar panels! (Various issues, like around light pollution and the fact that looking up at the proposed satellites with binoculars could cause eye damage... don't seem to be on their roadmap.) This is one idea that's going to age badly whether or not they actually launch anything, I suspect.

Battery tech is too boring, but seems more likely to manage long-term effectiveness.

[1] https://www.reflectorbital.com

Re: Auto-grading decade-old Hacker News discussions with hindsight

#168

Earlier quoted context omitted.

Yes, I wish we could serve static content more like bittorent, where your uri has an associate hash, and any intermediate router or cache could be an equivalent source of truth, with the final server only needing to play a role if nothing else has it. It is not possible right now to make hosting democratized/distributed/robust because there's no way for people to donate their own resources in a seamless way to keepin…

This is IPFS

In my experience from the couple of times I clicked an IPFS link years ago, it loaded for a long time and never actually loaded anything, failing the first "I wish we could serve static content" part.

If you make it possible for people to donate bandwidth you might just discover no one wants to.

Re: Auto-grading decade-old Hacker News discussions with hindsight

#169
post #76
post #38

Earlier quoted context omitted.

I love it when you share some insight about HN or internet communication for which you have relevant searches at the ready to explanations of the concept. A personal favourite is “the contrarian dynamic”. Do you have a list of those at the ready or do you just remember them? If you feel like sharing, what’s your process and is there a list of those you’d make public? I imagine having one would be useful, e.g. for onb…

I just remember them. Or forget them! The process is simply that moderation is super repetitive, so eventually certain pathways get engraved in one's memory. A lot of the time, though, I can't quite remember one of these patterns and I'm unable to dig up my past comments about it. That's annoying, in that particular way when your brain can feel something's there but is unable to retrieve it.

Well, you're #24 in this article's hall of fame, and the LLM thinks your moderation views stood the test of time. Perhaps it can already retrieve them for you.
Post reply on HN