Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

531–540 of 643 posts

Re: 2025: The Year in LLMs

#531

Earlier quoted context omitted.

Mastery of words is thinking? That's the crazy thing. Yes, in fact, it turns out that language encodes and embodies reasoning. All you have to do is pile up enough of it in a high-dimensional space, use gradient descent to model its original structure, and add some feedback in the form of RL. At that point, reasoning is just a database problem, which we currently attack with attention. No one had the faintest clue. E…

> ELIZA, ROFL. How'd ELIZA do at the IMO last year? What's funny is the failure to grasp any contextual framing of ELIZA. When it came out people were impressed by it's reasoning, it's responses. And in your line of defense it could think because it had mastery of words! But fast forward the current timeline 30 years. You will have been of the same camp that argued on behalf of ELIZA when the rest of the world was as…

No one was impressed with ELIZA's "reasoning" except for a few non-specialist test subjects recruited from the general population. Admittedly it was disturbing to see how strongly some of those people latched onto it.

Meanwhile, you didn't answer my question. How'd ELIZA do on the IMO? If you know a way to achieve gold-medal performance at top-level math and programming competitions without thinking, I for one am all ears.

Re: 2025: The Year in LLMs

#532

Earlier quoted context omitted.

That's why I gave you data! METR study was 16 people using Sonnet 3.5/3.7. Data I'm talking about is 10s of thousands of people and is much more up to date. Some counter examples to METR that are in the literature but I'll just say: "rigor" here is very difficult (including METR) because outcomes are high dimensional and nuanced, or ecological validity is an issue. It's hard to have any approach that someone wouldn't…

To be fair, I’ll take a non-biased 16 person study over “internal measures” from a MAANG company that burned 100s of billions on AI with no ROI that is now forcing its employees to use AI.

What do you think about the METR 50% task length results? About benchmark progress generally?

Re: 2025: The Year in LLMs

#533

Earlier quoted context omitted.

Please do provide some data for this "obvious value of coding agents". Because right now the only thing obvious is the increase in vulnerabilities, people claiming they are 10x more productive but aren't shipping anything, and some AI hype bloggers that fail to provide any quantitative proof.

Sure: at my MAANG company, where I watch the data closely on adoption of CC and other internal coding agent tools, most (significant) LOC are written by agents, and most employees have adopted coding agents as WAU, and the adoption rate is positively correlated with seniority. Like a lot of things LLM related (Simon Willison's pelican test, researchers + product leaders implementing AI features) I also heavily "vibe"…

We probably work at the same company, given you used MAANG instead of FAANG.

As one of the WAU (really DAU) you’re talking about, I want to call out a couple things: 1) the LOC metrics are flawed, and anyone using the agents knows this - eg, ask CC to rewrite the 1 commit you wrote into 5 different commits, now you have 5 100% AI-written commits; 2) total speed up across the entire dev lifecycle is far below 10x, most likely below 2x, but I don’t see any evidence of anyone measuring the counterfactuals to prove speed up anyways, so there’s no clear data; 3) look at token spend for power users, you might be surprised by how many SWE-years they’re spending.

Overall it’s unclear whether LLM-assisted coding is ROI-positive.

Re: 2025: The Year in LLMs

#534

Earlier quoted context omitted.

The problem I see, over and over, is that people pose poorly-formed questions to the free ChatGPT and Google models, laugh at the resulting half-baked answers that are often full of errors and hallucinations, and draw conclusions about the technology as a whole. Either that, or they tried it "last year" or "a while back" and have no concept of how far things have gone in the meantime. It's like they wandered into a m…

No, frankly it's the difference between actual engineers and hobbyists/amateurs/non-SWEs. SWEs are trained to discard surface-level observations and be adversarial. You can't just look at the happy path, how does the system behave for edge cases? Where does it break down and how? What are the failure modes? The actual analogy to a machine shop would be to look at whether the machines were adequate for their use case,…

I wish there was a way to discern posts from legit clever people from the not-so.

Its annoying to see posts from people who lag behind in intelligence and just dont get it - people learn at different rates. Some see way further ahead.

Re: 2025: The Year in LLMs

#535

Earlier quoted context omitted.

I think the confusion is people's misunderstanding of what 'new code' and 'new imagery' mean. Yes, LLMs can generate a specific CRUD webapp that hasn't existed before but only based on interpolating between the history of existing CRUD webapps. I mean traditional Markov Chains can also produce 'new' text in the sense that "this exact text" hasn't been seen before, but nobody would argue that traditional Markov Chains…

> It's pretty clear you don't have a solid background in generative models, because this is fundamentally what they do You don’t have a solid background. No one does. We fundamentally don’t understand LLMs, this is an industry and academic opinion. Sure there are high level perspectives and analogies we can apply to LLMs and machine learning in general like probability distributions, curve fitting or interpolations……

"You don’t have a solid background.

If you want to go around huffing and puffing your chest about a subject area, you kinda do fella. Credibility.

Re: 2025: The Year in LLMs

#536
post #67

Indeed. I don't understand why Hacker News is so dismissive about the coming of LLMs, maybe HN readers are going through 5 stages of grief? But LLM is certainly a game changer, I can see it delivering impact bigger than the internet itself. Both require a lot of investments.

"I can see it delivering impact bigger than the internet itself. Both require a lot of investments."

lol.... Just make sure you screenshot your post so you have a good reminder in a few years re. your predictive ability.

Re: 2025: The Year in LLMs

#537
post #520

Earlier quoted context omitted.

The early internet and smartphones (the Japanese ones, not iPhone) were definitely not "immediately" adopted by the mass, unlike LLM. If "immediate" usefulness is the metric we measure, then the internet and smartphones are pretty insignificant inventions compared to LLM. (of course it's not a meaningful metric, as there is no clear line between a dumb phone and a smart phone, or a moderately sized language model and…

Yeah the internet kind of started with ARPANET in 1969 and didn't really get going with the public till around 1999 so thirty years on. Here's a graph of internet takeoff with Krugman's famous quote of 1998 that it wouldn't amount to much being maybe the end of the skepticism https://www.contextualize.ai/mpereira/paul-krugmans-poor-pre... In common with AI there was probably a long period when the hardware wasn't rea…

Thats all irrelevant. Is/was there tremendous value to be had by being able to transport data? Of course. No doubt about it. Everything else got figured out and investments were made because of that.

The same line of thinking does not hold with LLMs given their non-deterministic nature. Time will tell where things land.

Re: 2025: The Year in LLMs

#538

I can’t get over the range of sentiment on LLMs. HN leans snake oil, X leans “we’re all cooked” —- can it possibly be both? How do other folks make sense of this? I’m not asking for a side, rather understanding the range. Does the range lead you to believe X over Y?

Well, this is the internet. Arguing about everything is its favorite pastime. But generally yes, I think back to Mongo/Node/metaverse/blockchain/IDEs/tablets and pretty much everything has had its boosters and skeptics, this is just more... intense. Anyway I've decided to believe my own eyes. The crowds say a lot of things. You can try most of it yourself and see what it can and can't do. I make a point to compare no…

Only those with great taste are well-equipped to make assertions about what we have infront of us.

The rest is all noise and personally I just block it out.

Re: 2025: The Year in LLMs

#539
post #361

I can’t get over the range of sentiment on LLMs. HN leans snake oil, X leans “we’re all cooked” —- can it possibly be both? How do other folks make sense of this? I’m not asking for a side, rather understanding the range. Does the range lead you to believe X over Y?

Truth lies in the middle. Yes LLM are an incredible piece of technology, and yes we are cooked because once again technologists and VC have no idea nor interest in understanding the long-term societal ramifications of technology. Now we are starting to agree that social media has had disastrous effects that have not fully manifested yet, and in the same breath we accept a piece of technology that promises to replace…

"that have not fully manifested yet"

This is not true..

"I find it utterly maddening how divorced STEM people have become from philosophical and ethical concerns of their work. I blame academia and the education system for creating this massive blind spot, and it is most apparent in echo chambers like HN that are mostly composed of Western-educated programmers with a degree in computer science. At least on X you get, among the lunatics, people that have read more than just books on algorithms and startups."

Steve Jobs had something to say about this. Shame hes gone.

Re: 2025: The Year in LLMs

#540
post #414

Earlier quoted context omitted.

This is not a great argument: > But it is hard to argue against the value of current AI [...] it is getting $1B dollar runway already. The psychic services industry makes over $2 billion a year in the US [1], with about a quarter of the population being actual believers. [2]. [1] The https://www.ibisworld.com/united-states/industry/psychic-ser... [2] https://news.gallup.com/poll/692738/paranormal-phenomena-met...

2022/2023: "It hallucinates, it's a toy, it's useless." 2024/2025: "Okay, it works, but it produces security vulnerabilities and makes junior devs lazy." 2026 (Current): "It is literally the same thing as a psychic scam." Can we at least make predictions for 2027? What shall the cope be then! Lemme go ask my psychic.

I suppose it's appropriate that you hallucinated an argument I did not make, attacked the straw man, and declared victory.
Post reply on HN