Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

581–590 of 643 posts

Re: 2025: The Year in LLMs

#581

Earlier quoted context omitted.

> exponential progress First you need to define what it means. What's the metric? Otherwise it's very much something you can argue about.

> What's the metric? Language model capability at generating text output. The model progress this year has been a lot of: - “We added multimodal” - “We added a lot of non AI tooling” (ie agents) - “We put more compute into inference” (ie thinking mode) So yes, there is still rapid progress, but these ^ make it clear, at least to me, that next gen models are significantly harder to build. Simultaneously we see a disti…

Next gen models are always hard to build, they are by definition pushing the frontier. Every generation of CPU was hard to build but we still had Moores law.

> Simultaneously we see a distinct narrowing between players (openai, deepseek, mistral, google, anthropic) in their offerings. Thats usually a signal that the rate of progress is slowing.

I agree with you on the fact in the first part but not the second part…why would convergence of performance indicate anything about the absolute performance improvements of frontier models?

> Remind me what was so great about gpt 5? How about gpt4 from from gpt 3? Do you even remember the releases? Yeah. I dont. I had to look it up.

3 -> 4 -> 5 were extraordinary leaps…not sure how one would be able to say anything else

> Just another model with more or less the same capabilities.

5 is absolutely not a model with more or less the same capabilities as gpt 4, what could you mean by this?

> “Mixed reception”

A mixed reception is an indication of model performance against a backdrop of market expectations, not against gpt 4…

> That is not what exponential progress looks like, by any measure.

Sure it is…exponential is a constant % improvement per year. We’re absolutely in that regime by a lot of measures

> The progress this year has been in the tooling around the models, smaller faster

Effective tool use is not somehow some trivial add on it is a core capability for which we are on an exponential progress curve.

> models with similar capabilities. Multimodal add ons that no one asked for, because its easier to add image and audio processing than improve text handling.

This is definitely a personal feeling of yours, multimodal models are not something no one asked for…they are absolutely essential. Text data is essential and data curation is non trivial and continually improving, we are also hitting the ceiling of internet text data. But yet we use an incredible amount of synthetic data for RL and this continues to grow……you guessed it, exponentially. and multimodal data is incredibly information rich. Adding multi modality lifts all boats and provides core capabilities necessary for open world reasoning and even better text data (e.g. understanding charts and image context for text).

Re: 2025: The Year in LLMs

#582
post #493
post #443

Earlier quoted context omitted.

It's only low cost for general usage chat users. If you are using it for anything beyond that, you are paying or sitting in a long queue (likely both). You may just be a little early to the renaissance. What happens when the models we have today run on a mobile device? The nokia 6110 was released 15 years after the first commercial cell phone.

Yes although even those people paying are likely still being subsidized and not currently paying the full cost. Interesting thought about current SOTA models running on my mobile device. I've given it some thought and I don't think it would change my life in any way. Can you suggest some way that it would change yours?

It will open access of llms to developers in the same way smart phones opened access to mobile general computing.

I really think most everyone misses the actual potential of llms. They aren't an app but an interface.

They are the new UI everyone has known they wanted going back as long as we've had computers. People wanted to talk to the computer and get results.

Think of the people already using them instead of search engines.

To me, and likely you, it doesn't add any value. I can get the same information at about the same speed as before with the same false positives to weed through.

To the person that couldn't use a search engine and filled the internet with easily answered questions before, it's a godsend. They can finally ask the internet in plain ole whatever language they use and get an answer. It can be hard to see, but this is the majority of people on this planet.

LLMs raise the floor of information access. When they become ubiquitous and basically free, people will forget they ever had to use a mouse or hunt for the right pixel to click a button on a tiny mobile device touch screen.

Re: 2025: The Year in LLMs

#583

You’re absolutely right! You astutely observed that 2025 was a year with many LLMs and this was a selection of waypoints, summarized in a helpful timeline. That’s what most non-tech-person’s year in LLMs looked like. Hopefully 2026 will be the year where companies realize that implementing intrusive chatbots can’t make better ::waving hands:: ya know… UX or whatever. For some reason, they think its helpful to distrac…

I took the good with the bad: the ai assisted coding tools are a multiplier, google ai overviews in search results are half baked (at best) and often just factually wrong. AI was put in the instagram search bar for no practical purpose etc.

Yeah totally. The point I’m trying to make, however, is that most people don’t code, so they didn’t get the multiplier, and only got the mediocre-to-bad, with a handful of them doing things like generating dumb images for a boost. I think that’s why a lot of people in the software business are utterly bewildered when customers aren’t jumping for joy when they release a new AI “feature.” I think a lot of what gets classified as cynical ceo enshittification is really people ignoring basic good design practices, like making sure you’re effectively helping customers solve an actual problem in a context and with methods they, at least, don’t hate. Especially on the smaller scale, like indie app developers who probably get more out of AI than most, they really think people are going to like new AI features simply because they’re new AI features. They’re very wrong.

Re: 2025: The Year in LLMs

#584

Earlier quoted context omitted.

What you’re describing is just competent engineering, and it’s already been applied to LLMs. People have been adversarial. That’s why we know so much about hallucinations, jailbreaks, distribution shift failures, and long-horizon breakdowns in the first place. If this were hobbyist awe, none of those benchmarks or red-teaming efforts would exist. The key point you’re missing is the type of failure. Search systems fai…

> If this were hobbyist awe, none of those benchmarks or red-teaming efforts would exist. Absolutely not true. I cannot express how strongly this is not true, haha. The tech is neat, and plenty of real computer scientists work on it. That doesn't mean it's not wildly misunderstood by others. > Concluding from those failure modes that this is just Clever Hans is not adversarial engineering. I feel like you're maybe mi…

Except of course it's not true lol. Horses are smart critters, but they absolutely cannot do arithmetic no matter how much you train them.

These things are not horses. How can anyone choose to remain so ignorant in the face of irrefutable evidence that they're wrong?

https://arxiv.org/abs/2507.15855

It's as if a disease like COVID swept through the population, and every human's IQ dropped 10 to 15 points while our machines grew smarter to an even larger degree.

Re: 2025: The Year in LLMs

#585

Earlier quoted context omitted.

> The answer is obvious. It ALREADY knew the truth. There’s no other logical way to explain this. I can think of several offhand. 1. The effect was never real, you've just convinced yourself it is because you want it to be, ie you Clever Hans'd yourself. 2. The effect is an artifact of how you measure "truth" and disappears outside that context ("It can be wildly off for certain things") 3. The effect was completely…

You asked for something concrete, so I’ll anchor every claim to either documented results or directly observable training mechanics. First, the claim that RLHF materially reduces hallucinations and increases factual accuracy is not anecdotal. It shows up quantitatively in benchmarks designed to measure this exact thing, such as TruthfulQA, Natural Questions, and fact verification datasets like FEVER. Base models and…

Your three alternatives don’t survive contact with this. Clever Hans fails because the effect generalizes. Measurement artifact fails because multiple independent metrics move together. Fraud fails because these results are reproduced across competing labs, companies, and open-source implementations.

He doesn't care. You might as well be arguing with a Scientologist.

Re: 2025: The Year in LLMs

#586

Earlier quoted context omitted.

Many more are employed while building it. And they will never stop building. It's modern version of rail. But instead of distances it will cover the area.

Will local folks get those jobs to build the data center? And if so, what happens to those builders once the data center is built?

> Will local folks get those jobs to build the data center?

Yes. At some point the demand will be so high that imported workers won't suffice and local population will need to be trained and hired.

> And if so, what happens to those builders once the data center is built?

They are going to be moved to a new place where the datacenters will need to be built next. Mobility if the workforce was often cited as one of the greatest strengths of US economy.

Re: 2025: The Year in LLMs

#588
post #94

[flagged]

Could you please stop posting dismissive, curmudgeonly comments? It's not what this site is for, and destroys what it is for. We want curious conversation here. https://news.ycombinator.com/newsguidelines.html

Who is we?

I want LLM astroturfers to have their reputations destroyed for pushing this idiocy on us

Re: 2025: The Year in LLMs

#589
post #50

[flagged]

This is extremely dismissive. Claude Code helps me make a majority of changes to our codebase now, particularly small ones, and is an insane efficiency boost. You may not have the same experience for one reason or another, but plenty of devs do, so "nothing happened" is absolutely wrong. 2024 was a lot of talk, a lot of "AI could hypothetically do this and that". 2025 was the year where it genuinely started to enter…

LLMs must be dismissed.

The dismissive tone is warranted.

Post reply on HN