Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

611–620 of 643 posts

Re: 2025: The Year in LLMs

#611

Earlier quoted context omitted.

How's the replication rate in that field? Last I heard it was below 50%. How can you think without tokens of some sort? That's half of the question that has to be answered by the linguists. The other half is that if language isn't necessary for reasoning, what is? We now know that a conceptually-simple machine absolutely can reason with nothing but language as inputs for pretraining and subsequent reinforcement. We d…

Read about linguistic history and make up your own mind, I guess. Or don’t, I don’t care. You’re dismissing a series of highly robust scientific results because they fail to validate your beliefs, which is highly irrational. I'm no longer interested in engaging with you.

I've read plenty of linguistics work on a lay basis. It explains little and predicts even less, so it hasn't exactly encouraged me to delve further into the field. That said, linguistics really has nothing to do with arguments with the Moon-landing deniers in this thread, who are the people you should really be targeting with your advocacy of rationality.

In other words, when I (seem to) dismiss an entire field of study, it's because it doesn't work, not because it does work and I just don't like the results.

Re: 2025: The Year in LLMs

#612
post #456

Earlier quoted context omitted.

Did Uber actually do a lot of capital investment? They don't own the cars, for example.

Uber nakedly broke the law and beat down labor, I'm honestly shocked none of the executives went to prison.

Uber didn’t beat down labor, they beat down capital, specifically the capital that owned (and lobbied for the existence of) taxi medallions

Re: 2025: The Year in LLMs

#613
post #109

Earlier quoted context omitted.

> in turn act irrationally it isn't irrational to act in self-interest. If LLM threatens someone's livelihood, it matters not that it helps humanity overall one bit - they will oppose it. I don't blame them. But i also hope that they cannot succeed in opposing it.

It's irrational to genuinely hold false beliefs about capabilities of LLMs. But at this point I assume around half of the skeptics are emotionally motivated anyway.

everybody is emotionally motivated, you included

Re: 2025: The Year in LLMs

#614

Earlier quoted context omitted.

> What's the metric? Language model capability at generating text output. The model progress this year has been a lot of: - “We added multimodal” - “We added a lot of non AI tooling” (ie agents) - “We put more compute into inference” (ie thinking mode) So yes, there is still rapid progress, but these ^ make it clear, at least to me, that next gen models are significantly harder to build. Simultaneously we see a disti…

Next gen models are always hard to build, they are by definition pushing the frontier. Every generation of CPU was hard to build but we still had Moores law. > Simultaneously we see a distinct narrowing between players (openai, deepseek, mistral, google, anthropic) in their offerings. Thats usually a signal that the rate of progress is slowing. I agree with you on the fact in the first part but not the second part…wh…

> exponential is a constant % improvement per year

I suppose of you pick a low enough exponent then the exp graph is flat for a long time and you're right, zero progress is “exponential” if you cherry pick your growth rate to be low enough.

Generally though, people understand “exponential growth” as “getting better/bigger faster and faster in an obvious way

> 3 -> 4 -> 5 were extraordinary leaps…not sure how one would be able to say anything else

They objectively were not.

The metrics and reception to them was very clear and overwhelming.

Youre spitting some meaningless revisionist BS here.

Youre wrong.

Thats all there is to it.

Re: 2025: The Year in LLMs

#615

Earlier quoted context omitted.

Yes, if all you did was replace current pre/mid/post training with a new (elusive holy grail) runtime continual learning algorithm, then it would definitely still just be a language model. You seem to be talking about it having TWO runtime continual learning algorithms, next-token and long-horizon RL, but of course RL is part of what we're calling an LLM. It's not obvious if you just did this without changing the lea…

No. There are no architectural changes and no "second runtime learning algorithm". There's just the good old in-context learning that all LLMs get from pre-training. RLVR is a training stage that pressures the LLM to take advantage of it on real tasks. "Runtime continual learning algorithm" is an elusive target of questionable desirability - given that we already have in-context learning, and "get better at SFT and R…

I'm not sure what you are saying. There are LLMs as exist today, and there are any number of changes one could propose to make to them.

The less you change, the more they stay the same. If you just add "more" RLVR (perhaps for a new domain - maybe chemistry vs math or programming?), then all you will get is an LLM that is better at acing chemistry reasoning benchmarks.

Re: 2025: The Year in LLMs

#616

Earlier quoted context omitted.

Next gen models are always hard to build, they are by definition pushing the frontier. Every generation of CPU was hard to build but we still had Moores law. > Simultaneously we see a distinct narrowing between players (openai, deepseek, mistral, google, anthropic) in their offerings. Thats usually a signal that the rate of progress is slowing. I agree with you on the fact in the first part but not the second part…wh…

> exponential is a constant % improvement per year I suppose of you pick a low enough exponent then the exp graph is flat for a long time and you're right, zero progress is “exponential” if you cherry pick your growth rate to be low enough. Generally though, people understand “exponential growth” as “getting better/bigger faster and faster in an obvious way ” > 3 -> 4 -> 5 were extraordinary leaps…not sure how one wo…

Doesn’t sound like you really seem to be interested in any sort of rational dialogue, metrics were “objectively” not better? What are you talking about of course they were have you even looked at benchmark progression for every benchmark we have?

You don’t understand what an exponential is or apparently what the benchmark numbers even are or possibly even how we actually measure model performance and the very real challenges and nuances involved but yet I’m “spitting some revisionist BS”. You have cited zero sources and are calling measured numbers “revisionist”.

You are also citing reception to models as some sort of indication of their performance, which is yet another confusing part of your reasoning.

I do agree that “metrics were were very clear” it just seems you don’t happen to understand what they are or what they mean.

Re: 2025: The Year in LLMs

#617

Earlier quoted context omitted.

"maybe a tiny bit better" is what you say when you've been tricked by snake oil salesman This shit has gotten worse since 2023.

> This shit has gotten worse since 2023. I would really appreciate it if people could be specific when they say stuff like this because it's so crazy out of line with all measurement efforts. There are an insane amount of serious problems with current LLM / agentic paradigms, but the idea that things have gotten worse since 2023 ? I mean come on.

You’re responding to a troll who just has a nasty, bitter axe to grind against AI. It’s honestly pretty sad and pathetic.

Re: 2025: The Year in LLMs

#618

Earlier quoted context omitted.

I’d put more faith in HN’s proclamations if it hadn’t widely been wrong about AI in 2023, 2024, and now 2025. Watching the tone shift here has been fascinating. As the saying goes, the only thing moving faster than AI advances right now is the speed at which HN haters move the goalposts…

Mmm. People who make AI their entire personality and brag that other people are too stupid to see what they see and soon they'll have to see the genius they're denying...does not make me think "oh, wow, what have I missed in AI".

[deleted]

Re: 2025: The Year in LLMs

#619

Earlier quoted context omitted.

> 5 years ago a typical argument against AGI was that computers would never be able to think because "real thinking" involved mastery of language which was something clearly beyond what computers would ever be able to do. Mastery of words is thinking? In that line of argument then computers have been able to think for decades. Humans don't think only in words. Our context, memory and thoughts are processed and occur…

Mastery of words is thinking? That's the crazy thing. Yes, in fact, it turns out that language encodes and embodies reasoning. All you have to do is pile up enough of it in a high-dimensional space, use gradient descent to model its original structure, and add some feedback in the form of RL. At that point, reasoning is just a database problem, which we currently attack with attention. No one had the faintest clue. E…

> Yes, in fact, it turns out that language encodes and embodies reasoning ... No one had the faintest clue

Funnily enough, they did, if you go back far enough. It's only the deconstructionists and the solipsists who had the audacity to think otherwise.

Re: 2025: The Year in LLMs

#620

Earlier quoted context omitted.

> ELIZA, ROFL. How'd ELIZA do at the IMO last year? What's funny is the failure to grasp any contextual framing of ELIZA. When it came out people were impressed by it's reasoning, it's responses. And in your line of defense it could think because it had mastery of words! But fast forward the current timeline 30 years. You will have been of the same camp that argued on behalf of ELIZA when the rest of the world was as…

No one was impressed with ELIZA's "reasoning" except for a few non-specialist test subjects recruited from the general population. Admittedly it was disturbing to see how strongly some of those people latched onto it. Meanwhile, you didn't answer my question. How'd ELIZA do on the IMO? If you know a way to achieve gold-medal performance at top-level math and programming competitions without thinking, I for one am all…

Does a prolog program think?
Post reply on HN