Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

631–640 of 643 posts

Re: 2025: The Year in LLMs

#631

Earlier quoted context omitted.

I'm saying that the kind of changes you propose aren't made by anyone, and might generally not be worth making. Because "better RLVR" is an easier and better pathway to actual cross-domain performance gains. If you could stabilize the kind of mess you want to make, you could put that effort into better RL objectives and get more return.

The mainstream LLM crowd aren't making these sorts of major changes yet, although some like DeepMind (the OG pushers of RL for AGI!) do acknowledge that a few more "transformer level" breakthoughs are necessary to reach what THEY are calling AGI, and others like LeCun are calling for more animal-like architectures. Anyways, regardless of who is currently trying to move beyond LLMs or not, it should be pretty obvious…

The claim I make is that LLMs can be AGI complete with pretty much zero architectural work. And none of the "brain-like" architectures are actually much better at "being a brain" than LLMs are - the issue isn't "the architecture is wrong", it's "we don't know how to train for this kind of thing".

"Fundamental limitations" aren't actually fundamental. If you want more learning than what "in-context" gives you? Teach the usual "CLI agent" LLM to make its own LoRAs and there goes that. So far, this isn't a bottleneck so pressing you'd want to resolve it by force.

LeCun is laughing stock nowadays, he didn't get kicked out of Meta for no reason.

Re: 2025: The Year in LLMs

#632

Earlier quoted context omitted.

Except of course it's not true lol. Horses are smart critters, but they absolutely cannot do arithmetic no matter how much you train them. These things are not horses. How can anyone choose to remain so ignorant in the face of irrefutable evidence that they're wrong? https://arxiv.org/abs/2507.15855 It's as if a disease like COVID swept through the population, and every human's IQ dropped 10 to 15 points while our ma…

(Continuing from my other post) The first thing I checked was "how did they verify the proofs were correct" and the answer was they got other AI people to check it, and those people said there were serious problems with the paper's methodology and it would not be a gold medal. https://x.com/j_dekoninck/status/1947587647616004583 This is why we do not take things at face value.

That tweet is aimed at Google. I don't know much about Google's effort at IMO, but OpenAI was the primary newsmaker in that event, and they reportedly did not use hints or external tools. If you have info to the contrary, please share it so I can update that particular belief.

Gemini 2.5 has since been superceded by 3.0, which is less likely to need hints. 2.5 was not as strong as the contemporary GPT model, but 3.0 with Pro Thinking mode enabled is up there with the best.

Finally, saying, "Well, they were given some hints" is like me saying, "LOL, big deal, I could drag a Tour peleton up Col du Galibier if I were on the same drugs Lance was using."

No, in fact I could do no such thing, drugs or no drugs. Similarly, a model that can't legitimately reason will not be able to solve these types of problems, even if given hints.

Re: 2025: The Year in LLMs

#633

Earlier quoted context omitted.

Except of course it's not true lol. Horses are smart critters, but they absolutely cannot do arithmetic no matter how much you train them. These things are not horses. How can anyone choose to remain so ignorant in the face of irrefutable evidence that they're wrong? https://arxiv.org/abs/2507.15855 It's as if a disease like COVID swept through the population, and every human's IQ dropped 10 to 15 points while our ma…

Or -- and hear me out -- that result doesn't mean what you think it does. That's the exact reason I mention the Clever Hans story. You think it's obvious because you can't come up with any other explanation, therefore there can't be another explanation and the horse must be able to do math. And if I can't come up with an explanation, well that just proves it, right? Those are the only two options, obviously. Except n…

Have you actually read the paper, or are you just waving it around?

I've spent a lot of time feeding similar problems to various models to understand what they can and cannot do well at various stages of development. Reading papers is great, but by the time a paper comes out in this field, it's often obsolete. Witness how much mileage the ludds still get out of the METR study, which was conducted with a now-ancient Claude 3.x model that wasn't at the top of the field when it was new.

Here, let me call a shot -- I bet this paper says LLMs fuck up on proofs like they fuck up on code. It will sometimes generate things that are fine, but it'll frequently generate things that are just irrational garbage.

And the goalposts have now been moved to a dark corner of the parking garage down the street from the stadium. "This brand-new technology doesn't deliver infallible, godlike results out of the box, so it must just be fooling people." Or in equestrian parlance, "This talking horse told me to short NVDA. What a scam."

Re: 2025: The Year in LLMs

#634

Earlier quoted context omitted.

The mainstream LLM crowd aren't making these sorts of major changes yet, although some like DeepMind (the OG pushers of RL for AGI!) do acknowledge that a few more "transformer level" breakthoughs are necessary to reach what THEY are calling AGI, and others like LeCun are calling for more animal-like architectures. Anyways, regardless of who is currently trying to move beyond LLMs or not, it should be pretty obvious…

The claim I make is that LLMs can be AGI complete with pretty much zero architectural work. And none of the "brain-like" architectures are actually much better at "being a brain" than LLMs are - the issue isn't "the architecture is wrong", it's "we don't know how to train for this kind of thing". "Fundamental limitations" aren't actually fundamental. If you want more learning than what "in-context" gives you? Teach t…

You keep using the term "AGI" without defining what you mean by it, other than implicity defining it as "whatever can be achieved without changing the Transformer architecture", which makes your "claim" just a definitional tautology, which is fine, but it does mean you are talking about something different than what I am talking about, which is also fine.

> And none of the "brain-like" architectures are actually much better at "being a brain" than LLMs are

I've no idea what projects you are referring to.

It would certainly be bizarre if the Transformer architecture, never designed to be a brain, turns out to be the best brain we can come up with, and equal to real brains which have many more moving parts, each evolved over millions of years to fill a need and improve capability.

Maybe you are smarter than Demis Hassabis, and the DeepMind team, and all their work towards AGI (their version, not yours) will be a waste of effort. Why not send him a note "hey, dumbass, Transfomers are all you need!" ?

Re: 2025: The Year in LLMs

#635

Earlier quoted context omitted.

Uber nakedly broke the law and beat down labor, I'm honestly shocked none of the executives went to prison.

Uber didn’t beat down labor, they beat down capital, specifically the capital that owned (and lobbied for the existence of) taxi medallions

No, you clearly have never talked to the workers at Uber (no not the devs, the drivers). Uber has disgustingly fought against unionization efforts, employee benefit efforts, and increasing wages. Such actions do not make you a good company, especially when the executives make millions while fighting against workers wanting a better life.

They are an evil company and the rot has been there since inception. This isn't even getting into their disgusting internal predator culture against women either.

Re: 2025: The Year in LLMs

#636

Earlier quoted context omitted.

> Will local folks get those jobs to build the data center? Yes. At some point the demand will be so high that imported workers won't suffice and local population will need to be trained and hired. > And if so, what happens to those builders once the data center is built? They are going to be moved to a new place where the datacenters will need to be built next. Mobility if the workforce was often cited as one of the…

So local people in town 1 who are getting these jobs to build the data center will then have to move to town 2 to build a data center there? What happens to the local people in town 2 who are also looking for construction jobs?

Local people in town 2 share the same fate that people in town 1 alread had. If there's not enough imported workers, from town 1 or elsewere people from town 2 will need to be trained and employed.

More and more data centers (and power sources) are going to be built at the same time so more and more workers will be needed. This is going to be THE job. I think there are going to be many similarities with the age when railroads were being developed. Hopefully with less worker deaths this time.

Re: 2025: The Year in LLMs

#637

Earlier quoted context omitted.

The claim I make is that LLMs can be AGI complete with pretty much zero architectural work. And none of the "brain-like" architectures are actually much better at "being a brain" than LLMs are - the issue isn't "the architecture is wrong", it's "we don't know how to train for this kind of thing". "Fundamental limitations" aren't actually fundamental. If you want more learning than what "in-context" gives you? Teach t…

You keep using the term "AGI" without defining what you mean by it, other than implicity defining it as "whatever can be achieved without changing the Transformer architecture", which makes your "claim" just a definitional tautology, which is fine, but it does mean you are talking about something different than what I am talking about, which is also fine. > And none of the "brain-like" architectures are actually much…

It would be certainly be bizarre if the 8086 architecture, never designed to be a foundation of all home, office and server computation, was the best CPU architecture ever made.

And it isn't. It's merely good enough.

That's what LLMs are. A "good enough" AI architecture.

By "AGI", I mean the good old "human equivalence" proxy. An AI that can accomplish any intellectual task that can be accomplished by a human. LLMs are probably sufficient for that. They have limitations, but not the kind that can't be worked around with things like sharply applied tool use. Which LLMs can be trained for, and are.

So far, all the weirdo architectures that try to replace transformers, or put brain-inspired features into transformers, have failed to live up to the promise. Which sure hints that the bottleneck isn't architectural at all.

Re: 2025: The Year in LLMs

#638

Earlier quoted context omitted.

You keep using the term "AGI" without defining what you mean by it, other than implicity defining it as "whatever can be achieved without changing the Transformer architecture", which makes your "claim" just a definitional tautology, which is fine, but it does mean you are talking about something different than what I am talking about, which is also fine. > And none of the "brain-like" architectures are actually much…

It would be certainly be bizarre if the 8086 architecture, never designed to be a foundation of all home, office and server computation, was the best CPU architecture ever made. And it isn't. It's merely good enough. That's what LLMs are. A "good enough" AI architecture. By "AGI", I mean the good old "human equivalence" proxy. An AI that can accomplish any intellectual task that can be accomplished by a human. LLMs a…

I'm not aware of any architectures that have tried to put "brain-inspired" features into Transformers, or much attempt to modify them at all for that matter.

The architectural Transformer tweaks that we've seen are:

- Various versions of attention for greater efficiency

- MOE vs dense for greater efficiency

- Mamba (SSM) + transformer hybrid for greater efficiency

None of these are even trying to fundamentally change what the Transformer is doing.

Yeah, the x86 architecture is certainly a bit of a mess, but as you say good enough, as long as what you want to do is run good old fashioned symbolic computer programs. However, if you want to run these new-fangled neural nets, then you'd be better off with a GPU or TPU.

> By "AGI", I mean the good old "human equivalence" proxy. An AI that can accomplish any intellectual task that can be accomplished by a human. LLMs are probably sufficient for that.

I think DeepMind are right here, and you're wrong, but let's wait another year or two and see, eh?

Re: 2025: The Year in LLMs

#639

Earlier quoted context omitted.

It’s possible this is correct. It’s also possible that people more experienced, knowledgable and skilled than you can see fundamental flaws in using LLMs for software engineering that you cannot. I am not including myself in that category. I’m personally honestly undecided. I’ve been coding for over 30 years and know something like 25 languages. I’ve taught programming to postgrad level, and built prototype AI system…

But right now these tools don’t work well. Great for vibe coding, prototyping, analysis, review, bouncing ideas. What are some of the models you've been working with?

All the major models from Anthropic, OpenAI, Google. I’ve probably used Gemini the least.

Re: 2025: The Year in LLMs

#640

Earlier quoted context omitted.

People keep saying stuff like this. That the improvements are so obvious and breathtaking and astronomical and then I go check out the frontier LLMs again and they're maybe a tiny bit better than they were last year but I can't actually be sure bcuz it's hard to tell. sometimes it seems like people are just living in another timeline.

You might want to be more specific because benchmarks abound and they paint a pretty consistent picture. LMArena "vibes" paint another picture. I don't know what you are doing to "check" the frontier LLMs but whatever you're doing doesn't seem to match more careful measurement... You don't actually have to take peoples word for it, read epoch.ai developments, look into the benchmark literature, look at ARC-AGI...

That's half the problem though. I can see benchmarks. I can see number go up on some chart or that the AI scores higher on some niche math or programming test, but those results don't seem to actually connect much to meaningful improvements in daily usage of the software when those updates hit the public.

That's where the skepticism comes in, because one side of the discussion is hyping up exponential growth and the other is seeing something that looks more logarithmic instead.

I realize anecdotes aren't as useful as numbers for this kind of analysis, but there's such a wide gap between what people are observing in practice and what the tests and metrics are showing it's hard not to wonder about those numbers.

Post reply on HN