Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

411–420 of 643 posts

Re: 2025: The Year in LLMs

#411
post #406

Most LLMs got worse in 2025. Only addicts and the type of computer gamer that feels drawn to complex setups, gamification and does not care about the end result will feel positive about the grift. 2025: The Year in Open Source? Nothing, all resources were tied up to debunk a couple of Python web developers who pose as the ultimate experts in LLMs.

In what way did they get worse?

I made you a dashboard of my 2025 writing about open-source that didn't include AI: https://simonwillison.net/dashboard/posts-with-tags-in-a-yea...

Re: 2025: The Year in LLMs

#412

Earlier quoted context omitted.

I will find this often-repeated argument compelling only when someone can prove to me that the human mind works in a way that isn't 'combining stuff it learned in the past'. 5 years ago a typical argument against AGI was that computers would never be able to think because "real thinking" involved mastery of language which was something clearly beyond what computers would ever be able to do. The implication was that t…

> 5 years ago a typical argument against AGI was that computers would never be able to think because "real thinking" involved mastery of language which was something clearly beyond what computers would ever be able to do. Mastery of words is thinking? In that line of argument then computers have been able to think for decades. Humans don't think only in words. Our context, memory and thoughts are processed and occur…

Mastery of words is thinking?

That's the crazy thing. Yes, in fact, it turns out that language encodes and embodies reasoning. All you have to do is pile up enough of it in a high-dimensional space, use gradient descent to model its original structure, and add some feedback in the form of RL. At that point, reasoning is just a database problem, which we currently attack with attention.

No one had the faintest clue. Even now, many people not only don't understand what just happened, but they don't think anything happened at all.

ELIZA, ROFL. How'd ELIZA do at the IMO last year?

Re: 2025: The Year in LLMs

#413

Earlier quoted context omitted.

> They are helping their users create things that didn't exist before. That is a derived output. That isn't new as in: novel. It may be unique but it is derived from training data. LLMs legitimately cannot think and thus they cannot create in that way.

That is a pedantic distinction. You can create something that didn't exist by combining two things that did exist, in a way of combining things that already existed. For example, you could use a blender to combine almond butter and sawdust. While this may not be "novel", and it may be derived from existing materials and methods, you may still lay claim to having created something that didn't exist before. For a more…

Pedantic and not true. The LLM has stochastic processes involved. Randomness. That’s not old information. That’s newly generated stuff.

Re: 2025: The Year in LLMs

#414
post #228

All these improvement in a single year, 2025. While this may seem obvious to those who follows along the AI / LLM news. It may be worth pointing out again ChatGPT was introduced to us in November 2022. I still dont believe AGI, ASI or Whatever AI will take over human in short period of time say 10 - 20 years. But it is hard to argue against the value of current AI, which many of the vocal critics on HN seems to have…

This is not a great argument: > But it is hard to argue against the value of current AI [...] it is getting $1B dollar runway already. The psychic services industry makes over $2 billion a year in the US [1], with about a quarter of the population being actual believers. [2]. [1] The https://www.ibisworld.com/united-states/industry/psychic-ser... [2] https://news.gallup.com/poll/692738/paranormal-phenomena-met...

2022/2023: "It hallucinates, it's a toy, it's useless."

2024/2025: "Okay, it works, but it produces security vulnerabilities and makes junior devs lazy."

2026 (Current): "It is literally the same thing as a psychic scam."

Can we at least make predictions for 2027? What shall the cope be then! Lemme go ask my psychic.

Re: 2025: The Year in LLMs

#415
post #366

> The reason I think MCP may be a one-year wonder is the stratospheric growth of coding agents. It appears that the best possible tool for any situation is Bash—if your agent can run arbitrary shell commands, it can do anything that can be done by typing commands into a terminal. I push back strongly from this. In the case of the solo, one-machine coder, this is likely the case - if you're exposing workflows or fixed…

The solution to that is Anthropic's Skills. Create a folder called skills/how-to-use-jira Add several Bash scripts with the right curl commands to perform specific actions Add a SKILL.md file with some instructions in how to use those scripts You've effectively flattened that MCP server into some Markdown and Bash, only the thing you have now is more flexible (the coding agent can adapt those examples to cover new th…

But that moves the burden of maintenance from the provider of the service to its users (and/or partially to intermediary in form of "skills registry" of sorts, which apparently is a thing now).

So maybe a hybrid approach would make more sense? Something like /.well-known/skills/README.md exposed and owned by the providers?

That is assuming that the whole idea of "skills" makes sense in practice.

Re: 2025: The Year in LLMs

#416

Earlier quoted context omitted.

> You really owe it to yourself to try them out. I've worked at multiple AI startups in lead AI Engineering roles, both working on deploying user facing LLM products and working on the research end of LLMs. I've done collaborative projects and demos with a pretty wide range of big names in this space (but don't want to doxx myself too aggressively), have had my LLM work cited on HN multiple times, have LLM based gith…

Over half of HN still thinks it’s a stochastic parrot and that it’s just a glorified google search. The change hit us so fast a huge number of people don’t understand how capable it is yet. Also it certainly doesn’t help that it still hallucinates. One mistake and it’s enough to set someone against LLMs. You really need to push through that hallucinations are just the weak part of the process to see the value.

The problem I see, over and over, is that people pose poorly-formed questions to the free ChatGPT and Google models, laugh at the resulting half-baked answers that are often full of errors and hallucinations, and draw conclusions about the technology as a whole.

Either that, or they tried it "last year" or "a while back" and have no concept of how far things have gone in the meantime.

It's like they wandered into a machine shop, cut off a finger or two, and concluded that their grandpa's hammer and hacksaw were all anyone ever needed.

Re: 2025: The Year in LLMs

#417
post #414

Earlier quoted context omitted.

This is not a great argument: > But it is hard to argue against the value of current AI [...] it is getting $1B dollar runway already. The psychic services industry makes over $2 billion a year in the US [1], with about a quarter of the population being actual believers. [2]. [1] The https://www.ibisworld.com/united-states/industry/psychic-ser... [2] https://news.gallup.com/poll/692738/paranormal-phenomena-met...

2022/2023: "It hallucinates, it's a toy, it's useless." 2024/2025: "Okay, it works, but it produces security vulnerabilities and makes junior devs lazy." 2026 (Current): "It is literally the same thing as a psychic scam." Can we at least make predictions for 2027? What shall the cope be then! Lemme go ask my psychic.

2022/2023: "Next year software engineering is dead"

2024: "Now this time for real, software engineering is dead in 6 months, AI CEO said so"

2025: "I know a guy who knows a guy who built a startup with an LLM in 3 hours, software engineering is dead next year!"

What will be the cope for you this year?

Re: 2025: The Year in LLMs

#418

Earlier quoted context omitted.

> exponential progress First you need to define what it means. What's the metric? Otherwise it's very much something you can argue about.

> What's the metric? Language model capability at generating text output. The model progress this year has been a lot of: - “We added multimodal” - “We added a lot of non AI tooling” (ie agents) - “We put more compute into inference” (ie thinking mode) So yes, there is still rapid progress, but these ^ make it clear, at least to me, that next gen models are significantly harder to build. Simultaneously we see a disti…

I don’t think the path was ever exponential but your claim here is almost as if the slow down hit an asymptote like wall.

Most of the improvements are intangible. Can we truly say how much more reliable the models are? We barely have quantitative measurements on this so it’s all vibes and feels. We don’t even have a baseline metric for what AGI is and we invalidated the Turing test also based on vibes and feels.

So my argument is that part of the slow down is in itself an hallucination because the improvement is not actually measurable or definable outside of vibes.

Re: 2025: The Year in LLMs

#419
post #366

Earlier quoted context omitted.

The solution to that is Anthropic's Skills. Create a folder called skills/how-to-use-jira Add several Bash scripts with the right curl commands to perform specific actions Add a SKILL.md file with some instructions in how to use those scripts You've effectively flattened that MCP server into some Markdown and Bash, only the thing you have now is more flexible (the coding agent can adapt those examples to cover new th…

But that moves the burden of maintenance from the provider of the service to its users (and/or partially to intermediary in form of "skills registry" of sorts, which apparently is a thing now). So maybe a hybrid approach would make more sense? Something like /.well-known/skills/README.md exposed and owned by the providers? That is assuming that the whole idea of "skills" makes sense in practice.

Yeah that's true, skill distribution isn't a solved problem yet - MCPs have a URL, which is a great way of making them available for people to start using without extra steps.

Re: 2025: The Year in LLMs

#420
post #395

With everything that we have done so far (our company) I believe by end of 2026 our software will be self improving all the time. And no it is not AI slop and we don't vibe code. There are a lot of practical aspects of running software and maintaining / improving code that can be done well with AI if you have the right setup. It is hard to formulate what "right" looks like at this stage as we are still iterating on t…

What exactly are your agents doing overnight? I often hear folks talk about their agents running for long periods of time but rarely talk about the outcomes they're driving from those agents.

We have a lot of grunt work scheduled overnight like finding bugs, creating tests where we don’t have good coverage or where we can improve, integrations, documentation work, etc.

Not everything gets accepted. There is a lot of work that is discarded and much more pending verification and acceptance.

Frankly, and I hope I don’t come as alarmist (judge for yourself from my previous comments on Hn and Reddit) we cannot keep up with the output! And a lot of it is actually good and we should incorporate it even partially.

At the moment we are figuring out how to make things more autonomous while we have the safety and guardrails in place.

The biggest issue I see at this stage is how to make sense of it all as I do not believe we have the understanding of what is happening - just the general notion of it.

I truly believe that we will reach the point where ideas matter more than execution, which what I would expect to be the case with more advanced and better applied AI.

Post reply on HN