Live data from Hacker News

2025: The Year in LLMs

simonwillison.net

131–140 of 643 posts

Re: 2025: The Year in LLMs

#131

Earlier quoted context omitted.

LLMs hold some real utility. But that real utility is buried under a mountain of fake hype and over-promises to keep shareholder value high. LLMs have real limitations that aren't going away any time soon - not until we move to a new technology fundamentally different and separate from them - sharing almost nothing in common. There's a lot of 'progress-washing' going on where people claim that these shortfalls will m…

Pretty much. What actually exists is very impressive. But what was promised and marketed has not been delivered.

Markets never deliver. That isnt new, i do think llms are not far off from google in terms of impact.

Search, as of today, is inferior to frontier models as a product. However, best case still misses expected returns by miles which is where the growsing comes from.

Generative art/ai is still up in the air for staying power but id predict it isnt going away.

Re: 2025: The Year in LLMs

#132
post #5

> Vendor-independent options include GitHub Copilot CLI, Amp, OpenHands CLI, and Pi ...and the best of them all, OpenCode[1] :) [1]: https://opencode.ai

Can OpenCode be used with the Claude Max or ChatGPT Pro subscriptions, i.e., without per-token API charges?

Re: 2025: The Year in LLMs

#133
post #5

> Vendor-independent options include GitHub Copilot CLI, Amp, OpenHands CLI, and Pi ...and the best of them all, OpenCode[1] :) [1]: https://opencode.ai

Can OpenCode be used with the Claude Max or ChatGPT Pro subscriptions, i.e., without per-token API charges?

Apparently it does work with Claude Max: https://opencode.ai/docs/providers/#anthropic

I don't see a similar option for ChatGPT Pro. Here's a closed issue: https://github.com/sst/opencode/issues/704

Re: 2025: The Year in LLMs

#134
post #67

Indeed. I don't understand why Hacker News is so dismissive about the coming of LLMs, maybe HN readers are going through 5 stages of grief? But LLM is certainly a game changer, I can see it delivering impact bigger than the internet itself. Both require a lot of investments.

The internet and smartphones were immediately useful in a million different ways for almost every person. AI is not even close to that level. Very to somewhat useful in some fields (like programming) but the average person will easily be able to go through their day without using AI. The most wide-appeal possibility is people loving 100%-AI-slop entertainment like that AI Instagram Reels product. Maybe I'm just too d…

A year after the iPhone came out… it didn’t have an App Store, barely was able to play video, barely had enough power to last a day. You just don’t remember or were not around for it.

A year after llms came out… are you kidding me?

Two years?

10 years?

Today, by adding an MCP server to wrap the same API that’s been around forever for some system, makes the users of that system prefer NLI over the gui almost immediately.

Re: 2025: The Year in LLMs

#135

Earlier quoted context omitted.

LLMs hold some real utility. But that real utility is buried under a mountain of fake hype and over-promises to keep shareholder value high. LLMs have real limitations that aren't going away any time soon - not until we move to a new technology fundamentally different and separate from them - sharing almost nothing in common. There's a lot of 'progress-washing' going on where people claim that these shortfalls will m…

Pretty much. What actually exists is very impressive. But what was promised and marketed has not been delivered.

I think the missing ingredient is not something the LLMs lack, but something we as developers don't do - we need to constrain, channel, and guide agents by creating reactive test environments around them. Not vibes, but hard tests, they are the missing ingredient to coding agents. You can even use AI to write most of these tests but the end result depends on how well you structured your code to be testable.

If you inherit 9000 tests from an existing project you can vibe code a replacement on your phone in a holiday, like Simon Willison's JustHTML port. We are moving from agents semi-randomly flailing around to constraint satisfaction.

Re: 2025: The Year in LLMs

#136
post #35

Earlier quoted context omitted.

My prediction: If we can successfully get rid of most software engineers, we can get rid of most knowledge work. Given the state of robotics, manual labor is likely to outlive intellectual labor.

I would have agreed with this a few months ago, but something Ive learned is that the ability to verify an LLMs output is paramount to its value. In software, you can review its output, add tests, on top of other adversarial techniques to verify the output immediately after generation. With most other knowledge work, I don't think that is the case. Maybe actuarial or accounting work, but most knowledge work exists at…

I also believe this - I think it will probably just disrupt software engineering and any other digital medium with mass internet publication (i.e. things RLVR can use). For the short term future it seems to need a lot of data to train on, and no other profession has posted the same amount of verifiable material. The open source altruism has disrupted the profession in the end; just not in the way people first predicted. I don't think it will disrupt most knowledge work for a number of reasons. Most knowledge professions have "credentials' (i.e. gatekeeping) and they can see what is happening to SWE's and are acting accordingly. I'm hearing it firsthand at least locally in things like law, even accounting, etc. Society will ironically respect these professions more for doing so.

Any data, verifiability, rules of thumb, tests, etc are being kept secret. You pay for the result, but don't know the means.

Re: 2025: The Year in LLMs

#137
post #67

Indeed. I don't understand why Hacker News is so dismissive about the coming of LLMs, maybe HN readers are going through 5 stages of grief? But LLM is certainly a game changer, I can see it delivering impact bigger than the internet itself. Both require a lot of investments.

It is an over correction because of all the empty promises of LLMs. I use Claude and chatgpt daily at work and am amazed at what they can do and how far they can come.

BUT when I hear my executive team talk and see demos of "Agentforce" and every saas company becoming an AI company promising the world, I have to roll my eyes.

The challenge I have with LLMs is they are great at creating first draft shiny objects and the LLMs themselves over promise. I am handed half baked work created by non technical people that now I have to clean up. And they don't realize how much work it is to take something from a 60% solution to a 100% solution because it was so easy for them to get to the 60%.

Amazing, game changing tools in the right hands but also give people false confidence.

Not that they are not also useful for non-technical people but I have had to spend a ton of time explaining to copywriters on the marketing team that they shouldn't paste their credentials into the chat even if it tells them to and their vibe coded app is a security nightmare.

Re: 2025: The Year in LLMs

#138
post #90

[flagged]

I appreciate his work for being more informative and organized than average AI-related content. Without his blogging, it would be a struggle to navigate the bombastic and narcissistic Twitter/Reddit posts for AI updates. The barrier to entry for AI reporting is so low that you just need to give a bit more care to be distinguished, and he is getting the deserved attention for doing exactly that in a systematical and disciplined manner. (I do believe many on HN are more than capable but not interested in doing the same.) Personally, I sometimes find his posts more congratulatory or trivial than I like, but I have learned to take what I want and ignore what I don’t.

Re: 2025: The Year in LLMs

#139
post #98

Earlier quoted context omitted.

I'm not sure what the issue is here but it's not ok to cross into personal attack on HN. We ban accounts that do that, so please don't do it again. https://news.ycombinator.com/newsguidelines.html

how is that a personal attack? a personal attack would be eg calling him a DC. all I did was point out the intellectual dishonesty of his argument. that's an attack on his intellectually dishonest argument, not his person. by all means go ahead and ban me

"I will not pretend you are engaging honestly" is well into the realm of personal attack, and you can't do that here.

Ditto for "I am very disappointed about your BULLSHIT" in the GP comment.

Re: 2025: The Year in LLMs

#140

Earlier quoted context omitted.

LLMs hold some real utility. But that real utility is buried under a mountain of fake hype and over-promises to keep shareholder value high. LLMs have real limitations that aren't going away any time soon - not until we move to a new technology fundamentally different and separate from them - sharing almost nothing in common. There's a lot of 'progress-washing' going on where people claim that these shortfalls will m…

Pretty much. What actually exists is very impressive. But what was promised and marketed has not been delivered.

Yes and most of the investment has been kind of post-GPT4 betting that things will get exponentially more impressive
Post reply on HN