Earlier quoted context omitted.
2022/2023: "Next year software engineering is dead" 2024: "Now this time for real, software engineering is dead in 6 months, AI CEO said so" 2025: "I know a guy who knows a guy who built a startup with an LLM in 3 hours, software engineering is dead next year!" What will be the cope for you this year?
The cope + disappointment will be knowing that a large population of HN users will paint a weird alternative reality. There are a multitude of messages about AI that are out there, some are highly detached from reality (on the optimistic and pessimistic side). And then there is the rational middle, professionals who see the obvious value of coding agents in their workflow and use them extensively (or figure out how t…
2025: The Year in LLMs
471–480 of 643 posts
Re: 2025: The Year in LLMs
#472Earlier quoted context omitted.
I used Gemini Pro, Claude Pro yesterday a couple of dozen times and basically have been daily. I have a project to convert my multiplayer XNA game from C# to Javascript and to add networking to the game-play using LLMs. They are far worse at it now than they were a year ago. They actually implemented the requirements (Though inaccurately) to the best of their ability a year ago. Especially Gemini. Now they don't even…
Weird. I would expect Gemini 3 Pro and Claude Opus 4.5 to run rings around Gemini 1.5 Pro and Claude Sonnet 3.5. How are you running them - regular chat interface or do you have them setup with Claude Code or Gemini CLI?
I am considering making a thread where I compel others to attempt to get what I'm trying to get out of it and show me their work.
The game is only around 25000-30000 LOC in C#.
Re: 2025: The Year in LLMs
#473Earlier quoted context omitted.
2022/2023: "Next year software engineering is dead" 2024: "Now this time for real, software engineering is dead in 6 months, AI CEO said so" 2025: "I know a guy who knows a guy who built a startup with an LLM in 3 hours, software engineering is dead next year!" What will be the cope for you this year?
The cope + disappointment will be knowing that a large population of HN users will paint a weird alternative reality. There are a multitude of messages about AI that are out there, some are highly detached from reality (on the optimistic and pessimistic side). And then there is the rational middle, professionals who see the obvious value of coding agents in their workflow and use them extensively (or figure out how t…
I'd imagine I'm not the only one who has a similar situation. Until all those people and processes can be swept away in favor of letting LLMS YOLO everything into production, I don't see how that changes.
Re: 2025: The Year in LLMs
#474All these improvement in a single year, 2025. While this may seem obvious to those who follows along the AI / LLM news. It may be worth pointing out again ChatGPT was introduced to us in November 2022. I still dont believe AGI, ASI or Whatever AI will take over human in short period of time say 10 - 20 years. But it is hard to argue against the value of current AI, which many of the vocal critics on HN seems to have…
It's a great tool, but right now it's only being used to feed the greed. >> Again, I guess no one knew AI would be as big as it is today, and it is only just started. People have been saying similar about self driving cars for years now. "AI" is another one of those expensive ideas that we'll get 85% of the way there and then to get the other 15% will be way more expensive than anyone will want to pay for. It's alrea…
But stuff like this im not sure I understand:
> It's a great tool, but right now it's only being used to feed the greed.
if its a great tool, then how is it _only_ being used to "feed the greed" and what do you mean by that?
Also I think folks are quick to make analogies to other points in history: "AI is like the dot com boom we're going to crash and burn" and "AI is like {self driving cars, crypto, etc} and the promises will all be broken, its all hype" but this removes the nuance: all of these things are extremely different with very specific dynamics that in _some_ ways may be similar but in many crucial and important ways are completely different.
Re: 2025: The Year in LLMs
#475Indeed. I don't understand why Hacker News is so dismissive about the coming of LLMs, maybe HN readers are going through 5 stages of grief? But LLM is certainly a game changer, I can see it delivering impact bigger than the internet itself. Both require a lot of investments.
> I don't understand why Hacker News is so dismissive about the coming of LLMs I find LLMs incredibly useful, but if you were following along the last few years the promise was for “exponential progress” with a teaser world destroying super intelligence. We objectively are not on that path. There is no “coming of LLMs”. We might get some incremental improvement, but we’re very clearly seeing sigmoid progress. I can’t…
> We might get some incremental improvement, but we’re very clearly seeing sigmoid progress.
again, if it is "very clear" can you point to some concrete examples to illustrate what you mean?
> I can’t speak for everyone, but I’m tired of hyperbolic rants that are unquestionably not justified (the nice thing about exponential progress is you don’t need to argue about it)
OK but what specifically do you have an issue with here?
Re: 2025: The Year in LLMs
#476Earlier quoted context omitted.
That must have been a long time back. Having lived through the time when web pages were served through CGI and mobile phones only existed in movies, when SVMs where the new hotness in ML and people would write about how weird NNs were, I feel like I've seen a lot more concrete progress in the last few decades than this year. This year honestly feels quite stagnant. LLMs are literally technology that can only reproduc…
> LLMs are literally technology that can only reproduce the past. Funny, I've used them to create my own personalized text editor, perfectly tailored to what I actually want. I'm pretty sure that didn't exist before. It's wild to me how many people who talk about LLM apparently haven't learned how to use them for even very basic tasks like this! No wonder you think they're not that powerful, if you don't even know ba…
Without you, there was nothing.
Re: 2025: The Year in LLMs
#477Earlier quoted context omitted.
> He's one of the most valuable writers on LLMs Is he, really? Most of his blog posts are little more than opportunistic, buttressing commentary on someone else's blog post or article, often with a bit of AI apologia sprinkled in (for example, marginalizing people as paranoid for not taking AI companies at their word that they aren't aggressively scraping websites in violation of robots.txt, or exfiltrating user data…
If you're not assuming good faith what are you assuming here? What's my motivation? "buttressing commentary on someone else's blog post" That's how link blogs work. I wrote more about my approach to that here: https://simonwillison.net/2024/Dec/22/link-blog/ (And yes, there I go again linking to something I've written from a comment. It's entirely relevant to the point I am making here. That's why I have a blog - so…
Are you really going to insult my and others' intelligence like this? Directly or indirectly, your motivation is money. You already offer monthly subscriptions to your blog, and you're clearly trying to build a monetizable brand for yourself as a leading authority on AI, especially as it pertains to software development.
Re: 2025: The Year in LLMs
#478Earlier quoted context omitted.
The cope + disappointment will be knowing that a large population of HN users will paint a weird alternative reality. There are a multitude of messages about AI that are out there, some are highly detached from reality (on the optimistic and pessimistic side). And then there is the rational middle, professionals who see the obvious value of coding agents in their workflow and use them extensively (or figure out how t…
Please do provide some data for this "obvious value of coding agents". Because right now the only thing obvious is the increase in vulnerabilities, people claiming they are 10x more productive but aren't shipping anything, and some AI hype bloggers that fail to provide any quantitative proof.
Like a lot of things LLM related (Simon Willison's pelican test, researchers + product leaders implementing AI features) I also heavily "vibe" check the capabilities myself on real work tasks. The fact of the matter is I am able to dramatically speed up my work. It may be actually writing production code + helping me review it, or it may be tasks like: write me a script to diagnose this bug I have, or build me a streamlit dashboard to analyze + visualize this ad hoc data instead of me taking 1 hour to make visualizations + munge data in a notebook.
> people claiming they are 10x more productive but aren't shipping anything, and some AI hype bloggers that fail to provide any quantitative proof.
what would satisfy you here? I feel you are strawmanning a bit by picking the most hyperbolic statements and then blanketing that on everyone else.
My workflow is now:
- Write code exclusively with Claude
- Review the code myself + use Claude as a sort of review assistant to help me understand decisions about parts of the code I'm confused about
- Provide feedback to Claude to change / steer it away or towards approaches
- Give up when Claude is hopelessly lost
It takes a bit to get the hang of the right balance but in my personal experience (which I doubt you will take seriously but nevertheless): it is quite the game changer and that's coming from someone who would have laughed at the idea of a $200 coding agent subscription 1 year ago
Re: 2025: The Year in LLMs
#479Earlier quoted context omitted.
The cope + disappointment will be knowing that a large population of HN users will paint a weird alternative reality. There are a multitude of messages about AI that are out there, some are highly detached from reality (on the optimistic and pessimistic side). And then there is the rational middle, professionals who see the obvious value of coding agents in their workflow and use them extensively (or figure out how t…
The nature of my job has always been fighting red tape, process, and stake holders to deploy very small units of code to production. AI really did not help with much of that for me in 2025. I'd imagine I'm not the only one who has a similar situation. Until all those people and processes can be swept away in favor of letting LLMS YOLO everything into production, I don't see how that changes.
Re: 2025: The Year in LLMs
#480Earlier quoted context omitted.
> The emergent phenomenon is that the LLM can separate truth from fiction when you give it a massive amount of data. I don't believe they can. LLMs have no concept of truth. What's likely is that the "truth" for many subjects is represented way more than fiction and when there is objective truth it's consistently represented in similar way. On the other hand there are many variations of "fiction" for the same subject…
They can and we have definitive proof. When we tune LLM models with reinforcement learning the models end up hallucinating less and becoming more reliable. Basically in a nut shell we reward the model when telling the truth and punish it when it’s not. So think of it like this, to create the model we use terabytes of data. Then we do RL which is probably less than one percent of additional data involved in the initia…
I can think of several offhand.
1. The effect was never real, you've just convinced yourself it is because you want it to be, ie you Clever Hans'd yourself.
2. The effect is an artifact of how you measure "truth" and disappears outside that context ("It can be wildly off for certain things")
3. The effect was completely fabricated and is the result of fraud.
If you want to convince me that "I threatened a statistical model with a stick and it somehow got more accurate, therefore it's both intelligent and lying" is true, I need a lot less breathless overcredulity and a lot more "I have actively tried to disprove this result, here's what I found"