Live data from Hacker News

The last six months in LLMs in five minutes

simonwillison.net

301–310 of 631 posts

Re: The last six months in LLMs in five minutes

#301

> The coding agents got really good It's since november 2025, the so called "inflection point", that I'm still wondering for who coding agents become "really good". All I observe they got better at tool call and answering questions about big codebases, especially if the question has a vague pattern to search, and they're superuseful for that! For generating production code even with a lot of steering and baby sitting…

> It's since november 2025, the so called "inflection point", that I'm still wondering for who coding agents become "really good".

The answer is "for lots of people, but not you".

You're doing a vague impression of being fair and even-handed, arguing for non-polarization, but underlying everything you're saying is an obvious attitude of poralizing superiority: That _your_ personal experience with AI is the real truth. That _your_ codebase is more intricate and more challenging than what other people are doing. That everyone else is being led by a "marketing hype train".

Re: The last six months in LLMs in five minutes

#302

Earlier quoted context omitted.

You do realize that you're complaining about the Claude Code TUI, right? That's not what this product is; merely a tool it uses.

So why has your tool completely broken the Claude Code UI then? Can't you see in the gif? It's completely broken. My Claude doesn't look like that. Neither does anyone else's.

Claude Code will automatically "dumb" the TUI down a bit when it can't properly detect certain terminal capabilities, to avoid potential font rendering issues.

Likely there are some terminal caps that aren't being properly preserved inside of the sandbox. It's never bothered me since the agent itself works fine.

Re: The last six months in LLMs in five minutes

#303

Earlier quoted context omitted.

Take it up with Anthropic. It's actually their billion-dollar TUI product you're commenting on. The problem with being such a naysayer is that you're entirely disconnected from what's going on. You haven't tried an agent like Claude Code and experienced it for yourself, so you don't recognise what it looks like when it's in front of you.

I have tried Claude code. It doesn't look like that! I don't know what the project is. All I see is a TUI that looks completely broken. Go and use Claude Code right now. Does it look like that? Random underscores all over the page. No it doesn't.

It can look like that in certain conditions. The question is why are you so eager to give critique on unrelated work, appearing in a demo screencap, to someone who didn't produce it?

Re: The last six months in LLMs in five minutes

#304

Earlier quoted context omitted.

You claim "very high quality" but can't even get the basic UI working properly. You wrap tmux and a container in 2k lines of code and claim quality, I think the comment above was aimed at this claim.

The UI is working properly. Interfering with Anthropic's UI, or any of the other agent harness' UIs it supports, would be madness incarnate. I also strongly suspect that you'd only taken a cursory glance at the top of the readme prior to passing judgment.

[deleted]

Re: The last six months in LLMs in five minutes

#305

Earlier quoted context omitted.

The polarization comes from the very disparate coding experiences and output quality that different people find when using these tools. For example, I've had the opposite experience of yours, generating very high quality work using Claude (such as https://github.com/kstenerud/yoloai ). Just in dealing with all the bugs and idiosyncrasies in the technologies I'm using, the agent has been a godsend in discovering and c…

> The polarization comes from the very disparate coding experiences and output quality that different people find when using these tools. Not just when using tools, also when using humans. The frame of reference of what is considered 'production code' differs immensely between organizations, teams and people. The code I get from LLM's is usually much better than what I get from my peers. Maybe not one shot, but after…

> The code I get from LLM's is usually much better than what I get from my peers

Then you should seriously question for who you're working for imo.

> It also isn't lazy.

It is indeed lazy in my experience, as in being overly zealous when creating useless test cases and ignoring the important ones. I don't want it to test a sum I want to know a test that can "guarantee" me that a further change doesn't break existing code. And producing this high quality in tests is HARD, and requires a lot of steering with agents. This culture of tests code coverage is just wrong, the best code base I worked with had code coverage only on the net percent of code that matters, the rest is covered by for static type checking and integration tests

Re: The last six months in LLMs in five minutes

#306
post #26

December 2025 was the breakthrough for me. January Claude was euphoric, ChatGPT was up there. February Gemini cooked for a second there. March amazing. April the big bad nerf. May GPT 5.5 is just pure bliss altough 2x limits temporarily, not sure about Claude it's sort of okay still not as good as it felt before, slowly increasing limits with more compute and rebuilding good will.

I think Opus 4.6 at its peak was the "how can anyone not get that this is good" for me. Then the nerf, and the massive uplift in tokens for 4.7, a model which I find lazy and prone to hallucinate. It's probably time to try GPT5.5. Like many I'm pretty heavily invested in the anthropic ecosystem at this point, which I suppose gives another strong reason to make the switch.

The openclaw ban pushed me over to 5.5 for some daily usage. I feel like Opus and 5.5 are good at very different things. 5.5 can be too literal, and it does not have as much of a ‘creative’ bent whether that’s toward design, UI/UX, interpreting vague instructions, etc. So, in that way, Opus had sort of spoiled me.

On the other hand, this year I’ve been in the habit of using codex as a bug finder / audit layer, where it shines, and I can tell you, Opus makes a lot of mistakes, and as we all know struggles with laziness — and has gotten good at encoding that laziness into the codebase (// Per instructions, pass this test by default) where it can live for a long time. So, Opus had spoiled me, but more with its ability to sketch holistically than its ability to put out perfect codebases.

Upshot - it was good to switch horses for a while, as you mention. Slightly different skill sets there. And I still reach for claude especially for initial design. But right now the daily driver is 5.5 / xhigh fast mode, and it’s very capable.

Re: The last six months in LLMs in five minutes

#307

Earlier quoted context omitted.

That’s really impressive, and slightly worrying for creatives involved in film, animation or modelling.

I wouldn't be that concerned that animation is going anywhere. Both outputs look really off, especially around the feet.

In a serious creative tool you would also want a lot more creative input. At a minimum the ability to steer the animation with skeletons that feed into a control net, or something like that. And the ability to control the look and feel and create much more consistent characters. Both things that exist in good tooling, but both things that create work that will keep animators employed. But it will dramatically reduce the number of animators needed to reach a given level of "good enough".

And looking at the trajectory of the animation industry, I don't think increases in productivity will be used to raise the quality of the animation if the alternative is to just pay fewer animators

Re: The last six months in LLMs in five minutes

#308
post #44

Earlier quoted context omitted.

I remember this very clearly myself. Before opus 4.5, I was doing a lot of hand holding and was coding a lot myself, but I have not written code since that day more or less. I did write some stuff myself just to learn how the enigma encryption machine worked, so wrote myself to learn. But professionally, I stopped coding in November.

How do you justify your salary given that you're just using a tool that any of us could use for $20 an hour in your role?

Never to feed the trolls ... but, how does my carpenter deserve $100 an hour when he is using an electric drill and power saw I can get at Home Deepo for $100 bucks?

Most good developers are not employed because just because they can code well.

What is over is: fizzbuzz and trivial CS algorithm regurgitation as a gate.

Re: The last six months in LLMs in five minutes

#309

Earlier quoted context omitted.

I have tried Claude code. It doesn't look like that! I don't know what the project is. All I see is a TUI that looks completely broken. Go and use Claude Code right now. Does it look like that? Random underscores all over the page. No it doesn't.

It can look like that in certain conditions. The question is why are you so eager to give critique on unrelated work, appearing in a demo screencap, to someone who didn't produce it?

I don't know what you're talking about.

His tool wraps Claude and breaks the TUI. What's so hard to understand?

That's valid critique. What world have I woke up in today?

Re: The last six months in LLMs in five minutes

#310

Last 6 months is humanity losing control of LLMs. - Memory market cornering which mitigated the adoption of local AI despite great open model being released. - Fast penetration of IP exfiltrating tools in companies world-wide. - Developers producing more code that they can read. - Autonomous agents killing Open Source by siphoning the attention economy - Autonomous agents destroyed online communities (including HN) -…

[flagged]
Post reply on HN