Live data from Hacker News

The last six months in LLMs in five minutes

simonwillison.net

371–380 of 631 posts

Re: The last six months in LLMs in five minutes

#371
post #364

Earlier quoted context omitted.

Claude Code will automatically "dumb" the TUI down a bit when it can't properly detect certain terminal capabilities, to avoid potential font rendering issues. Likely there are some terminal caps that aren't being properly preserved inside of the sandbox. It's never bothered me since the agent itself works fine.

I think you can fix that by setting an environment variable (regarding the terminal?) but it was a while since I checked. (I was running Claude as a subprocess and had similar issues.) Also this reminds me of a principle I learned from a mentor. "People are visual buyers. If it looks good, people will think the code is good." Unfortunately it doesn't matter whose fault the janky TUI is, people will see that and assoc…

It's more along the lines of: Anyone with an axe to grind will find something to grind it on.

Early stage products will have some rough edges. We've seen that in Docker, Kubernetes, AWS, Azure, LXC, KVM, etc. And people griped and raged about the sheer incompetence of the maintainers and utter lack of quality, but they still used those tools even before the rough edges were polished away and folks finally settled down.

The less one pays for something, the more entitled one feels to whinge and heap on abuse.

I've been down this road so much now that it's no biggie if a few Karens want to blow off steam at my expense. I'm not above exposing their silliness though ;-)

Re: The last six months in LLMs in five minutes

#372

> The coding agents got really good It's since november 2025, the so called "inflection point", that I'm still wondering for who coding agents become "really good". All I observe they got better at tool call and answering questions about big codebases, especially if the question has a vague pattern to search, and they're superuseful for that! For generating production code even with a lot of steering and baby sitting…

I believe by now we know exactly what it's good at and what it's terrible at.

The problem is that our CEO's fear of the future that pushes them to peculiar decisions that objectively make no sense (cf the infamous discussion of the Microsoft employee on Github that couldn't force its agent to do the proper thing).

It's not the first time I witness this kind of discrepancy and probably not the last, I just learned to adapt to it.

Re: The last six months in LLMs in five minutes

#373
post #187

Earlier quoted context omitted.

I have no idea what you're trying to say. If anyone really can vibe code then programming salaries are pretty much guaranteed to come down. The critical question is whether it really is true that anyone can do it, or if it still requires rare skill.

are you a programmer? it 100% requires skill. AI or not. i'm trying to say there's levels to this. if you don't agree then you don't agree. but i can buy commodity tools for any skill and that doesn't make me professional grade at that skill.

Yes I am. The vibe coding I've tried didn't work very well so I agree that still required my skill. But I also don't have access to the latest models and supposedly they're a lot better (see this article for example!).

So is it possible for non-programmers to vibe code if they have the latest models? If not now, what about in a few years?

AI is clearly a different class of tool to something like a welder.

Re: The last six months in LLMs in five minutes

#374

Earlier quoted context omitted.

The UI is working properly. Interfering with Anthropic's UI, or any of the other agent harness' UIs it supports, would be madness incarnate. I also strongly suspect that you'd only taken a cursory glance at the top of the readme prior to passing judgment.

I did not much more than a cursory glance too, but found "./sandbox/create.go", a ~1300 lines long file with so much duplication even within just itself that I stopped counting. Now it was a long time ago I did Go professionally, but I'm also in the camp of "That doesn't really count as high-quality", although I know for a fact you can get quality code out of LLMs, but I don't think that's a good showcase of that.

> I did not much more than a cursory glance too, but found "./sandbox/create.go", a ~1300 lines long file with so much duplication even within just itself that I stopped counting.

Really? What duplication did you actually find? I count a few small ones in buildMounts and ReadPrompt, maybe 20 lines or so, but hardly anything worthy of such an epithet.

Admittedly, the parsing & escaping code and some utility functions could be moved outside to shrink the file, but otherwise I'm having trouble finding issues with the code.

Re: The last six months in LLMs in five minutes

#376
post #367
post #259

Earlier quoted context omitted.

> AI is a tool. Use it appropriately Yes, but no room is made for people who see no use for it. There is a forced-consensus that this technology is useful, which I have to combat against at work. We teach in a very different environment, but your use sounds typical of my colleagues. "I ask it for suggestions and pick one", but nobody seems to wonder about what is lost when we shrink the horizon of what we will teach…

> Yes, but no room is made for people who see no use for it. There is a forced-consensus that this technology is useful, which I have to combat against at work. This is the crux of the issue -- The technology is useful. Using it appropriately is probably the thing that people are ignoring, but you're conflating one and the other in your comment. It is not useful to you in this case, and complain that it is an overall…

I mean only that I see no use for it myself, in my own work. I'm sure there are people working in roles around me who believe they get some use out of AI doing their work for them, and they will have to answer to auditors when they find problems with their work, or when someone is killed.

To me, as a non-techie person, it feels as if people who work in software believe that because their work can be done by AI, everyone else's can, too. Or that this would be better, simply because it proposes a technological solution to human work — it is taken as read that a solution which uses cool sounding computers and data farms is better than one done by humans with a pen and a pad and life experience. They don't have to justify this belief, because the money is on their side.

Re: The last six months in LLMs in five minutes

#378

Earlier quoted context omitted.

Sure give it a go, perhaps it will work better now with frontier models, I haven't tried it in a while (this was a year ago, things have improved since then). I'm not sure what tests for having amazing graphics, gameplay, input, UI, sounds, etc would look like, but it would be interesting to see the results!

okay hold my beer. both claude and codex running now. EDIT: both agents took about 20 minutes. I used that exact prompt in a clean directory for each, and then said "deploy to netlify" - so a total of two prompts. Codex: https://astounding-bavarois-27b5a2.netlify.app Claude: http://strong-hotteok-91dfb0.netlify.app Netlify is having trouble claiming the Claude project, so if you need a password it's "My-Drop-Site" FY…

Nice, very retro (looking at the codex one)!

Claude one doesn't really work (collision detection was the problem I had before too), but fairly close.

Yes when I tried previously I had a few gameplay issues in frogger and I couldn't manage to one-shot this sort of thing at the time (a year ago), so last year definitely saw some good progress at this sort of thing. The asteroids game I was very happy with though, had a very cool retro feel and was wireframe only. Wasn't so keen on the code produced as it had a patchwork feel to it.

Re: The last six months in LLMs in five minutes

#379
post #65

Earlier quoted context omitted.

"That's a higher level of abstraction" No, it's not because it's seen 'anatomy' for Pelicans, Animals - even how it's represented in Animals. If you try to get the AI to actually decompose it and start to 'draw pelicans' in very obscure ways, it will immediately fail. Try to get the AI to draw the pelican form a very odd angle - like underneath, to the right, one wing extended, one wing not ... 0% chance. Precisely b…

> Try to get the AI to draw the pelican form a very odd angle - like underneath, to the right, one wing extended, one wing not ... 0% chance. Proof by existence? https://gist.github.com/nlothian/50241d34a654fcf0caa280d4475... Looks pretty good to me. ChatGPT in "Thinking" model. Edit: I've added the Opus version on the same link.

? That's evidence that it does not work.

Neither of those are from 'under' they both look either front or top?

Imagine yourself under the ducks feet, looking up at an oblique angle - wings as I suggested. The AI won't do that, it has no reference for dimensionality.

Re: The last six months in LLMs in five minutes

#380

Earlier quoted context omitted.

I don't want to offend (it's AI coded anyway :)) but that does not scream "high quality" to me. The headline gif on that repo just paints a terrible picture. It can't draw a box correctly, there's random underscores all over the screen. The UI itself is just incredibly incoherent. I don't even know what I'm looking at. Like, no it doesn't seem like very high quality work... It just seems like a vibe coded tool. Edit:…

Take it up with Anthropic. It's actually their billion-dollar TUI product you're commenting on. The problem with being such a naysayer is that you're entirely disconnected from what's going on. You haven't tried an agent like Claude Code and experienced it for yourself, so you don't recognise what it looks like when it's in front of you.

[deleted]
Post reply on HN