Live data from Hacker News

The last six months in LLMs in five minutes

simonwillison.net

201–210 of 631 posts

Re: The last six months in LLMs in five minutes

#202
post #88

Earlier quoted context omitted.

I remember this very clearly myself. Before opus 4.5, I was doing a lot of hand holding and was coding a lot myself, but I have not written code since that day more or less. I did write some stuff myself just to learn how the enigma encryption machine worked, so wrote myself to learn. But professionally, I stopped coding in November.

It is sad. I like programming, if I couldn't do it and had to write text (which I do hate, I'm not a writer) it would be make quite a sad world.

Nothing stopping you from doing that in a post-LLM world

Re: The last six months in LLMs in five minutes

#204

Earlier quoted context omitted.

Claude in Office was a tipping point for nontechnical folks around me. Everyone’s slides decks are immaculate now. Finance isn’t needing nearly as much BI help. It’s pretty impressive.

I find it really troubling finance are relying on LLMs (word generators!) for financial analysis - I mean I guess it means there will never be any annoying gaps in the data.

Depends on how it’s done.

I use it a lot now for knocking up grafana charts etc. It’s not so much that the LLM is feeding the numbers through. You can still use real tools to analyse and summarise the numbers, it’s just much quicker at driving them.

As ever with data analysis, two things will continue to be true. Real insights come from spotting something that looks off and digging into it deeper. Secondly, it’s really easy to connect data in a misleading way.

I’ve had a Claude analysis handed to me this morning including a summary list of actions we’re going to take next which falls into this very trap.

The insights you’ll get from your data will only be as deep as the curiosity of the person at the helm.

Re: The last six months in LLMs in five minutes

#205
post #172

Earlier quoted context omitted.

How do you justify your salary given that you're just using OSS compiler/editor any of us could use for free in your role ? AI just changed how I edit code - I still see coworkers (senior developers) failing with Claude/Codex and get stuck when there are trivial solutions if you understand the full problem space. Right now AI is just a productivity tool.

Can you share how you use it to edit code? I‘ve seen a couple approaches, curious what you are doing: 1. Spec -> plan -> code (all agent driven, maybe with grill-me or ultraplan) 2. Handwritten spec -> agent driven plan -> agent driven code 3. Agent driven spec -> vibed code -> Fix by handholding until ok-ish 4. Vibed throwaway prototypes -> extract useful patterns -> rewrite with handholding 5. Generate file structu…

Usually I describe the problem, explore a bit with LLM iteratively. Then I switch to creating a plan when I have enough insight (and the LLM has it in context/same session as exploration), specifying all the things I'm trying to accomplish.

Then I just iterate with LLM - I let it start writing stuff in YOLO mode and check on what it's doing in the code steering it in the direction I want.

Usually the code LLM generates will work but is kind of garbage - but I can easily steer it towards better implementations.

Sometimes using an LLM is theoretically slower than hand-rolling - if I just sat down and focused I could outperform the iteration and the waiting, especially considering how stupid agents are at running expensive builds/test suites (with a bunch of explicit instructions in skills/claude/agents.md). But the practical improvement of going with LLM is that you have a bunch of thinking traces saved as a part of your iteration proces - it's really easy to get back into flow. This is a huge productivity win for me given how many interruptions I have in my work day. Like so many people like to point out - writing code ends up being less and less of your time as you level up in your career.

Re: The last six months in LLMs in five minutes

#206
post #65

Earlier quoted context omitted.

"That's a higher level of abstraction" No, it's not because it's seen 'anatomy' for Pelicans, Animals - even how it's represented in Animals. If you try to get the AI to actually decompose it and start to 'draw pelicans' in very obscure ways, it will immediately fail. Try to get the AI to draw the pelican form a very odd angle - like underneath, to the right, one wing extended, one wing not ... 0% chance. Precisely b…

> Try to get the AI to draw the pelican form a very odd angle - like underneath, to the right, one wing extended, one wing not ... 0% chance. Proof by existence? https://gist.github.com/nlothian/50241d34a654fcf0caa280d4475... Looks pretty good to me. ChatGPT in "Thinking" model. Edit: I've added the Opus version on the same link.

Those are just awful compared to the side view of a pelican on a bike.

Re: The last six months in LLMs in five minutes

#207

I asked Gemini for a video of 'pelican riding a unicycle in hyde park' - I was blown away by the output: https://gemini.google.com/share/55e250c99693

I'm surprised by Grok as well:

https://grok.com/imagine/post/8d1eab88-737f-4d46-ba92-9b6502...

Interesting that it does better at making the pelican peddle in the video generation than in image generation.

Re: The last six months in LLMs in five minutes

#208
post #26

December 2025 was the breakthrough for me. January Claude was euphoric, ChatGPT was up there. February Gemini cooked for a second there. March amazing. April the big bad nerf. May GPT 5.5 is just pure bliss altough 2x limits temporarily, not sure about Claude it's sort of okay still not as good as it felt before, slowly increasing limits with more compute and rebuilding good will.

I think Opus 4.6 at its peak was the "how can anyone not get that this is good" for me. Then the nerf, and the massive uplift in tokens for 4.7, a model which I find lazy and prone to hallucinate. It's probably time to try GPT5.5. Like many I'm pretty heavily invested in the anthropic ecosystem at this point, which I suppose gives another strong reason to make the switch.

I only used Claude first time in April, previously only ChatGPT and Gemini. And I struggle to see what the hype is all about - yes it seems a tiny bit smarter than the pack, but on the 20$ subscription it runs out of tokens in 5-20 minutes, and then you need to wait 3-4h.

ChatGPT 5.5 seems capable, although a bit stingy with “thinking” compared to earlier models, and I never run into session limits.

Re: The last six months in LLMs in five minutes

#209
> The coding agents got really good

It's since november 2025, the so called "inflection point", that I'm still wondering for who coding agents become "really good".

All I observe they got better at tool call and answering questions about big codebases, especially if the question has a vague pattern to search, and they're superuseful for that! For generating production code even with a lot of steering and baby sitting?

Absolutely not, not quite there not even close in my experience.

But we should stop talking about 1s and 0s, especially with marketing hype trains, there exist a gradient of capabalities that agents have that really depends on the intricacies of the codebase you're working on, I think everyone has yet to discover how to better apply these tools in their day to day work.

But that totally collides with the current narrative, that flattens out our work to be always the same and that can be automated easily in each case, it's not!

That's why the debate is so polizered imo, there isn't a shared experience

Re: The last six months in LLMs in five minutes

#210

Does this guy have a "publish to front page of HN" button on his blog editor?

HN has a mechanism that causes popular blogs to stay popular.

It's a winner-takes-all karma prize for being first to post the article.

This causes a rush of people to post.

HN has a mechanism by which duplicate submissions count as upvotes toward the first submission.

This is a positive feedback for the desire to be first, which increases duplicate submissions and in turn the karma reward.

This effect means that good blogs stay well upvoted. This isn't altogether a bad thing, but it does mean some blogs require a string of poorly received posts before that effect wears off and people no longer rush to be first.

One way to fix this would be to attribute all karma to user simonw himself ( and do similar where attribution to an HN user is known. )

Post reply on HN