I just ran this through a simple change I’ve asked Sonnet 4 and Opus 4.1, and it fails too. It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I worry everyone is chasing benchmarks to the detriment of general performance. Or the next token weight for the incorrect c…
> It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I don't understand why this kind of thing is useful. Do the thing yourself and move on. For every one problem like this, AI can do 10 better/faster than I can.
Claude Sonnet 4.5
261–270 of 819 posts
Re: Claude Sonnet 4.5
#262Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
I had a complete shocker with all of Claude, GitHub Copilot, and ChatGPT when trying to prototype an iOS app in Swift around 12 months ago. They would all really struggle to generate anything usable, and making any progress was incredibly slow due to all the problems I was running into. This was in stark contrast to my experience with TypeScript/NextJS, Python, and C#. Most of the time output quality for these was at…
Re: Claude Sonnet 4.5
#263ChatGPT even does zip file downloads, packaging up all your files.
Re: Claude Sonnet 4.5
#264Re: Claude Sonnet 4.5
#265Earlier quoted context omitted.
All these are useless metrics. It doesn't say anything meaningful on the quality of your life. I would be more interested in knowing if he can now retire in next 5 years instead of waiting another 15? Or do he now just just get to work for 2 hours and enjoy the remaining 6 hours doing meaningful things apart from staring at a screen.
Not everyone hates their job and gets no satisfaction from it. Some of us relish doing something useful and getting paid for it.
Re: Claude Sonnet 4.5
#266Anecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s de…
Re: Claude Sonnet 4.5
#267Earlier quoted context omitted.
GPT-5 is like the guy on the baseball team that's really good at hitting home runs but can't do basic shit in the outfield. It also consistently gets into drama with the other agents e.g. the other day when I told it we were switching to claude code for executing changes, after badmouthing claude's entirely reasonable and measured analysis it went ahead and decided to `git reset --hard` even after I twice pushed back…
> it went ahead and decided to `git reset --hard` even after I twice pushed back on that idea So this is something I've noticed with GPT (Codex). It really loves to use git. If you have it do something and then later change your mind and ask it to undo the changes it just made, there's a decent chance it's going to revert to the previous git commit, regardless of whether that includes reverting whole chunks of code i…
Re: Claude Sonnet 4.5
#268I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…
Why did you have access to a preview?
Thanks for all your work, Simon! You're my favorite journalist in this space and I really appreciate your tone.
Re: Claude Sonnet 4.5
#269Earlier quoted context omitted.
All these are useless metrics. It doesn't say anything meaningful on the quality of your life. I would be more interested in knowing if he can now retire in next 5 years instead of waiting another 15? Or do he now just just get to work for 2 hours and enjoy the remaining 6 hours doing meaningful things apart from staring at a screen.
Not everyone hates their job and gets no satisfaction from it. Some of us relish doing something useful and getting paid for it.