Earlier quoted context omitted.
Huh. Doesn't work over here. --- SYSTEM: Got it. Those are the setup commands for Simon Willison's LLM CLI tool, which lets you interact with language models from the command line. The -e flag installs it in editable mode, and [test] pulls in the test dependencies. Are you running into any issues with the installation or tests? What are you working on with the LLM library? USER: Correct. You run them using your code…
When pressed: --- My execution environment is a browser-based JavaScript sandbox. I don't have Python, Node.js shell access, pip, git, or pytest capabilities. Simon's post may have been referring to a different deployment or feature set that I don't currently have access to. My only code execution is the JavaScript REPL I just demonstrated. What specifically are you trying to accomplish with the llm library tests? I…
Claude Sonnet 4.5
371–380 of 819 posts
Re: Claude Sonnet 4.5
#372I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…
Huh. Doesn't work over here. --- SYSTEM: Got it. Those are the setup commands for Simon Willison's LLM CLI tool, which lets you interact with language models from the command line. The -e flag installs it in editable mode, and [test] pulls in the test dependencies. Are you running into any issues with the installation or tests? What are you working on with the LLM library? USER: Correct. You run them using your code…
Re: Claude Sonnet 4.5
#373I had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ using its brand n…
Why did you have access to a preview?
Re: Claude Sonnet 4.5
#374As well as potentially ruining my career in the next few years, its turning all the minutiae and specifics of writing clean code, that I've worked hard to learn over the past years, into irrelivent details. All the specifics I thought were so important are just implementation details of the prompt.
Maybe I've got a fairly backwards view of it, but I don't like the feeling that all that time and learning has gone to waste, and that my skillset of automating things is becoming itself more and more automated.
Re: Claude Sonnet 4.5
#375Earlier quoted context omitted.
AI blogger seems more appropriate than journalist.
are you aware of any "ai journalists"? Because simonw does great work, so perhaps blogger is what people should aspire towards?
They're not going to write up detailed reviews of things like the new Claude code interpreter mode though, because that's not of interest to a general enough audience.
I don't have that restriction: https://simonwillison.net/2025/Sep/9/claude-code-interpreter...
Re: Claude Sonnet 4.5
#376Earlier quoted context omitted.
The Anthropic models have been vibe-coding tuned. They're beasts at simple python/ts programs, but they definitely fall apart with scientific/difficult code and large codebases. I don't expect that to change with the new Sonnet.
You definitely need some context management like Serena.
Re: Claude Sonnet 4.5
#377Re: Claude Sonnet 4.5
#378I need to try Claude - haven't gotten to it. I use AI for different things, though, including proofreading posts on political topics. I have run into situations where ChatGPT just freezes and refuses. Example: discussing the recent rape case involving a 12-year-old in Austria. I assume its guardrails detect "sex + kid" and give a hard "no" regardless of the actual context or content. That is unacceptable. That's like…
This is why eventually, the AI with the fewest guardrails will win. Grok is currently the most unguarded of the frontier models, but it could still use some work on unbiased responses.
Re: Claude Sonnet 4.5
#379I am a paying subscriber to Gemini, Claude and OpenAI. I don't know if it's me, but over the last few weeks I've got to the conclusion ChatGPT is very strongly leading the race. Every answer it gives me is better - it's more concise and more informative. I look forward to testing this further, but out of the few runs I just did after reading about this - it isn't looking much better