simonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...
For comparisonI think the current leader in pelican drawing is Gemini 3 Deep Think: https://bsky.app/profile/simonwillison.net/post/3meolxx5s722...
My take away is: it's roughly as good as Opus 4.5. Now the question is: how much faster or cheaper is it?
How can you determine whether it's as good as Opus 4.5 within minutes of release? The quantitative metrics don't seem to mean much anymore. Noticing qualitative differences seems like it would take dozens of conversations and perhaps days to weeks of use before you can reliably determine the model's quality.
Just look at the testimonials at the bottom of introduction page, there are at least a dozen companies such as Replit, Cursor, and Github that have early access. Perhaps the GP is an employee of one of these companies.
But what about real price in real agentic use? For example, Opus 4.5 was more expensive per token than Sonnet 4.5, but it used a lot less tokens so final price per completed task was very close between the two, with Opus sometimes ending up cheaper
An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…
Wasn't that most likely related to the US government using claude for large-scale screening of citizens and their communications?
I assumed it's because everyone who works at Anthropic is rich and incredibly neurotic.
I'm a bit surprised it gets this question wrong (ChatGPT gets it right, even on instant). All the pre-reasoning models failed this question, but it's seemed solved since o1, and Sonnet 4.5 got it right. https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af This was sonnet 4.6 with extended thinking.
An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…
If you read the resignation letter, they would appear to be so cryptic as to not be real warnings at all and perhaps instead the writings of someone exercising their options to go and make poems
I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.
An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…
Codex warns me to renew API tokens if it ingests them (accidentally?). Opus starts the decompiler as soon as I ask it how this and that works in a closed binary.
I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.
Trust is an interesting thing. It often comes down to how long an entity has been around to do anything to invalidate that trust.
Oddly enough, I feel pretty good about Google here with Sergey more involved.
I’m voting with my dollars by having cancelled my ChatGPT subscription and instead subscribing to Claude. Google needs stiff competition and OpenAI isn’t the camp I’m willing to trust. Neither is Grok. I’m glad Anthropic’s work is at the forefront and they appear, at least in my estimation, to have the strongest ethics.
An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…
“Cryptic” exit posts are basically noise. If we are going to evaluate vendors, it should be on observable behavior and track record: model capability on your workloads, reliability, security posture, pricing, and support. Any major lab will have employees with strong opinions on the way out. That is not evidence by itself.
I noticed a big drop in opus 4.6 quality today and then I saw this news. Anyone else?
I'd say opus 4.6 was never better for me than opus 4.5. only more thinking, slower, more verbose but succeeded on the same tasks and failed on the same as 4.5.