5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/ crushes with a 77.3% in Terminal Bench. The shortest lived lead in less than 35 minutes. What a time to be alive!
The lack of broad benchmark reports in this makes me curious: Has OpenAI reverted to benchmaxxing? Looking forward to hearing opinions once we all try both of these out
Claude Opus 4.6
291–300 of 1001 posts
Re: Claude Opus 4.6
#292Re: Claude Opus 4.6
#293The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
[flagged]
Re: Claude Opus 4.6
#294Earlier quoted context omitted.
> A year or more ago, I read that both Anthropic and OpenAI were losing money on every single request even for their paid subscribers This gets repeated everywhere but I don't think it's true. The company is unprofitable overall, but I don't see any reason to believe that their per-token inference costs are below the marginal cost of computing those tokens. It is true that the company is unprofitable overall when you…
The reports I remember show that they're profitable per-model, but overlap R&D so that the company is negative overall. And therefore will turn a massive profit if they stop making new models.
Re: Claude Opus 4.6
#295Earlier quoted context omitted.
One aspect of this is that apparently most people can't draw a bicycle much better than this: they get the elements of the frame wrong, mess up the geometry, etc.
Absolutely. A technically correct bike is very hard to draw in SVG without going overboard in details
Re: Claude Opus 4.6
#296The bicycle frame is a bit wonky but the pelican itself is great: https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
[flagged]
Would you mind sharing which benchmarks you think are useful measures for multimodal reasoning?
Re: Claude Opus 4.6
#297Re: Claude Opus 4.6
#298Earlier quoted context omitted.
The cost per token served has been falling steadily over the past few years across basically all of the providers. OpenAI dropped the price they charged for o3 to 1/5th of what it was in June last year thanks to "engineers optimizing inferencing", and plenty of other providers have found cost savings too. Turns out there was a lot of low-hanging fruit in terms of inference optimization that hadn't been plucked yet. >…
> "engineers optimizing inferencing" are we sure this is not a fancy way of saying quantization?
Re: Claude Opus 4.6
#299Earlier quoted context omitted.
Anthropic has perhaps the most embarrassing status page history I have ever seen. They are famous for downtime. https://status.claude.com/
As opposed to other companies which are smart enough not to report outages.
Re: Claude Opus 4.6
#300This is the first model to which I send my collection of nearly 900 poems and an extremely simple prompt (in Portuguese), and it manages to produce an impeccable analysis of the poems, as a (barely) cohesive whole, which span 15 years. It does not make a single mistake, it identifies neologisms, hidden meaning, 7 distinct poetic phases, recurring themes, fragments/heteronyms, related authors. It has left me completel…
This sounds wayyyy over the top for a mode that released 10 mins ago. At least wait an hour or so before spewing breathless hype.