Earlier quoted context omitted.
What happened to you?
AI fried brains, unfortunately.
GPT-5.3-Codex
61–70 of 634 posts
Re: GPT-5.3-Codex
#62https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
Re: GPT-5.3-Codex
#63I think Anthropic rushed out the release before 10am this morning to avoid having to put in comparisons to GPT-5.3-codex! The new Opus 4.6 scores 65.4 on Terminal-Bench 2.0, up from 64.7 from GPT-5.2-codex. GPT-5.3-codex scores 77.3.
I do not trust the AI benchmarks much, they often do not line up with my experience. That said ... I do think Codex 5.2 was the best coding model for more complex tasks, albeit quite slow. So very much looking forward to trying out 5.3.
I use 5.2 Codex for the entire task, then ask Opus 4.5 at the end to double check the work. It's nice to have another frontier model's opinion and ask it to spot any potential issues.
Looking forward to trying 5.3.
Re: GPT-5.3-Codex
#64That was fast! I really do wonder whats the chain here. Did Sam see the Opus announcement and DM someone a minute later?
Re: GPT-5.3-Codex
#65It's so difficult to compare these models because they're not running the same set of evals. I think literally the only eval variant that was reported for both Opus 4.6 and GPT-5.3-Codex is Terminal-Bench 2.0, with Opus 4.6 at 65.4% and GPT-5.3-Codex at 77.3%. None of the other evals were identical, so the numbers for them are not comparable.
It's better on a benchmark I've never heard of!? That is groundbreaking, I'm switching immediately!
Looking at the Opus model card I see that they also have by far the highest score for a single model on ARC-AGI-2. I wonder why they didn't advertise that.
Re: GPT-5.3-Codex
#66Re: GPT-5.3-Codex
#67So can I use this from Opencode? Because Anthropic started to enforce their TOS to kill the Opencode integration
What you can't do is pretend opencode is claude code to make use of that specific claude code subscription.
Re: GPT-5.3-Codex
#68Earlier quoted context omitted.
How so?
Its kind of a suck up that more or less confirms the beef stories that were floating around this past week. In case you missed it. For example: Nvidia's $100 billion OpenAI deal has seemingly vanished - Ars Technica https://arstechnica.com/information-technology/2026/02/five-... Specifically this paragraph is what I find hilarious. > According to the report, the issue became apparent in OpenAI’s Codex, an AI code-gen…
They should design their own hardware, then. Somehow the other companies seem to be able to produce fast-enough models.
Re: GPT-5.3-Codex
#69Re: GPT-5.3-Codex
#70Earlier quoted context omitted.
I do not trust the AI benchmarks much, they often do not line up with my experience. That said ... I do think Codex 5.2 was the best coding model for more complex tasks, albeit quite slow. So very much looking forward to trying out 5.3.
5.2 Codex became my default coding model. It “feels” smarter than Opus 4.5. I use 5.2 Codex for the entire task, then ask Opus 4.5 at the end to double check the work. It's nice to have another frontier model's opinion and ask it to spot any potential issues. Looking forward to trying 5.3.