GPT-5.3-Codex
81–90 of 634 posts
Re: GPT-5.3-Codex
#82Re: GPT-5.3-Codex
#83,,GPT‑5.3-Codex is the first model we classify as High capability for cybersecurity-related tasks under our Preparedness Framework , and the first we’ve directly trained to identify software vulnerabilities. While we don’t have definitive evidence it can automate cyber attacks end-to-end, we’re taking a precautionary approach and deploying our most comprehensive cybersecurity safety stack to date. Our mitigations inc…
https://www.nbcnews.com/tech/tech-news/openai-releases-chatg...
Re: GPT-5.3-Codex
#84Terminal Bench 2.0 | Name | Score | |---------------------|-------| | OpenAI Codex 5.3 | 77.3 | | Anthropic Opus 4.6 | 65.4 |
yea but i feel like we are over the hill on benchmaxxing, many times a model has beaten anthropic on a specific bench, but the 'feel' is that it is still not as good at coding
not saying there's a better way but both suck
Re: GPT-5.3-Codex
#85> We are working to safely enable API access soon.
Re: GPT-5.3-Codex
#86When 2 multi billion giants advertise same day, it is not competition but rather a sign of struggle and survival. With all the power of the "best artificial intelligence" at your disposition, and a lot of capital also all the brilliant minds, THIS IS WHAT YOU COULD COME UP WITH? Interesting
What's funny is that most of this "progress" is new datasets + post-training shaping the model's behavior (instruction + preference tuning). There is no moat besides that.
Re: GPT-5.3-Codex
#87So can I use this from Opencode? Because Anthropic started to enforce their TOS to kill the Opencode integration
Re: GPT-5.3-Codex
#88I remember when AI labs coordinated so they didn't push major announcements on the same day to avoid cannibalizing each other. Now we have AI labs pushing major announcements within 30 minutes .
Re: GPT-5.3-Codex
#89It's so difficult to compare these models because they're not running the same set of evals. I think literally the only eval variant that was reported for both Opus 4.6 and GPT-5.3-Codex is Terminal-Bench 2.0, with Opus 4.6 at 65.4% and GPT-5.3-Codex at 77.3%. None of the other evals were identical, so the numbers for them are not comparable.
Re: GPT-5.3-Codex
#90It is absurd to release 5.3-Codex before first releasing 5.3. Also, there is no reason for OpenAI and Anthropic to be trying to one-up each other's releases on the same day. It is hell for the reader.
Because Claude Code is stealing the thunder so OpenAI is focusing on coding now.