Live data from Hacker News

GPT-5.3-Codex

openai.com

81–90 of 634 posts

Re: GPT-5.3-Codex

#83

,,GPT‑5.3-Codex is the first model we classify as High capability for cybersecurity-related tasks under our Preparedness Framework , and the first we’ve directly trained to identify software vulnerabilities. While we don’t have definitive evidence it can automate cyber attacks end-to-end, we’re taking a precautionary approach and deploying our most comprehensive cybersecurity safety stack to date. Our mitigations inc…

Is "high-capability" a stronger or weaker claim than "team of phd-level experts"?

https://www.nbcnews.com/tech/tech-news/openai-releases-chatg...

Re: GPT-5.3-Codex

#84
post #76

Terminal Bench 2.0 | Name | Score | |---------------------|-------| | OpenAI Codex 5.3 | 77.3 | | Anthropic Opus 4.6 | 65.4 |

yea but i feel like we are over the hill on benchmaxxing, many times a model has beaten anthropic on a specific bench, but the 'feel' is that it is still not as good at coding

'feel' is no more accurate

not saying there's a better way but both suck

Re: GPT-5.3-Codex

#85
gpt-5.3-codex isn't available on the API yet. From TFA:

> We are working to safely enable API access soon.

Re: GPT-5.3-Codex

#86
post #74

When 2 multi billion giants advertise same day, it is not competition but rather a sign of struggle and survival. With all the power of the "best artificial intelligence" at your disposition, and a lot of capital also all the brilliant minds, THIS IS WHAT YOU COULD COME UP WITH? Interesting

What's funny is that most of this "progress" is new datasets + post-training shaping the model's behavior (instruction + preference tuning). There is no moat besides that.

"post-training shaping the models behavior" it seems from your wording that you find it not that dramatic. I rather find the fact that RL on novel environments providing steady improvements after base-model an incredibly bullish signal on future AI improvements. I also believe that the capability increase are transferring to other domains (or at least covers enough domains) that it represents a real rise in intelligence in the human sense (when measured in capabilities - not necessarily innate learning ability)

Re: GPT-5.3-Codex

#87
post #23

So can I use this from Opencode? Because Anthropic started to enforce their TOS to kill the Opencode integration

OpenAI models in general, yes - `opencode auth login`, select OpenAI, then ChatGPT Pro/Plus. I just checked and 5.3-codex isn't available in opencode yet, but I assume it will be soon.

Re: GPT-5.3-Codex

#88

I remember when AI labs coordinated so they didn't push major announcements on the same day to avoid cannibalizing each other. Now we have AI labs pushing major announcements within 30 minutes .

Wouldn't that be illegal ? i.e. cartel to collude like that ?

Re: GPT-5.3-Codex

#89

It's so difficult to compare these models because they're not running the same set of evals. I think literally the only eval variant that was reported for both Opus 4.6 and GPT-5.3-Codex is Terminal-Bench 2.0, with Opus 4.6 at 65.4% and GPT-5.3-Codex at 77.3%. None of the other evals were identical, so the numbers for them are not comparable.

I usually wait to see what ArtificialAnalysis says for a direct comparison.

Re: GPT-5.3-Codex

#90

It is absurd to release 5.3-Codex before first releasing 5.3. Also, there is no reason for OpenAI and Anthropic to be trying to one-up each other's releases on the same day. It is hell for the reader.

Because Claude Code is stealing the thunder so OpenAI is focusing on coding now.

Yeah, Claude Code is what everyone is talking about these days and since OpenAI has always been the spending driver being 2nd or 3rd fiddle just isn't acceptable if they're gonna justify it.
Post reply on HN