Earlier quoted context omitted.
'feel' is no more accurate not saying there's a better way but both suck
Speak for yourself. I've been insanely productive with Codex 5.2. With the right scaffolding these models are able to perform serious work at high quality levels.
GPT-5.3-Codex
151–160 of 634 posts
Re: GPT-5.3-Codex
#152Earlier quoted context omitted.
yea but i feel like we are over the hill on benchmaxxing, many times a model has beaten anthropic on a specific bench, but the 'feel' is that it is still not as good at coding
'feel' is no more accurate not saying there's a better way but both suck
Re: GPT-5.3-Codex
#153Earlier quoted context omitted.
I do not trust the AI benchmarks much, they often do not line up with my experience. That said ... I do think Codex 5.2 was the best coding model for more complex tasks, albeit quite slow. So very much looking forward to trying out 5.3.
Yeah, these benchmarks are bogus. Every new model overfits to the latest overhyped benchmark. Someone should take this to a logical extreme and train a tiny model that scores better on a specific benchmark.
It's not just over-fitting to leading benchmarks, there's also too many degrees of freedom in how a model is tested (harness, etc). Until there's standardized documentation enabling independent replication, it's all just benchmarketing .
Re: GPT-5.3-Codex
#154The most important question: Can it do Svelte now?
Re: GPT-5.3-Codex
#155Earlier quoted context omitted.
More importantly, this is the early steps of a model self improving itself. Do we still think we'll have soft take off?
This has already been going on for years. It's just that they were using GPT 4.5 to work on GPT 5. All this announcement mean is that they're confident enough in early GPT 5.3 model output to further refine GPT 5.3 based on initial 5.3. But yes, takeoff will still happen because of this recursive self improvement works, it's just that we're already past the inception point.
Re: GPT-5.3-Codex
#156Pelican seems much worse than the Opus 4.6 one (though the bicycle is more accurate): https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
Re: GPT-5.3-Codex
#157Earlier quoted context omitted.
'feel' is no more accurate not saying there's a better way but both suck
The variety of tasks they can do and will be asked to do is too wide and dissimilar, it will be very hard to have a transversal measurement, at most we will have area specific consensus that model X or Y is better, it is like saying one person is the best coder at everything, that does not exist.
Like can the model take your plan and ask the right questions where there appear to be holes.
How wide of architecture and system design around your language does it understand.
How does it choose to use algorithms available in the language or common libraries.
How often does it hallucinate features/libraries that aren't there.
How does it perform as context get larger.
And that's for one particular language.
Re: GPT-5.3-Codex
#158,,GPT‑5.3-Codex is the first model we classify as High capability for cybersecurity-related tasks under our Preparedness Framework , and the first we’ve directly trained to identify software vulnerabilities. While we don’t have definitive evidence it can automate cyber attacks end-to-end, we’re taking a precautionary approach and deploying our most comprehensive cybersecurity safety stack to date. Our mitigations inc…
I wonder if this will continue to be the case.
Re: GPT-5.3-Codex
#159Earlier quoted context omitted.
yea but i feel like we are over the hill on benchmaxxing, many times a model has beaten anthropic on a specific bench, but the 'feel' is that it is still not as good at coding
'feel' is no more accurate not saying there's a better way but both suck
I’d feel unscientific and broken? Sure maybe why not.
But at the end of the day I’m going to choose what I see with my own two eyes over a number in a table.
Benchmarks are a sometimes useful to. But we are in prime Goodharts Law Territory.
Re: GPT-5.3-Codex
#160I think Anthropic rushed out the release before 10am this morning to avoid having to put in comparisons to GPT-5.3-codex! The new Opus 4.6 scores 65.4 on Terminal-Bench 2.0, up from 64.7 from GPT-5.2-codex. GPT-5.3-codex scores 77.3.
I do not trust the AI benchmarks much, they often do not line up with my experience. That said ... I do think Codex 5.2 was the best coding model for more complex tasks, albeit quite slow. So very much looking forward to trying out 5.3.