Earlier quoted context omitted.
Too soon to tell, give it a billion tokens before we make up our minds
Oh boy, you are far from what it requires, we are probably talking 3B+, but note that this is just codex, obviously codex is also doing automatic adversarial with the regular zoo (gemini-3.1-pro-preview, opus-4.6/4.7, gpt-5.3-codex, minimax-2.7, glm-5.1, mimo-2 (now 2.5) and so-on, you get the gist) :)
OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
81–90 of 174 posts
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#82Earlier quoted context omitted.
Yes Opus 4.7 fast (no reasoning) did a worst job than Sonnet 4.6 high (with reasoning) according to Gemini 3.1 Pro evaluation.
Your table doesn't indicate reasoning vs non-reasoning, or reasoning level
The models not availble on copilot were tested through opencode (max reasoning) and deepseek v4 was tested through Cline (with max reasoning too).
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#83All the AI players definitely seem to be trying to claw more money out of their users at the moment.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#84Earlier quoted context omitted.
The doctor would be responsible for the accuracy of their translation tool, something they can't verify but you expect them to use?
What's the alternative then ? -> You are in China, you go to emergency, nobody speaks your language Move hands ? DeepSeek is better than using hands, even Baidu Translate, ChatGPT or whatever you find. Other solutions are theoretically nice on paper but almost delusional. An imperfect solution is better than no solution. == Similarly, a deaf-person is theorically better with a certified interpreter that can talk with…
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#85Exactly double the cost of GPT 5.4 - $5 per MTok input, $0.50 cached, $30 output. All the AI players definitely seem to be trying to claw more money out of their users at the moment.
30/180 usd on Openrouter. Did I miss something?
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#86Earlier quoted context omitted.
I disagree, it improved enormously especially at staying consistent for long-tasks, I have a task running for 32 days (400M+ tokens) via Codex and that's only since gpt-5.4
That’s actually crazy, what kind of task is that? And is that a recurring kind of task like some analysis, or coding related?
I think users really underestimate the capabilities of "AI" when using the right tooling/combinations of models and procedures (and loops), that's talking with 2 decades of dev behind me, genuinely I'm not on phase with people saying it produces slop of any kind, at this stage, it's mostly the fault of the prompter (or the prompter not having enough tokens to do mass adversarial), but clearly, I can genuinely state that the code produced is overall the SAME quality as I would by being extremely meticulous.
I'm like a bot following 30+ threads concurrently, sometimes it's fun, sometimes it feels like playing casino, sometimes it's boring, but this is truly an insane era if you have the funding for it, obviously we stack many MANY accounts in rotation 24/7, equivalent in API cost by myself is about 100K$+ (a month) but we pay only a fraction of that cost thanks to the plans.
PS: I have 8 monitors in front of me to manage all that (portable monitors stacked together).
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#87Earlier quoted context omitted.
Has that task accomplished anything yet?
I think the OP is in for a rude surprise when the task is “finished”.
“You're really not going to like it," observed Codex.
"Tell us!"
"All right, said Codex. "The answer to your Great Question..."
"Yes...!"
"Is..." said Codex, and paused.
"Yes...!"
"Is..."
"Yes...!!!...?"
"Forty-two," said Codex, with infinite majesty and calm.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#88Earlier quoted context omitted.
I think the OP is in for a rude surprise when the task is “finished”.
It will go somewhat like this: “You're really not going to like it," observed Codex. "Tell us!" "All right, said Codex. "The answer to your Great Question..." "Yes...!" "Is..." said Codex, and paused. "Yes...!" "Is..." "Yes...!!!...?" "Forty-two," said Codex, with infinite majesty and calm.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#89Earlier quoted context omitted.
Seems like benchmark for how good a model is for vibe coding. Your prompt is extremely slim yet you score it on a bunch of features.
Yes, the prompt is slim by design. I might be wrong, but the point was to see what the model can do "on it's own". The eval prompt is quite extensive: https://github.com/guilamu/llms-wordpress-plugin-benchmark/b...
I personally develop with very detailed spec, and I don’t want nothing more and nothing less compared to the spec.
I found 5.4/5.5 much better at following spec while Opus makes some things up, which aligns with your benchmark but that makes 5.4/5.5 better for me while worse for you.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#90Exactly double the cost of GPT 5.4 - $5 per MTok input, $0.50 cached, $30 output. All the AI players definitely seem to be trying to claw more money out of their users at the moment.
(Note, that stops being true at higher reasoning levels, where our observed total cost goes up ~2-3x.)