Live data from Hacker News

GPT-5.6

openai.com

341–350 of 1001 posts

Re: GPT-5.6

#341
Just used terra ultra for exactly one prompt in codex and it ate through my full 5h window in about 10mns (20$ plan). The results look pretty good though. Luckily I have had my chatGPT subscription for a while and have a bunch of resets available (nice compared to anthropic).

Assuming I take the 5x plan it would give me about an hour of active sessions with terra ultra (maybe ultra is not good value regarding tokens?), not even using Sol yet. Does everyone using codex use the 200$ plan?

I normally use the 100$ anthropic plan and barely ever reach the usage limit.

Re: GPT-5.6

#342
post #212

Earlier quoted context omitted.

Agreed. GPT 5.5 will come up with more straightforward solutions with far fewer tokens than Claude. Also, the usage limits are much more generous for Codex than Claude Code for the same monthly plan.

Last time I used Codex it would make loads of assumptions, often quite big ones, without asking. Did they fix that, as that for me was what actually made codex worse.

Did you use plan mode?

Re: GPT-5.6

#343
post #87
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

There has been a lot of chatter ever since the Mythos scores had been release that SWEbench pro had major contamination and that Mythos had memorized many questions that lacked the context to be solvable on their own. And now with OpenAI saying a large number of the questions are broken, I think it's worth taking that single outlier benchmark with some salt when the overall trend is that 5.6 is very competitive with Mythos at about half the price.

Re: GPT-5.6

#344
post #303

Earlier quoted context omitted.

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and r…

> Lead with conclusion. I would presume (perhaps falsely?) that an instruction like this would lead to the model presenting a conclusion not supported by the evidence, and potentially backtracking as it then tries to justify said conclusion. Yes, if deliberation happens, the model should figure out what it wants to say during that phase; but if you're using auto mode, the model is not going to be doing any deliberati…

Why would auto mode turn off thinking?

Re: GPT-5.6

#345

Just used terra ultra for exactly one prompt in codex and it ate through my full 5h window in about 10mns (20$ plan). The results look pretty good though. Luckily I have had my chatGPT subscription for a while and have a bunch of resets available (nice compared to anthropic). Assuming I take the 5x plan it would give me about an hour of active sessions with terra ultra (maybe ultra is not good value regarding tokens?…

Do you know if you used sol/terra/luna?

Re: GPT-5.6

#346

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

“Bro” spent most of his career in the wilderness because everybody thought ML/NN/etc were a dead end.

I’d not wager against him having at one one more break though architecture before he retires.

Re: GPT-5.6

#347
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I had great results combining the two. If you (or your employer) can afford then you can ping-pong the models in the plan phase (not really ping-pong as humans should get a say too) and then let one implement and the other review. I got better results working this way than just to stick to a single model.

Re: GPT-5.6

#348
I think the most interesting part of this is that OpenAI is going way easier on the classifiers than Anthropic. They explicitly state that many defensive cybersecurity uses are supported and implicitly criticize Anthropic's stance on Fable's uses by saying that overblocking cyber requests is itself a major security risk as more AI models continue to advance in intelligence. I have so many questions as to what is going on on a game theoretic level in the AI space in the past two months, it seems like multiple actors have realized their incentives are really quite different than they originally thought.

Re: GPT-5.6

#349
post #260

>> approximately 700,000 A100e GPU hours of black-box automated red teaming Amusing that they use A100e as the reference point to sound impressive. Different ways you could make that conversion, but based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point), that's something like 200hr on a GB300 NVL72 rack. Not nothing either, but far less astounding sounding than 700k hrs.

Wait, what do you mean? 700k A100e hours are equal to 200 hours of a GB300 NVL72 rack? One GB300 NVL72, 72-GPU rack has equal processing power to 3500 A100e GPUs?

> based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point)

The A100 doesn't have hardware FP4, and you'd be running a quantized model with some accuracy loss but unless this was natively trained on FP4*

* to add another layer, they own the model and could apply tons of post-training techniques to reduce that accuracy loss and probably already do

Re: GPT-5.6

#350
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

like others said in the thread: much less drama and i'll add much less attitude from the company and the models, overall i'm having much calmer experience with codex, hope it stays that way
Post reply on HN