Live data from Hacker News

GPT-5.6

openai.com

121–130 of 1001 posts

Re: GPT-5.6

#121
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

I'll agree and expand on "weird restrictions" -- I used to check the claude usage graphs multiple times a day to see where I'm at on my weekly budget. With gpt 5.5 I don't think I'm working differently but haven't felt the need to check anything because I think I've hit my limit... once? on some egregious edge case scenario iirc

Re: GPT-5.6

#123
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

omp is really good.

I have one non technical people in my firm using it. One is using it to assist with editing books, basically using it to gather up manuscripts from e-mail / Google Doc etc. submissions, and then switch models between a cheap one and Opus (for actually analysing the manuscript).

The other non-technical person has done really surprising things with it AI, like a long-running GPT 5.5 Pro chat session which is basically her expense tracker - it has an .xlsx file "carried" in the chat, and she just tells ChatGPT (or scans a receipt) whenever she has a new expense, and then prompts it in natural language when she needs a report. I'm looking forward to seeing what she can do with omp.

Re: GPT-5.6

#124

Wow, the "Agents' Last Exam" graph looks unreal!

I mean the y axis is deceptive to make it seem like greater gains since it starts at 30%, when in reality the differences aren't great.

Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable.

What about the graph looked so unreal to you?

Re: GPT-5.6

#125
Looks like a great set of models, but there are about 20 different thinking/model levels here in this family and they are very complex to pick the right one for the task

E.g. for GeneBench Pro, it looks like you would always use GPT-5.6 Sol over Terra/Luna, its pareto optimal.

For Agents Last Exam, you would maybe want Luna, then Terra, then Luna, then Sol as you increasingly budget for tasks.

I feel that there may need to be a new auto mode in many of these cases. It selects the best model and thinking given a particular problem.

Feels like it's going to have to go that way eventually, because here we have about 20 different model and thinking levels you could use, and they're not obvious which ones are right for the given use case.

Re: GPT-5.6

#126
post #29

Earlier quoted context omitted.

Claude Code is a massively bloated agent harness. Try Pi: https://pi.dev/

Pi is so “unbloated” that it’s extra effort to use. You can decide how much work to put into it. I get the trade off. But this is a big jump from CC. I’d recommend some middle ground like opencode.

It's worth trying out OpenCode, then oh-my-pi, and also the commercial harnesses like Codex. (I haven't yet bothered to try Antigravity and have no interest in Gemini-cli now that it's not available except on expensive plans.)

pi is also worth tinkering with, particularly if you have an eye towards automating some things.

Re: GPT-5.6

#127
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I've been using Claude Code, Codex, Gemini (now Antigravity) at the same time for half year now, ever since I dipped my toe into agentic coding. I'd say in general Claude Code and Codex are equally powerful, Gemini is lagging behind.

One thing I appreciate with Codex is, OpenAI nowadays sometimes just gives you quota resets you can bank, so when you use up weekly quota before the week ends, you could just reset the quota, to continue using Codex. I've been much less anxious about Codex quota because of this perk. I just used one reset in the bank yesterday, and still have 3 resets left. Whereas with Claude, when you've used 95% quota 3 days before the week ends, you'd be much more anxious.

On the other hand, Claude Code's /remote-control mechanism is extremely helpful when I am running it in the cloud and wants to monitor it or control it on my phone. Codex currently doesn't support this kind of usage. Codex only allows you to use your phone to connect to a session on your desktop, not in the cloud.

Re: GPT-5.6

#128

Earlier quoted context omitted.

Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.

Anyone know what the deal is with the resets?

They've discovered it's a good marketing strategy. Whenever there's an outage, or a new launch, there's often a reset with it, which helps keep people engaged with OAI / Tibo and reduces churn.

They've also introduced banked resets, which are really clever. If you have a $200/month plan and three banked resets, you're not churning because you will overweight giving up those resets (loss aversion theory).

Re: GPT-5.6

#129
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

With the exception of Fable which is going away anyway, Codex is better especially after the last couple Opus releases. It’s also no longer slower than Claude. You get much more generous usage from the 20x plan. And you get far better uptime. If benchmarks and early tester impressions are accurate, you also get access to Fable level capability at greater speed and lower cost (included in subscription).

> Fable which is going away anyway

$2 says nah. You can't take Fable away in a week where GPT-5.6 and Grok 4.5 launch, if you want to hold on to customers.

Post reply on HN