Live data from Hacker News

GPT-5.6

openai.com

311–320 of 1001 posts

Re: GPT-5.6

#311
post #97

Earlier quoted context omitted.

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

Nudged by this thread, I've decided to switch from Claude to Codex for a bit to see what happens. But...I immediately became lost in their marketing vortex of confusion on plans and pricing. Anyone care to tell me which plan I should be using? On the other side I use the $100 Claude Code plan. We actually have a "Business" ChatGPT subscription already, which seems to be $50/mo/seat. OpenAI's web site offers a set of…

Test-drive it with an individual Pro account (5x or 20x) for a month. Download the Codex CLI client from https://github.com/openai/codex and auth it in the browser via the URL it provides. Set the model to 5.6-Sol and effort to max.

Re: GPT-5.6

#312

I wish model launches were like proper product releases it's impossible to _try_ it out on release! it's not on their codex subscription, or the web/mobile chatgpt interfaces, or aws bedrock, etc. I just cant find a working endpoint with the latest model after they announce

GPT-5.6 Terra just showed up in Codex for me.

Re: GPT-5.6

#313
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

The harness is so much better than cc which is a buggy mess. Gpt is also way faster than Claude. I’ve been using gpt for a while now and I know a lot of people that swapped away from Anthropic for multiple reasons. However - fable still seems to be the best coding agent, it’s just slow and the harness sucks. So I still use it in some rare cases like to review codex. I’m hoping 5.6 lets me drop it entirely.

Re: GPT-5.6

#314

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. When has this ever not been the case? I don't think this is a GPT 5.6 specialty!

There was a fad a while back of building insanely long prompts - tens of thousands of tokens - including having models write prompts for themselves. I always thought it was counterproductive, especially if you're going to use the prompt more than a couple of times. (That said, the e.g. Claude Code system prompt is insanely long, so if you genuinely have a lot of information to provide maybe it's beneficial. Like, shorter is better, but you don't want to be under-specified.)

Re: GPT-5.6

#315
post #142
post #97

Earlier quoted context omitted.

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.

> Codex is faster but you always have to correct it because it got something wrong

this has been my experience with Codex as well, and I have to fix its mistakes every single time. But recently, I literally threw away three hours of work because it kept adding hundreds of lines to my code base. When I restarted the entire work using Fable and Opus, it was like night and day.

Re: GPT-5.6

#316

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

[flagged]

Re: GPT-5.6

#317
Huh, a good alternative just as anthropic's 50% weekly subscription subsidy is ending this weekend. Time to see if it's benchmaxxed or actually a strong leap over GPT5.5.

They also seem to really not care about alignment, or care about it in the wrong way. It's entirely missing in the blogpost and there are some concerning bits in the model card, seemingly treating CoT controllability as something to be "investigated" rather than the warning sign it's supposed to be.

Re: GPT-5.6

#318
post #37

"GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back." This one is really promising, as i…

Computer-use is a big limitation that my 2015 Macbook Pro cannot handle. I find the Codex cli says it looks at the end output artifact but so often it fails to refine it into acceptable form. If it could use my computer screen and visual inputs for review, it might be able to actually design documents/powerpoints/etc. I'm juicing everything I can out of the 11 year old laptop and I'm honestly impressed at what it can still do.

Re: GPT-5.6

#319
post #9
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Use a harness that doesn't lock you into a moat, like OpenCode.

You can use Codex with any endpoint compatible with OpenAI Response API[1], like llama.cpp.

[1]: https://unsloth.ai/docs/basics/codex

Re: GPT-5.6

#320
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

IME it entirely depends on your work. I find myself using both daily for different things.

Codex with GPT 5.5 is much better at general SWE tasks but Claude Code with Opus is far better at complex reasoning tasks like reading and summarizing research papers, replicating experiments, identifying research gaps and proposing interesting follow ups.

Post reply on HN