Its an extremely capable model. I think the way we need to approach works shifts again. We need to get our harnesses/workflows to let it gather some momentum on the first couple rounds but then we also need to structure it so that it can slingshot and accomplish the long range goal.
GPT-5.6
241–250 of 1001 posts
Re: GPT-5.6
#242Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right
Re: GPT-5.6
#243I wish model launches were like proper product releases it's impossible to _try_ it out on release! it's not on their codex subscription, or the web/mobile chatgpt interfaces, or aws bedrock, etc. I just cant find a working endpoint with the latest model after they announce
Re: GPT-5.6
#244Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
They're different models with different philosophies behind them. This is anecdotal with a user group of 1, but in my experience:
Claude has a stronger personality and is more creative. If you give it vague instructions, it's better at filling in the blanks with reasonable ideas.
GPT-5.5 is better at following instructions. If you know exactly what you want, it will do it without going off the rails. It's also less likely to imply that you're dumb, but I don't really care about that. Some people do.
Re: GPT-5.6
#2458% on ARC-AGI-3, they actually got some traction going...
before today all the contestants were capped at $10k
Re: GPT-5.6
#246GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.
Richard Sutton specifically states that the search has to be smart. We know that the brain uses recurrent connections and is shallow. I think a lot more money has to go into architecture. Feed Forward transformers can only scale so far
Re: GPT-5.6
#247Earlier quoted context omitted.
Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
JEPA is just getting started
Re: GPT-5.6
#248Re: GPT-5.6
#249Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
The answer is it depends. Claude's generally better at frontend and debugging tasks, while Codex is stronger at backend features and exploratory work. They have very different coding styles and thus very different strengths.
Re: GPT-5.6
#250Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.