Live data from Hacker News

GPT-5.6

openai.com

241–250 of 1001 posts

Re: GPT-5.6

#241
prompts -> loops -> slingshots?

Its an extremely capable model. I think the way we need to approach works shifts again. We need to get our harnesses/workflows to let it gather some momentum on the first couple rounds but then we also need to structure it so that it can slingshot and accomplish the long range goal.

Re: GPT-5.6

#242
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

I figured pi itself would be the best harness because it's barebones and you make it what you want. omp is to pi what doom is to emacs is what lazyvim is to neovim.

Re: GPT-5.6

#243

I wish model launches were like proper product releases it's impossible to _try_ it out on release! it's not on their codex subscription, or the web/mobile chatgpt interfaces, or aws bedrock, etc. I just cant find a working endpoint with the latest model after they announce

The announcement says they're rolling it out over the next 24 hours or so. I think it's reasonable to do a slow-roll-out release for one of the most used products on the internet.

Re: GPT-5.6

#244
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

> does it really matter anymore?

They're different models with different philosophies behind them. This is anecdotal with a user group of 1, but in my experience:

Claude has a stronger personality and is more creative. If you give it vague instructions, it's better at filling in the blanks with reasonable ideas.

GPT-5.5 is better at following instructions. If you know exactly what you want, it will do it without going off the rails. It's also less likely to imply that you're dumb, but I don't really care about that. Some people do.

Re: GPT-5.6

#246
post #234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

I mean, theoretically you can solve every finitary problem with a brute force solution...

Richard Sutton specifically states that the search has to be smart. We know that the brain uses recurrent connections and is shallow. I think a lot more money has to go into architecture. Feed Forward transformers can only scale so far

Re: GPT-5.6

#247
post #226

Earlier quoted context omitted.

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.

Are you joking? They spend billions of dollars training LLMs to get a 7.8% on arc agi 3 whereas DINO models are near sota in image classification, provide meaningful embeddings to the point where image segmentation is just PCA. The spend on DINO cannot be more than five million (correct me if I'm wrong)

JEPA is just getting started

Re: GPT-5.6

#249
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

The answer is it depends. Claude's generally better at frontend and debugging tasks, while Codex is stronger at backend features and exploratory work. They have very different coding styles and thus very different strengths.

Any actual data backing this up? Or is this just your personal experience?

Re: GPT-5.6

#250
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

Even less drama with open models like GLM.
Post reply on HN