GPT-5.6
271–280 of 1001 posts
Re: GPT-5.6
#272"GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
I am on Plus subscription and see Terra and Luna in Codex, but no sign of Sol. Will it be available only on Pro plans?
UPD from announcement: "The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
Re: GPT-5.6
#273Earlier quoted context omitted.
Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…
I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.
I don't like OpenAI as a company, but they appear to have QA, and that is probably enough to get me to switch.
Re: GPT-5.6
#274Earlier quoted context omitted.
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
Are you joking? They spend billions of dollars training LLMs to get a 7.8% on arc agi 3 whereas DINO models are near sota in image classification, provide meaningful embeddings to the point where image segmentation is just PCA. The spend on DINO cannot be more than five million (correct me if I'm wrong) JEPA is just getting started
Re: GPT-5.6
#275Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…
If anything the online optics have been bad for Anthropic for the last half year. OpenAI doesn't have optics issues, from my point of view they simply have the issue that they are the least trustworthy player at the frontier. The way they pivoted from their original mission is truly breathtaking, especially coming in gloatingly to take the government contract when Anthropic got kicked out for insisting the government does not use their systems for mass surveillance or autonomous weapons systems. You understand what that means, right? OpenAI models are now actively used/developed for mass surveilance and/or autonomous weapons systems.
I know there are plenty here who seem to value their own ability to use these models cheaply above all other considerations. Then OpenAI is a great choice, and much less restrictive than Anthropic. But their problem is not on the optics. It's on the substance.
Re: GPT-5.6
#276Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
I can't tell the difference between Fable and GPT 5.5. I tried Fable while it was in trial $20 mode, used up my whole quota, and it was great, but as soon as I went back to GPT 5.5, everything was the same. But what I love about Openai is that they still let you hook OTHER harnesses up to a subscription. My Pi setup has been built up for a few months now into exactly what I want and moving over to CC or even Codex is…
Re: GPT-5.6
#277GPT 5.5 has a tendency to write English calques and non-idiomatic prose in other languages. Although that can be somewhat tamed with detailed instructions and a corpus of confusing terms, the model’s output often reads like a literal translation rather than native prose. Since I notice these issues most clearly in languages I know well, it makes me reluctant to trust the model’s output in languages in which I’m less proficient.
Ironically, ChatGPT began as a simple text-generation tool, but much of its offerings and benchmarks now focus on coding and agentic workflows, while leaving behind what made it notable in the first place.
Re: GPT-5.6
#278>> approximately 700,000 A100e GPU hours of black-box automated red teaming Amusing that they use A100e as the reference point to sound impressive. Different ways you could make that conversion, but based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point), that's something like 200hr on a GB300 NVL72 rack. Not nothing either, but far less astounding sounding than 700k hrs.
Re: GPT-5.6
#279GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.
Or a breakthrough in algorithms etc.
The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.
Re: GPT-5.6
#280Earlier quoted context omitted.
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
Are you joking? They spend billions of dollars training LLMs to get a 7.8% on arc agi 3 whereas DINO models are near sota in image classification, provide meaningful embeddings to the point where image segmentation is just PCA. The spend on DINO cannot be more than five million (correct me if I'm wrong) JEPA is just getting started