Live data from Hacker News

GPT-5.6

openai.com

331–340 of 1001 posts

Re: GPT-5.6

#331
post #234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

These aren’t raw base models they are the result of a ton of RLHF and various adjustments.

Bitter lesson wildly overstated in this context.

Re: GPT-5.6

#332
post #234

Earlier quoted context omitted.

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

This isn’t really how it works anymore. Agents rely heavily on tool use and the agentic harness to perform tasks. Pre-training is no longer very effective.

I thought models werent allowed tools on arc-agi?

Re: GPT-5.6

#333

Earlier quoted context omitted.

I totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm. Great catch.

The SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.

SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.

Re: GPT-5.6

#335
post #17

I haven't tried an OpenAI model for a long time, but with Fable going to API pricing soon this might be enough to get me to try codex.

Seeing how Anthropomorphic just reset usage quotas back to 0 and the other day extended Fable sub inclusion by a few days, I have a feeling they might not drop Fable out of sub after all, because like you I would most definitely take a long good look at codex at that point.

Re: GPT-5.6

#336

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> can better infer the user’s underlying goal and intended level of work

This is a trap.

It's the optimistic fallacy that poisons all "consumer scale" machine learning products and what's going to effectively ruin these models as they keep chasing it in the same way that web queries were ruined, social media feeds were ruined, and media recommenders were ruined.

For the vendor, optimizing metrics across their whole user base, they always see positive technological progress as their system gets better at making assumptions and accumulating user engagement scores in aggregate. But for the individual user, most of which has some weird tail intent/interest and some of whom have many weird tail intent/interests, the experience quietly but catastrophically degrades. Output/results become more generic, more divergent with the underspecified "weird tail" intent, and more stubbornly hard to ever wrangle towards that "weird tail" altogether.

We've been watching this cycle happen for 20 years now and it's proving hard for anybody to escape because it works so well for the trillion dollar company driving it forward. But while each step might feel ergonomic and welcome to individual users, there's a frog boiling enshitification at play.

In pursuit of output quality and capability (rather than simply the vendor's user count), what we need rather than "makes better guesses" is "presses for more clarity", even where it feels kind of annoying.

Even among human professionals, one of the first hurdles of breaking out of junior tier work is gaining the confidence to press your colleagues and clients to be more specific in their thoughts and expressions despite their desire to have you do it all for them. But they're often coming to you with incomplete, muddy, and conflicting ideas for which there is no safe and correct assumption that you might just run with, and it's your expertise (i.e. relevant "intelligence") that's critical to bringing attention to that. To achieve professional progression, you need to learn to do that and to not just optimize appeasing the ambiguous client/colleague today in exchange for mutual expense tomorrow. To avoid enshitification, which is probably not possible, we need these models to be learning that too.

Re: GPT-5.6

#337

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

Falsifying Yann Lecun isn't exactly a priority for anyone seriously working in this space.

Re: GPT-5.6

#338
post #151
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

It's not clear replies to this thread aren't openAI employees or incentivized influencers, but every benchmark has gpt-5.5 underperforming opus 4.8, sometimes by as much as 10%. Can they all be wrong/paid-off?

[deleted]

Re: GPT-5.6

#339
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I recommend trying Codex too. In fact, I recommend running them side-by-side if you have the budget, e.g. have both independently plan the same feature or implement in a different worktree, or have them critique each other's work.

I personally find GPT-5.5 to be a better programmer than Opus 4.8, it is extremely thorough, but I don't like the code it generates ("austere"), and find Opus 4.8 to write more "human friendly" code. The programming comments GPT-5.5 makes is pretty awful where-as Opus 4.8 is good. I feel like Opus 4.8 is better at grasping my intention than GPT-5.5, and honestly find GPT-5.5 to be kind of "autistic". I do prefer the language (not the writing) of GPT-5.5, as I find the philosophical flowery language of Opus 4.8 kind of annoying.

I have only managed to try Fable 5 a little bit, which feels like a much more generally smarter version of Opus 4.8, that is much better a programming and grasping your intention, and I think even the intention of your code, and is _really_ good at spotting bugs or problems with logic in your code. It feels wicked smart but is extemely expensive. It feels smart in the sense like it has a "bigger brain" and is much more sensitive to subtleties/details.

These are different "brains", have different "personalities", etc. I think the best thing is to develop a feeling for it yourself.

Post reply on HN