Live data from Hacker News

GPT-5.6

openai.com

961–970 of 1001 posts

Re: GPT-5.6

#961

Earlier quoted context omitted.

Seeing how Anthropomorphic just reset usage quotas back to 0 and the other day extended Fable sub inclusion by a few days, I have a feeling they might not drop Fable out of sub after all, because like you I would most definitely take a long good look at codex at that point.

It's not just the API pricing either, there's also the constant uncertainty. They pull the model then put it back up, they say the model is going away then suddenly it's not. And then there's the fact Fable is barely usable because it randomly downgrades to Opus out of nowhere whenever it thinks about exploits. It's definitely good that Anthropic's feeling the pressure. Anthropic has worn out their welcome with this…

> And then there's the fact Fable is barely usable because it randomly downgrades to Opus out of nowhere whenever it thinks about exploits.

I suspect it's not what you meant, but it's definitely not random and is very deliberate. Just today I got it to reliably trigger the "safety" filter with (drumroll) having it list the weight keys of a 300M parameter ModernBERT-derived model. Their "safety" classifier must be matching one of the key names in there and trigger their "this is a frontier model" anti-competitive filter[1] (even though it's just a tiny 300M parameter model, four orders of magnitude smaller than the frontier).

[1]: https://news.ycombinator.com/item?id=48464732

Fortunately once you know how it works (i.e. dumb keyword classifier) it's easy-ish to get around: just rename the keys so that it doesn't contain the naughty keyword. (At least as long as it doesn't trigger on something in its own thinking trace, which needs... more creative workarounds.)

Re: GPT-5.6

#962
post #905

I've been testing Sol/Terra/Luna now since yesterday, running complex evals on all of them and I feel a bit... mixed on how they perform. The eval is an agent that runs a set of tools and a prompt we can tune separately for different models. The OpenAI version of the prompt was specifically tuned based on their guide[0]. Then we let Opus to run another agent that acts as a user, trying to solve a problem (anonymized…

I thought I was an OpenAI fanboy, but version 5.6 isn’t for me. Sol Ultra just keeps working and checking, and working and checking again, but it can’t even correct minor errors that aren’t a problem for 5.5 xhigh. I’ve rolled Codex back to 5.5 for now.

The problem is they nerfed 5.5 about a month ago. The change was immediately visible to me: context compacting started to happen about 3x as frequently.

I think 5.6 is still not even close to 5.5 xhigh pre-nerf.

Re: GPT-5.6

#964
post #840

Earlier quoted context omitted.

Given that both Gemini 3.5 Flash (high) and Gemini 3 Flash Preview (medium) beat GPT-5.6 Sol (high) for correctness and score in your benchmarks I don’t trust them at all. The rest of the ranking also doesn’t make sense, like GPT-5.3-Codex (medium) performs better than Claude Opus 4.8 (medium) yeah sure

It's because the benchmark is not coding-only. Gemini models tend to have most knowledge for most domains, and are one of the most intelligent overall. You can check other benchmarks too, on specific categories, those models still beat other SOTA models. Regarding Opus, Anthropic models often fail to follow instructions, formatting requirements or simply refuse to answer questions (i.e. Fable). The issue with Gemini…

100% agree with your statement that gemini models have most knowledge in domains. They shine in so many generic topics

Re: GPT-5.6

#965
post #905

I've been testing Sol/Terra/Luna now since yesterday, running complex evals on all of them and I feel a bit... mixed on how they perform. The eval is an agent that runs a set of tools and a prompt we can tune separately for different models. The OpenAI version of the prompt was specifically tuned based on their guide[0]. Then we let Opus to run another agent that acts as a user, trying to solve a problem (anonymized…

Yes, this is exactly why I said yesterday that OpenAI does benchmaxxing, seemingly quite a bit [1]. I got a flurry of downvotes for it at first, but I think people came around to it once they tried the model like you did.

Ultimately I think the issue is that OpenAI is under tremendous pressure to perform, but GPT-6 is not ready yet, so they had to push GPT-5 to its limits, and the only way they could do it was with really heavy RLHF, which has its shortcomings. Like, it is super obvious that Sol, Terra and Luna are all heavily biased towards working on a problem relentlessly because that's what their reward functions emphasized. That pushes up their scores in some benchmarks but does not translate to actual intelligence and capability.

[1]https://news.ycombinator.com/item?id=48849454

Re: GPT-5.6

#966

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

This is the first I have herd of this benchmark. Can someone explain how it in any way indicates how close we are to "AGI"? Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9... It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it…

> Can someone explain how it in any way indicates how close we are to "AGI"?

I think it is historical name. At some point when benchmarking was very undeveloped, this was targeting abstract reasoning and generalization, hence AGI.

Re: GPT-5.6

#967
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it. But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of…

Still fun to watch models trying their best :). I think "money spent" metric in this test is becoming the most interesting to watch. It is like looking at RT cost of "Hello, world!"...

Re: GPT-5.6

#968
post #893

I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…

I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…

> to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty?

you have to start using AI to achieve something which was way more challenging before, then you will feel lots of fulfillment and inspiration.

Re: GPT-5.6

#970

From my first tests today, it is a workhouse. It can scan my whole code base, optimize every part, with a greater level of autonomy than other tools. This is insane, we are living at the best time.

The literal worst time. I prefer meritocracies.

You think aristrocrats are using GPT and succeeding?
Post reply on HN