Live data from Hacker News

GPT-5.6

openai.com

641–650 of 1001 posts

Re: GPT-5.6

#641

Earlier quoted context omitted.

I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it. But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of…

I think it's still useful in a "hello world" sort of way. It means you actually tried the new model.

Honestly that's the main value I get from it myself - making a pelican means I have to figure out API keys and how to talk to the provider, or how to run it locally for the local models.

Re: GPT-5.6

#642
post #520

Earlier quoted context omitted.

What destructive actions are you afraid of in particular? Honestly the models are pretty smart, I let the agents go --yolo and nothing bad has ever happened (yet) that couldn't be solved with git.

I'm not concerned about the code it's working on, but rather anything else - modifying files outside of the project dir (e.g. incorrect tool call), modifying system configuration, doing something bad on the internet, etc.

[deleted]

Re: GPT-5.6

#643

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

I'm surprised it is that low. Are not all top AI labs "cheating" and workaround LLMs's low sample efficiency by hiring people to generate more data points - similar problems with answers, so they can train models on those and improve scores? A good benchmark for general intelligence probably should be a complete black box, no sample data given/leaked at all.

Re: GPT-5.6

#644

We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.

Not at all, we love them all with Chinese labs. And wish them to continue competing and not winning. That is how we get best models, lower prices and better availability.

Re: GPT-5.6

#645
Maybe it’s a bug but on iOS individual paid Pro account - I can no longer see which model is being used nor select which model I want.

Re: GPT-5.6

#646

Earlier quoted context omitted.

If you conceptualize this as “there is an appropriate amount of brevity for each situation” then it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount. My view is that popular models by default output wildly excessive amounts of prose for nearly every use case, so if this changes in a new model that’s a pure win.

The models don't get better, except when a new one is released. Their performance depends solely on the model training before release and how well you curate the context you feed it. That's it. Contrary to popular belief these things are not intelligent.

[dead]

Re: GPT-5.6

#647
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

At least Anthropic doesn't bend down to Pentagon/Trump administration

Re: GPT-5.6

#648

Earlier quoted context omitted.

Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.

I don't know. If they were training on this, I feel like they would be able to get the shape of a bike frame right; it's a pretty simple polygon, and a lot of the bike frames that are getting generated would be impossible to steer.

I mean, humans can't draw bikes!

https://themagnet.substack.com/p/why-is-it-so-hard-to-draw-a...

Re: GPT-5.6

#650
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

This is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.

It's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work.

It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust it to do a task without deferring to Opus and that's really annoying at times. I want to know what I'm getting up front, not after the fact.

Post reply on HN