Earlier quoted context omitted.
I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it. But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of…
I think it's still useful in a "hello world" sort of way. It means you actually tried the new model.
GPT-5.6
641–650 of 1001 posts
Re: GPT-5.6
#642Earlier quoted context omitted.
What destructive actions are you afraid of in particular? Honestly the models are pretty smart, I let the agents go --yolo and nothing bad has ever happened (yet) that couldn't be solved with git.
I'm not concerned about the code it's working on, but rather anything else - modifying files outside of the project dir (e.g. incorrect tool call), modifying system configuration, doing something bad on the internet, etc.
Re: GPT-5.6
#643GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
Re: GPT-5.6
#644We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
Re: GPT-5.6
#645Re: GPT-5.6
#646Earlier quoted context omitted.
If you conceptualize this as “there is an appropriate amount of brevity for each situation” then it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount. My view is that popular models by default output wildly excessive amounts of prose for nearly every use case, so if this changes in a new model that’s a pure win.
The models don't get better, except when a new one is released. Their performance depends solely on the model training before release and how well you curate the context you feed it. That's it. Contrary to popular belief these things are not intelligent.
Re: GPT-5.6
#647Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.
Re: GPT-5.6
#648Earlier quoted context omitted.
Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.
I don't know. If they were training on this, I feel like they would be able to get the shape of a bike frame right; it's a pretty simple polygon, and a lot of the bike frames that are getting generated would be impossible to steer.
https://themagnet.substack.com/p/why-is-it-so-hard-to-draw-a...
Re: GPT-5.6
#649Really wanna see it in DeepSWE benchmark
Re: GPT-5.6
#650Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!
This is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.
It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust it to do a task without deferring to Opus and that's really annoying at times. I want to know what I'm getting up front, not after the fact.