Live data from Hacker News

GPT-5.6

openai.com

841–850 of 1001 posts

Re: GPT-5.6

#841
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

I recently asked Claude to help me choose a single MOSFET (transistor) for a specific use case in a mundane circuit. The safety triggered and it ended the conversation and refused to continue. Gemini has also done the same thing to me. Looks like the big players got very spooked by the temporary Trump admin ban on Mythos and they all locked down way too hard.

Re: GPT-5.6

#842
post #597

GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are still running, uploading them soon too) EDIT: The high variants are there too, enjoy the hamsters[0]. [0]: https://aibenchy.com/showcase/?q=gpt-5.6

Given that both Gemini 3.5 Flash (high) and Gemini 3 Flash Preview (medium) beat GPT-5.6 Sol (high) for correctness and score in your benchmarks I don’t trust them at all. The rest of the ranking also doesn’t make sense, like GPT-5.3-Codex (medium) performs better than Claude Opus 4.8 (medium) yeah sure

One example where the order seems correct, is this SVG generation test:

https://aibenchy.com/showcase/?q=Gemini+3.5%2Cgpt+5.6%2C+5.3...

You can see that most Gemini 3.5 generations are more correct than 5.6 Sol (the net is in the middle of the table, hamster seems reasonable and not deformed, etc.)

Re: GPT-5.6

#843
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

I recently asked Claude to help me choose a single MOSFET (transistor) for a specific use case in a mundane circuit. The safety triggered and it ended the conversation and refused to continue. Gemini has also done the same thing to me. Looks like the big players got very spooked by the temporary Trump admin ban on Mythos and they all locked down way too hard.

It triggered for me on a completely pedestrian game design prompt a couple of days ago. I’ve sent feedback and continued with Opus, but that was really unexpected

Re: GPT-5.6

#845
post #601

Earlier quoted context omitted.

Not always, in some cases, changing to a higher reasoning makes the AI doubt itself too much, and skip over the correct answer by overcomplicating the problem and polluting the context. It would be nice to see on which categories of problems the extra thinking makes it better and on which it makes it worse.

This shows up in OpenAI's graphs on their announcement page. There is a peak performance datapoint in the graphs past which (to the right on the graph indicating more resources spent) peformance declines. And it's on every graph on that page!

And in my tests, that point of "overthinking" depends on the problem's complexity, so it's not necessarily that using "xhigh" is always bad or good.

Re: GPT-5.6

#846

Earlier quoted context omitted.

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

For sure verbal diarrhea can be a problem. I think there's a difference between a generic instructions e.g. "be brief" and contextual guidance: "I am an experienced software developer with a recent undergraduate degree in pure mathematics. Be terse, I will ask questions if I need clarification."

[dead]

Re: GPT-5.6

#847
post #709

Earlier quoted context omitted.

Could we see the prompts, though?

I did a quick comparison of models a couple of months ago by giving it a MYTXTADV.BAS file and give them all the same prompt to create a sprite-based version of a text adventure game I wrote in Basic over 30 years ago. It was interesting to see where the approaches were similar and where they diverged.

For local coding agents, I test them by asking them to recreate the BASIC game Taipan (from 1979) in Python. So far one got close, making a terminal version with ASCII menus ala Drug Wars for the TI-83.

Re: GPT-5.6

#848

Earlier quoted context omitted.

This is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.

It's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work. It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust…

I asked whether an outdoors mosquito trap product (via a screenshot) would negatively impact other insect species in my garden and it refused. Though quick internet search did reveal that it would harm and trap many other species of harmless insects.

Re: GPT-5.6

#849

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

the size comparison didnt occur to me. I assumed the names were just random nice sounding focus-groupped marketing names.

Re: GPT-5.6

#850
post #665

Earlier quoted context omitted.

I think they’re saying it’s irrelevant now, possibly because it’s less likely to trail off on meandering thought bubbles.

Does anyone else feel each model is like watching your kids grow up. They we're bubbly and fun and weird, you needed to tell them to sit down and be quiet. Now if you tell them too much they go mute or stop telling you important information. Oh intelligence!

Yeah, don’t do that.
Post reply on HN