Live data from Hacker News

GPT-5.6

openai.com

881–890 of 1001 posts

Re: GPT-5.6

#881
post #794

I really appreciate the focus on intelligence WITH token efficiency. I'd like to see that become the trend. Smartest per token metrics. Least tokens to accomplish the task above a certain success level. Most of my tasks would benefit from efficiency / token, but switching models constantly, and trying to guess the right model and effort level takes up too much of my processing.

> I really appreciate the focus on intelligence WITH token efficiency that is a polite way of saying "I don't believe AGI is coming anytime soon".

I honestly interpreted and agreed with the version "this saves money".

Re: GPT-5.6

#882

We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.

personally I hope any company involved with child slaughter ends up crashing and burning, i say this because both those companies are buddies with the us department of war (who helped annihilate a school the other day)

thank you for being a voice of reason. I am with you.

Re: GPT-5.6

#883

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

We have it slightly ahead of Fable in our multi-agent coding evaluations. Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol. Fable scores slightly lower, but with an average solution size of 12.2 KB. Data at…

Seems quite kind to Gemini models.

Re: GPT-5.6

#884

Earlier quoted context omitted.

This is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.

It's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work. It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust…

I asked him about sharks to be able to answer my kids question and it got triggered somehow. Then again when I asked it if my code had bugs or vulnerabilities before I commit.

At some point just kill the thing, it's not able to work properly as it is.

Re: GPT-5.6

#885
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I have them talk to each other via tmux to great effect on complex changes. Its great for auditing changes as work is done.

Re: GPT-5.6

#886

Earlier quoted context omitted.

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

Notice how neither him, nor Ilya, nor Mira shipped anything relevant recently It's telling

Not sure how Mira gets into the same sentence as Yann and Ilya.

As far as the lack of shipping, they're scientists and what we're doing now with LLMs is more "engineering."

Re: GPT-5.6

#888

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

isnt native english speaker

"i really wish this thing in my non native language was easier to decipher"

huh? if you dont know the words then read them in your native language. Sol/Terra/Luna are immediately unambiguous to an english speaker with any sense.

Re: GPT-5.6

#890

We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.

personally I hope any company involved with child slaughter ends up crashing and burning, i say this because both those companies are buddies with the us department of war (who helped annihilate a school the other day)

damn so brave
Post reply on HN