I really appreciate the focus on intelligence WITH token efficiency. I'd like to see that become the trend. Smartest per token metrics. Least tokens to accomplish the task above a certain success level. Most of my tasks would benefit from efficiency / token, but switching models constantly, and trying to guess the right model and effort level takes up too much of my processing.
> I really appreciate the focus on intelligence WITH token efficiency that is a polite way of saying "I don't believe AGI is coming anytime soon".
GPT-5.6
881–890 of 1001 posts
Re: GPT-5.6
#882We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
personally I hope any company involved with child slaughter ends up crashing and burning, i say this because both those companies are buddies with the us department of war (who helped annihilate a school the other day)
Re: GPT-5.6
#883GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
We have it slightly ahead of Fable in our multi-agent coding evaluations. Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol. Fable scores slightly lower, but with an average solution size of 12.2 KB. Data at…
Re: GPT-5.6
#884Earlier quoted context omitted.
This is a major reason why I and a number of biologists I've talked to have canceled their anthropic accounts recently. Not working is not working.
It's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work. It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust…
At some point just kill the thing, it's not able to work properly as it is.
Re: GPT-5.6
#885Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Re: GPT-5.6
#886Earlier quoted context omitted.
Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
Notice how neither him, nor Ilya, nor Mira shipped anything relevant recently It's telling
As far as the lack of shipping, they're scientists and what we're doing now with LLMs is more "engineering."
Re: GPT-5.6
#887Re: GPT-5.6
#888I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…
"i really wish this thing in my non native language was easier to decipher"
huh? if you dont know the words then read them in your native language. Sol/Terra/Luna are immediately unambiguous to an english speaker with any sense.
Re: GPT-5.6
#889Re: GPT-5.6
#890We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
personally I hope any company involved with child slaughter ends up crashing and burning, i say this because both those companies are buddies with the us department of war (who helped annihilate a school the other day)