Earlier quoted context omitted.
> Intent understanding This will totally make it brain damaged over a certain tasks. Sort of like the same brain damage that prompted OpenAI project managers to destroy ChatGPT.app today.
Can you elaborate?
GPT-5.6
711–720 of 1001 posts
Re: GPT-5.6
#712And 5.6 Luna ($0.21) is also impressive, cheaper than GLM 5.2 ($0.37) with higher intelligence.
Re: GPT-5.6
#713Earlier quoted context omitted.
On the one hand: yes, pelicans on bikes are definitely in the training set at this point. On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.
I sort of agree, but within the same model I expect the reasoning effort to be reflected in the quality of output and that's basically how it played out. When you're comparing different models, then it's just who benchmaxxed the best and there's not a lot of value there.
Re: GPT-5.6
#714Earlier quoted context omitted.
> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.
The human brain has 80 billion neurons and a 100 trillion synapses. I think you're underselling the processing power of that warm chunk of meat. The real message of the last 15 years has actually been the opposite: if you throw enough processing power at it, intelligence emerges.
Re: GPT-5.6
#715Earlier quoted context omitted.
Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.
The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.
Re: GPT-5.6
#716Re: GPT-5.6
#717Earlier quoted context omitted.
Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
Mythos can do some amazing things (I'm assuming, I've never seen it). A young child can learn to control its body without reading any books on dynamical systems and kinematics. Mythos cannot learn to control a humanoid robot after sucking in every piece of data Anthropic can get their hands on.
Re: GPT-5.6
#718Earlier quoted context omitted.
If you conceptualize this as “there is an appropriate amount of brevity for each situation” then it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount. My view is that popular models by default output wildly excessive amounts of prose for nearly every use case, so if this changes in a new model that’s a pure win.
The models don't get better, except when a new one is released. Their performance depends solely on the model training before release and how well you curate the context you feed it. That's it. Contrary to popular belief these things are not intelligent.
Not quite. The hosting side can change reasoning budgets (or re-assign what terms like "high" means), temperature and other decoding parameters, output length limits, finetune internal "hidden" prompt, latency optimizations, finetune attention algorithms, even change quantization - all still serving as the same model.
We know (or suspect) Anthropic frequently nerfs models while keeping their name and version the same.
Re: GPT-5.6
#719Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.