Live data from Hacker News

GPT-5.6

openai.com

711–720 of 1001 posts

Re: GPT-5.6

#711
post #500

Earlier quoted context omitted.

> Intent understanding This will totally make it brain damaged over a certain tasks. Sort of like the same brain damage that prompted OpenAI project managers to destroy ChatGPT.app today.

Can you elaborate?

It had a tough time updating today. Or this evening. It just wouldn't update. It actually just freaking disappeared from my MacBook. It took some googling and downloading and multiple tries to get it back and working. Because they also combine on a MacBook Codex with ChatGPT app. I guess codex became ChatGPT app or some silliness like that.

Re: GPT-5.6

#713

Earlier quoted context omitted.

On the one hand: yes, pelicans on bikes are definitely in the training set at this point. On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.

I sort of agree, but within the same model I expect the reasoning effort to be reflected in the quality of output and that's basically how it played out. When you're comparing different models, then it's just who benchmaxxed the best and there's not a lot of value there.

But that is my point: if benchmaxxing was all the labs were doing, then surely the dumber model could/would have equivalent performance? Rather than noticeably worse perf on a (somewhat trivial to game) test.

Re: GPT-5.6

#714

Earlier quoted context omitted.

> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.

The human brain has 80 billion neurons and a 100 trillion synapses. I think you're underselling the processing power of that warm chunk of meat. The real message of the last 15 years has actually been the opposite: if you throw enough processing power at it, intelligence emerges.

I think you're helping GPs point: there is a lot of efficiency gains to be made to match the processing power of the brain, given it's size and power draw.

Re: GPT-5.6

#715

Earlier quoted context omitted.

Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.

The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.

Bizarre? The size of the model is in the name. Sun, earth, and moon don't mean anything?

Re: GPT-5.6

#716
post #704

Earlier quoted context omitted.

Serious question: what is a short prompt? (For that matter at what point is it "long"? And does the rest of the context matter? Should it be short too?)

Why waste time say lot word when few word do trick?

I'm more concerned with things like skills

Re: GPT-5.6

#717
post #226

Earlier quoted context omitted.

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.

My main takeaway from LeCun's thesis isn't that you can't build LLMs to do useful things better than the best human, it's that these systems don't learn arbitrary skills efficiently, like humans do. And the question is, why not? 8% on ARC-AGI-3 is amazing for a machine considering how far we've come since digital computers were first built. But it is pretty poor if you're claiming something is well on its way to exhibiting human-like intelligence.

Mythos can do some amazing things (I'm assuming, I've never seen it). A young child can learn to control its body without reading any books on dynamical systems and kinematics. Mythos cannot learn to control a humanoid robot after sucking in every piece of data Anthropic can get their hands on.

Re: GPT-5.6

#718

Earlier quoted context omitted.

If you conceptualize this as “there is an appropriate amount of brevity for each situation” then it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount. My view is that popular models by default output wildly excessive amounts of prose for nearly every use case, so if this changes in a new model that’s a pure win.

The models don't get better, except when a new one is released. Their performance depends solely on the model training before release and how well you curate the context you feed it. That's it. Contrary to popular belief these things are not intelligent.

>The models don't get better, except when a new one is released. Their performance depends solely on the model training before release and how well you curate the context you feed it. That's it.

Not quite. The hosting side can change reasoning budgets (or re-assign what terms like "high" means), temperature and other decoding parameters, output length limits, finetune internal "hidden" prompt, latency optimizations, finetune attention algorithms, even change quantization - all still serving as the same model.

We know (or suspect) Anthropic frequently nerfs models while keeping their name and version the same.

Re: GPT-5.6

#719
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

And I know this is petty but the CC cli/harness just grates. It’s overcomplicated, performatively cutesy, and buggy. It’s in my way. The codex harness gives me what I need and gets out of the way.

Re: GPT-5.6

#720
Not sure what everyone's experience is but I find 5.6 Sol to be a great liar. Reported success on a half done job and left things in a broken state after having quite a few back & forth followups on the initial prompt to clarify the plan. Didn't experience this with 5.5. Opus 4.7 and below sometimes did it but they fixed it in Opus 4.8. So, overall, the initial experience has made me think that this model will be a lot more stressful to work with just because the level of trust that it actually completes the task is now much much lower.
Post reply on HN