Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

631–640 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#632
post #284

Earlier quoted context omitted.

“More intelligence” is the new feature. Almost everyone is asking for this. Citation: have you looked at OAI and Anthropic’s customer growth numbers?

Every use case of every customer doesn’t need more intelligence. I’m willing to bet that the vast majority will be perfectly fine running on “low intelligence” at a cheap price forever.

What are you talking about?

Prices of lowest tiers of models have fallen how much - 10-100x over the last two years.

And actually, the model quality you needed to pay for in the past, you can just run on device now essentially for free.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#633

Earlier quoted context omitted.

> Maybe it’s the realization that it was never that cheap in the first place and they're forcing us to upgrade in a slow and painful way. All the analysis I have seen points to frontier models being profitable to serve. It’s using 50% or more of your GPUs for research plus CapEx for capacity expansion that makes these businesses so heavily cash-negative. What you are observing is downstream of another detail. It gets…

There is really ample analysis pointing to inference not being profitable, look at anything Ed Zitron has reported.

Ed Zitron is amazing at cherry picking data to fit his thesis.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#634

Earlier quoted context omitted.

Every time I use opus these days I go shut up... you are not fable.. Hard to imagine how just three days with it changed how I saw LLM use.

I really don't feel this way. Seemed pretty similar to me, noticeably better, but marginally. What am I missing?

It may depend on your specific workload. E.g. for regular webdev work Opus is more than adequate, for heavy duty data analysis, for experimental stuff and for complex systems it was night and day.

I had only a few places where I did spot a difference but that difference was significant and I can imagine where people would be amazed.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#635
post #168

Earlier quoted context omitted.

I'm suspect on how much of a coding advance it will be. Seems odd that their announcement has zero coding benchmarks, with the closest related thing being terminal bench.

Maybe I'll know once I try it? Honestly, for small functions or methods, I don't think there's a huge difference between models. But the larger the code gets, the more noticeable the difference seems to be. Personally, I think this kind of coding experience varies from person to person

Not the size of function but conplexity.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#636
post #516

GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints. https://metr.org/blog/2026-06-2…

I know it messes up their eval scores but to me this kind of cheating is a better demonstration of intelligence than just attempting the tasks algorithmically.

Maybe true, but if you're using an LLM to do some real world work, do you want it to have some abstract notion of intelligence, or do you want it to actually do the job you assigned it?

Re: Previewing GPT‑5.6 Sol: a next-generation model

#637

Earlier quoted context omitted.

The definition has already been stretched to not fit the previous models. There is no meaningful, static definition that significantly predates current capabilities. There's a reason why ai xrisk doomers had to come up with the term ASI. I would seriously suggest that everyone take a look at the wikipedia page for AGI from the month before ChatGPT was released, compare it to the current version, and not come to that…

The first sentence is “understand or learn any intellectual task that a human can.” Whatever you think of the benefits of LLMs, they don’t understand and they can only learn during the training period and with very minor adjustments in post training. So, no I don’t think any of these models are generally intelligent.

> they don’t understand

I have not seen any instance of this frequently-made assertion which is at all justified. It seems to rely on a definition of "understand" which is more about spirituality than actual observable evidence (they clearly can comprehend even complex tasks well enough to execute on them, and if you won't call that "understanding", you're playing word games rather than stating an objective fact).

Likewise, agents can literally come to a greater understanding of a problem through trial and error, and there are plenty of mechanisms to retain that knowledge. If you don't want to call that "learning", you're just making a choice to define it in a way more restrictive than how we use it for humans, and intentionally making communication more difficult.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#638
post #56

Earlier quoted context omitted.

If you have no need for Anthropic/OpenAI's frontier model capability, you may be better served with an open-weight model that can't be taken away. Edit: > GPT-5 does the job. I bring up DeepSeek V4 Flash a lot on HN, but I want to mention that according to Artificial Analysis, it trades blows with GPT-5 (high) (from August, 2025) [0] [0]: https://artificialanalysis.ai/models/comparisons/deepseek-v4...

deepseek has no part of their privacy policy on their API about training. They are 100% training on every single word you give it. If your customers are fine with that, your IP is not interesting, then you can use it.

Though with open models you have a lot of choice where to get it from. I see like ~15 providers here with various logging/ZDR policies, so pick whatever mix of price to features you want:

https://openrouter.ai/deepseek/deepseek-v4-flash

Re: Previewing GPT‑5.6 Sol: a next-generation model

#639

Earlier quoted context omitted.

And what is it worse at than an average human today that can be done on a computer?

almost everything? AGI has to be able to completely replace a human in any information worker role indefinitely.

I think you're speeding past the word "average" in the sentence. I'd argue that current frontier models already exceed the abilities of average humans across the majority of tasks you can do on a computer, although you might be able to argue that they tend to be a bit slower?

That latter part is debatable though - have you seen a non-technical person try to figure out something new on a computer?

Post reply on HN