Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

301–310 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#301
post #254

Did GPT-5.6 Sol Ultra decide the terrible colors for the benchmark graphs?

I remember them using these chart colours during the 5 launch, maybe even 4.1 back in the day. Don’t know why, maybe its their CI manual that’s been generated by gpt-3.5-turbo…

Re: Previewing GPT‑5.6 Sol: a next-generation model

#302
post #255

Earlier quoted context omitted.

He is not an ML researcher or engineer, he is a passionate AI enthusiast blogger. He mostly does SVGs and other low effort checks (sometimes with major flaws, as people have pointed out a few times in the HN comments). Properly evaluating the model across all fronts requires a deep understanding of LLMs, how they work, the trade offs behind new architectures and the relevant research papers. It also takes a lot of ti…

He created Django, what do you mean he's not an engineer? Also 'low-effort??' his posts are extremely in-depth, clearly very thought through with a significant amount of time and energy. Additionally he does perform multifaceted checks across LLMs in many of his other blog posts.

> ML researcher or engineer

The charitable reading is that they meant “ML researcher or ML engineer” with the latter meaning, I guess, an engineer who works on developing LLMs not just using them.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#303
post #44

I think GPT writes code the best. How well will it write in version 5.6? It gives me chills. Recently, I went head-to-head with GPT on nearly 2,000 lines of code, and GPT's solution was superior and faster. I even referenced multiple codebases on GitHub while trying, but they were incomparable to GPT. So using GPT brings both fear and excitement. The fear comes from realizing that this level of code is now the averag…

Is it possible for you to provide examples? What were you trying to solve? What was your solution and why was GPT's solution superior and faster?

> ... why was GPT's solution superior and faster?

Not saying that's the case with OP, but I've found folks sometimes just rationalize it so [0] as they're paying top dollar for it (especially, when compared to may be less capable but affordable models).

[0] https://en.wikipedia.org/wiki/Choice-supportive_bias

Re: Previewing GPT‑5.6 Sol: a next-generation model

#304
post #59

Earlier quoted context omitted.

Unless you are hosting it yourself on your own infrastructure it absolutely can be taken away.

>Unless you're running Linux yourself, it can absolutely be taken away.

Yes. The difference is obviously that full, fat Linux runs on a superset of anything a layperson would call a computer, and can be built from source on roughly the same set of hardware. Running the full, fat Deepseek (as in the 1.6T model, unquantized) is too big to run on anything a layperson would call a computer, and being able to actually build it is even harder.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#305

How are they able to compare with Fable when Fable was only available for three days?

Terminalbench numbers are publicly available. What is more interesting, why is that the only benchmark they highlight. Maybe 5.6 isn’t that far ahead of Fable 5 in DeepSWE and FrontierCode (which I consider the most useful and close to my evals + subjective experience)…

Re: Previewing GPT‑5.6 Sol: a next-generation model

#306

I hate not being able to use the latest models. There needs to be a much faster resolution to whatever is happening with the federal government.

How else is this administration going to make money?!? How dare you...if they do not accept bribes...what is there left for them? This is a premium buy...First one to beat competition gets the worm. So, you pay Trump, trump gives you access...then you pay subscription to SAMA lol.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#307
post #56

Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…

If you have no need for Anthropic/OpenAI's frontier model capability, you may be better served with an open-weight model that can't be taken away. Edit: > GPT-5 does the job. I bring up DeepSeek V4 Flash a lot on HN, but I want to mention that according to Artificial Analysis, it trades blows with GPT-5 (high) (from August, 2025) [0] [0]: https://artificialanalysis.ai/models/comparisons/deepseek-v4...

It’s my daily driver in opencode

Re: Previewing GPT‑5.6 Sol: a next-generation model

#308

I can’t help but think that these benchmarks are completely fake. Sam even posted a benchmark on X a couple days ago of how the ‘complete version’ of 5.5 cyber was already ahead of Mythos apparently. This just feels like absolutely fake nonsense. The impact of Mythos on the industry was clear and in front of everyone’s eyes. The amount of vulnerabilities Mozilla fixed. The vulnerabilities and exploits Anthropic showc…

Well if they are posting fraudulent benchmarks, that's a good sign to invest in their IPO. It's pure downside protection: IPO does well, profit. IPO does poorly, concrete evidence of pre-IPO fraud.

I personally don't think it's likely that OpenAI would post completely fake numbers in this pre-IPO period, but if you do, this is an opportunity.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#310

"We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." This seems like it would be the largest and first closed-source model Cerebras has offered till date

Codex Spark models already run on Cerebras
Post reply on HN