Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

191–200 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#191

Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…

[deleted]

Re: Previewing GPT‑5.6 Sol: a next-generation model

#193
post #4

"Next generation model" If it was the next generation, why isn't it a major version change..?

AFAIK there is no difference between "generation" and "version". Version naming/numbering depends on how good it turns out to be, and competition. If the competition releases something then you need to push something out too. Calling it 5.6 creates the least possible expectations, and therefore more potential for positive feedback. The Sol/Terra/Luna naming is interesting. I wonder what Anthropic are considering for…

You gotta check out the new ChatGPT 6.3 Betelgeuse bro

Re: Previewing GPT‑5.6 Sol: a next-generation model

#195

Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…

OpenAI also announced two days ago that they're starting to make Cerebras style chips themselves [0], will be interesting to see how fast SotA model inference will be by the end of the year. [0]: https://openai.com/index/openai-broadcom-jalapeno-inference-...

Cerebras is different than what jalapeno is.

Jalepeno is for mass scale inference.

Cerebras is extremely expensive and difficult to scale, hence the limited release.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#196
post #93

Earlier quoted context omitted.

You mean Llanfairpwllgwyngyllgogerychwyrndrobwllllantysiliogogogoch ?

What is happening. I feel like I'm getting an aneurysm reading these comments.

It's the name of a place in Wales, which has made it a running joke for decades!

Re: Previewing GPT‑5.6 Sol: a next-generation model

#197

Earlier quoted context omitted.

For all intents and purposes you'll be able to move an open weight model wherever you want. I really dislike this rhetoric, you sound like the FSF guys who are like "you're not free until you're running coreboot with zero binary blobs". Sure they have a point but also, most people are fine running regular linux.

Most FSF guys actually have very nuanced views on the topic and you’re doing everyone a disservice by reducing it to an extremist sound bite.

Thankfully he didn't say that they're all like that. Instead he pointed out the few that are as a well known example of similar behavior.

If you reread the comment with a fresh mind you'll notice that you misunderstood what he wrote

Re: Previewing GPT‑5.6 Sol: a next-generation model

#198

Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…

For comparison, openrouter says opus 4.8 is ~55 tokens/s and fast mode is ~102. 750 tokens/s for their largest model is going to be nuts

Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750!

I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...

Re: Previewing GPT‑5.6 Sol: a next-generation model

#199

Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…

OpenAI also announced two days ago that they're starting to make Cerebras style chips themselves [0], will be interesting to see how fast SotA model inference will be by the end of the year. [0]: https://openai.com/index/openai-broadcom-jalapeno-inference-...

I don't see any indications that OpenAI is doing wafer-scale work.

I tend to doubt they would. Cerebras notably doesn't have a kv, is wildly high bandwidth, but within/across the chip, not able to dump/restore kv super well. I doubt openai is going to build something that is as expensive to run. Also, wafer-scale is absurdly hard & weird to pull off, so I doubt that would be their first foray.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#200
post #67

> Additionally, we’re introducing a new `ultra` mode that goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work. I'm curious about how does this work? Do the subagents also get to use the same tools? Will the client be flooded with tool calls? Why extra pricing for a new "model" when the same thing can happen in the client with more controls? And if it's an army of subagent…

I'm shocked they didn't use subagents already. Maybe they're just talking about their web deployment being unified with codex?

Deep Research has been using the Orchestrator -> Subagents -> Synthesizer loop since the beginning. It's just strange that they'd put a loop benchmark next to actual model benchmarks.

Maybe it's a tune of the base model that works especially well with the subagent loop?

Post reply on HN