Live data from Hacker News

Previewing GPT‑5.6 Sol: a next-generation model

openai.com

771–780 of 797 posts

Re: Previewing GPT‑5.6 Sol: a next-generation model

#771
post #240

Earlier quoted context omitted.

I'm skeptical of how fast "up to" 750t/s really means. Maybe if they make it extremely expensive so it frees up enough capacity? GPT‑5.3‑Codex‑Spark currently runs on Cerebras chips and it's giving me around 150t/s. Still relatively very fast, but nowhere near the 1,000t/s they claimed at launch. (Also it's not a very good model.) That said, I'm super bought in to faster models being better for most use cases than sm…

If it's 150 t/s, that's barely faster than Nvidia GPUs who are batching a lot more and are a lot more cost effective. Add in the Groq piece and Nvidia claims it can do 400 tokens/s.

I assume it’s just oversubscribed. I’m sure it “can” go faster. But yeah that was my point.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#772
post #581

Earlier quoted context omitted.

is this specifically in codex? have been trying to use the models for months on opencode then pi but it says chatgpt subscriptions don't have access to it - i was under the assumption that OpenAI doesn't lock down their models based on harness a la Claude Code

What plan are you on? It is only available to Pro users.

Plus. No wonder - i suspected this but i couldnt find any docs. Side note, how are you liking Pro? I have really been considering getting Pro recently, but not sure if its more worth to just switch to openrouter. I feel like my usage currently barely outstrips the plus usage limits and Pro would be too much, and using Openrouter by default would mean I would have a lot more leeway to run more random lighter workloads without worrying about using up my limit, but I'll really miss GPT 5

Re: Previewing GPT‑5.6 Sol: a next-generation model

#775

Earlier quoted context omitted.

I’ve been using 1,000 t/s on a near frontier model for a month now. It’s very useful for agentic coding. It does require new approaches for me personally since I get a lot less time to think or read its output.

Which model and how can you achieve that speed, if you don't mind me asking?

MiMo 2.5 Pro UltraSpeed. Requires a brief application with Xioami and a day or two to get approved.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#776
post #746

Earlier quoted context omitted.

You have to understand that the median human is terrible at (almost) everything. Humans, the only examples of general intelligence we know, are economically valuable precisely because they can train themselves to specialise at a (relatively) narrow task over time. You don’t measure how good a coding model is by how well it programs relative to Doctors, or how well it can prove theorems relative to baristas, or how we…

> Humans, the only examples of general intelligence we know Our intelligence only seems "general" to us, because we're viewing it through our own eyes. Our "intelligence" is specialized to our survival, and we're terrible at most tasks outside that scope.

We operate and think about subjects like Higher Topos Theory, Information Geometry and Algebraic Topology, which are several layers of abstractions removed from anything that can be termed as a skill “specialised to our survival”.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#777

Earlier quoted context omitted.

>and why price it day 1 at 1/2 the cost of Fable? Why would they price it the same as Fable it it doesn't cost the same as Fable ?

That's half my point - Anthropic's remarks suggest that is Fable significantly bigger (hence more costly to run) than Opus, so it is priced accordingly, but GPT 5.6 priced the same as 5.5 is one datapoint that suggests they are the same size.

Yeah...and why would being the same size mean anything ? This is par the course for Open AI. They've always been cheaper and likely smaller than Opus models even when they weren't much if any worse.

Re: Previewing GPT‑5.6 Sol: a next-generation model

#778
post #502

Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…

> second to last There's a word for this that you should never pass up an opportunity to use: penultimate. (You should also never pass up the opportunity to use "defenestrate," but it sadly does not apply here.)

>> penultimate

Oh that's a word I haven't heard in a while. And I know it mainly thanks to Monty Python great sketch:

https://m.youtube.com/watch?v=l9Aj7W3g1qo

Re: Previewing GPT‑5.6 Sol: a next-generation model

#779

Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…

Gemini has done the same thing, gemini-3.5-flash is 15x more expensive for input tokens than gemini-2.0-flash. They are forcing us up the pricing ladder by deprecating the old models....

Re: Previewing GPT‑5.6 Sol: a next-generation model

#780

Earlier quoted context omitted.

That's half my point - Anthropic's remarks suggest that is Fable significantly bigger (hence more costly to run) than Opus, so it is priced accordingly, but GPT 5.6 priced the same as 5.5 is one datapoint that suggests they are the same size.

Yeah...and why would being the same size mean anything ? This is par the course for Open AI. They've always been cheaper and likely smaller than Opus models even when they weren't much if any worse.

> why would being the same size mean anything

Because these companies are still 100% scale-pilled, and each new generation of model is bigger than the previous one, even if active parameters (MoE) is growing slower than total parameters. There is also scaling in inference-time compute which adds to the cost just the same. Fable is rumored to use variable compute which will increase the cost for inputs that get routed for more compute.

Maybe as you suggest the OpenAI pricing is unrelated to size/cost, but if so it seems a very aggressive move given that models that cost more to train and run need to generate more revenue to be profitable.

Post reply on HN