Earlier quoted context omitted.
I'm skeptical of how fast "up to" 750t/s really means. Maybe if they make it extremely expensive so it frees up enough capacity? GPT‑5.3‑Codex‑Spark currently runs on Cerebras chips and it's giving me around 150t/s. Still relatively very fast, but nowhere near the 1,000t/s they claimed at launch. (Also it's not a very good model.) That said, I'm super bought in to faster models being better for most use cases than sm…
If it's 150 t/s, that's barely faster than Nvidia GPUs who are batching a lot more and are a lot more cost effective. Add in the Groq piece and Nvidia claims it can do 400 tokens/s.
Previewing GPT‑5.6 Sol: a next-generation model
771–780 of 797 posts
Re: Previewing GPT‑5.6 Sol: a next-generation model
#772Earlier quoted context omitted.
is this specifically in codex? have been trying to use the models for months on opencode then pi but it says chatgpt subscriptions don't have access to it - i was under the assumption that OpenAI doesn't lock down their models based on harness a la Claude Code
What plan are you on? It is only available to Pro users.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#773Re: Previewing GPT‑5.6 Sol: a next-generation model
#774Re: Previewing GPT‑5.6 Sol: a next-generation model
#775Earlier quoted context omitted.
I’ve been using 1,000 t/s on a near frontier model for a month now. It’s very useful for agentic coding. It does require new approaches for me personally since I get a lot less time to think or read its output.
Which model and how can you achieve that speed, if you don't mind me asking?
Re: Previewing GPT‑5.6 Sol: a next-generation model
#776Earlier quoted context omitted.
You have to understand that the median human is terrible at (almost) everything. Humans, the only examples of general intelligence we know, are economically valuable precisely because they can train themselves to specialise at a (relatively) narrow task over time. You don’t measure how good a coding model is by how well it programs relative to Doctors, or how well it can prove theorems relative to baristas, or how we…
> Humans, the only examples of general intelligence we know Our intelligence only seems "general" to us, because we're viewing it through our own eyes. Our "intelligence" is specialized to our survival, and we're terrible at most tasks outside that scope.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#777Earlier quoted context omitted.
>and why price it day 1 at 1/2 the cost of Fable? Why would they price it the same as Fable it it doesn't cost the same as Fable ?
That's half my point - Anthropic's remarks suggest that is Fable significantly bigger (hence more costly to run) than Opus, so it is priced accordingly, but GPT 5.6 priced the same as 5.5 is one datapoint that suggests they are the same size.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#778Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…
> second to last There's a word for this that you should never pass up an opportunity to use: penultimate. (You should also never pass up the opportunity to use "defenestrate," but it sadly does not apply here.)
Oh that's a word I haven't heard in a while. And I know it mainly thanks to Monty Python great sketch:
Re: Previewing GPT‑5.6 Sol: a next-generation model
#779Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…
Re: Previewing GPT‑5.6 Sol: a next-generation model
#780Earlier quoted context omitted.
That's half my point - Anthropic's remarks suggest that is Fable significantly bigger (hence more costly to run) than Opus, so it is priced accordingly, but GPT 5.6 priced the same as 5.5 is one datapoint that suggests they are the same size.
Yeah...and why would being the same size mean anything ? This is par the course for Open AI. They've always been cheaper and likely smaller than Opus models even when they weren't much if any worse.
Because these companies are still 100% scale-pilled, and each new generation of model is bigger than the previous one, even if active parameters (MoE) is growing slower than total parameters. There is also scaling in inference-time compute which adds to the cost just the same. Fable is rumored to use variable compute which will increase the cost for inputs that get routed for more compute.
Maybe as you suggest the OpenAI pricing is unrelated to size/cost, but if so it seems a very aggressive move given that models that cost more to train and run need to generate more revenue to be profitable.