Earlier quoted context omitted.
For comparison, openrouter says opus 4.8 is ~55 tokens/s and fast mode is ~102. 750 tokens/s for their largest model is going to be nuts
Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750! I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...
Previewing GPT‑5.6 Sol: a next-generation model
221–230 of 797 posts
Re: Previewing GPT‑5.6 Sol: a next-generation model
#222This seems like it would be the largest and first closed-source model Cerebras has offered till date
Re: Previewing GPT‑5.6 Sol: a next-generation model
#223It appears that between GLM-5.2 and GPT-5.6, anthropic is feeling the heat, atleast in the bang-for-the-buck heuristic?
Re: Previewing GPT‑5.6 Sol: a next-generation model
#224Re: Previewing GPT‑5.6 Sol: a next-generation model
#225Other than the worst naming I have ever seen (Sol / Terra / Luna), the pricing is still expensive: > GPT‑5.6 is priced per 1M tokens across three model sizes: > Sol is $5 input / $30 output; > Terra is $2.50 input / $15 output > Luna is $1 input / $6 output. The OpenAI casino has never been more ready to take your money on gambling even more tokens.
With the $200/month plan I’ve never ran into any limits or issues. The product can be used every day for extensive sessions and development. What is everyone doing that makes them talk about tokens versus dollars?
Re: Previewing GPT‑5.6 Sol: a next-generation model
#226Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…
Yeah, this is the classic silicon valley strategy of selling at a loss and then once they have captured the market inflate prices. See Uber, Netflix, etc.
Feels like they are just pulling in as much as they can whilst competing on capabilities instead. At which point its a case of who can last the longest.
Doesn't feel like Uber/Netflix.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#227Re: Previewing GPT‑5.6 Sol: a next-generation model
#228Earlier quoted context omitted.
For comparison, openrouter says opus 4.8 is ~55 tokens/s and fast mode is ~102. 750 tokens/s for their largest model is going to be nuts
Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750! I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...
Re: Previewing GPT‑5.6 Sol: a next-generation model
#229Re: Previewing GPT‑5.6 Sol: a next-generation model
#230Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…
Yep this is a glimpse into the future of 500+ t/s, which is in my opinion the next big thing that validates Jevon's paradox (the models are already smart enough)