Earlier quoted context omitted.
Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750! I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...
The ChatGPT subscription gives you access to the -spark model(s) in Codex which are blazing fast (but pretty dumb) which I think runs on Cerebras hardware too.
Previewing GPT‑5.6 Sol: a next-generation model
581–590 of 797 posts
Re: Previewing GPT‑5.6 Sol: a next-generation model
#582Re: Previewing GPT‑5.6 Sol: a next-generation model
#583Earlier quoted context omitted.
And more importantly those 10 million tokens/s should cost fractions of a penny. Tokens need to be dirt cheap so I hope they build out massive solar+battery powered data centers asap.
No anything but wasteful, weak, expensive, environmentally harmful solar. Nuclear is the only path forward for superior energy production, at least until we figure out fusion.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#584Are we starting to see the 'we just realized that 100,000,000 GPU's later, 2+2 isn't the magic number, no matter how many times we calculate it' hit home?
Re: Previewing GPT‑5.6 Sol: a next-generation model
#585Earlier quoted context omitted.
Yeah, that's the point, right? With tool calling the LLM becomes code. So instead of asking it to write an accounting software, you can hire the LLM to be your accountant.
But you'd still need code if you need something done in a consistent way.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#586Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…
The company is valued like they broke open the grail, when in reality it's more like they bought a Cybertruck, got it stuck in the mud, and realized "You know what this thing does better than all other cars... shovel mud"
I'm shorting Cerebras with margin to virtually zero.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#587Earlier quoted context omitted.
> I can think of the tedious task of finding certain functionality within a codebase. I usually can't beat an AI agent harness at this task today. Yup, I remember "racing" the AIs to figure things out in codebases just a year ago. Today, I have no chance. Whether it is due to degraded reasoning capabilities on my part or better models, I don't know.
AI is always going to be able to write a grep statement faster and more accurately than a human
Re: Previewing GPT‑5.6 Sol: a next-generation model
#588Re: Previewing GPT‑5.6 Sol: a next-generation model
#589Earlier quoted context omitted.
Wow.. what?! How is this so fast?! Where can I read more?
Funnily enough, pasting your comment straight into Jimmy leads to a... Funnily suboptimal answer that does not answer the question. As someone else already contributed, this is driven by a Canadian startup taalas that basically makes chips that are llms, so everything is very fast but also, baked into the chip. Once this kind of stuff is a commodity in like 10 years, our world will be very, very different.
This chip would still be 272mm2 on N2 which is an eye-watering $30k/wafer and bigger than a 9950x or Nvidia 5070.
This just isn't feasible. Some of the latest-gen LLMs seem to have 5-10T parameters or about 1000x more. I don't know that taping out just one chip makes economic sense let alone the 300-1000 chips required for a cutting-edge model. Things like continuing education so your model knows about the latest NPM packages or world news is super important, but seems like it would require new chips.
There are a TON of uses for an 8B parameter models on the edge, but this is WAY too big to put on the edge of anything. Something like a 10mm2 100m parameter voice model might be feasible on the edge, but only for expensive devices, but most of those are TSMC 28nm (up to 29MTr/mm2) or GF FDX22 (up to 40MTR/mm2) which would increase the AI chip to the point where it would absolutely dominate the BOM.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#590Easily the most interesting part of this announcement is buried in the second to last paragraph: "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity." 750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new vers…
https://mikeveerman.github.io/tokenspeed/?rate=750&mode=thin... This is what 750tps looks like, I guess.