Live data from Hacker News

Accelerating GPT-5.6 Sol Ultrafast

cerebras.ai

111–120 of 295 posts

Re: Accelerating GPT-5.6 Sol Ultrafast

#111
post #81

Good news for Intel and AMD. Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.

why intel and amd ? these are cerebras wafers? i know people are joking about the sol ultrafast prices (its unlikely to be accessible for average joes) but this shows scaling wafer cores works for inference boost which makes me very excited, sol ultrafast will be as slow as it will get if that makes sense. at these token speeds , we will see a much deeper economic impact.

Right now the bottle neck is not the CPU, so people aren‘t spending big $ on them. But with this ultra fast mode, CPU becomes a bigger part of the bottleneck and thus Intel and AMD can charge more $$$.

Re: Accelerating GPT-5.6 Sol Ultrafast

#112
post #77
post #72

Earlier quoted context omitted.

Output from Cerebras with GPT model is 750 tokens per second. Don’t blink. (Chatjimmy has 14,200 TPS.)

ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.

The knowledge of ChatJimmy is terrible. Even Qwen on my iPhone is better.

Re: Accelerating GPT-5.6 Sol Ultrafast

#114
post #44

Earlier quoted context omitted.

No quality compromise/degradation is something I have had this industry, including especially OpenAI, claim multiple times in the past and I have more than once been able to verify that it was in fact not the case. Examples being gpt-3.5-turbo vs text-davinci-003, GPT-4-Turbo and all the other post training checkpoints they had under one name (which was a major bug bear for me back then witnessing degradations with n…

No one claimed gpt-3.5-turbo doesn't have any degradation over davinci-003. In fact it was quite obvious that gpt-3.5 had way less knowledge but more post trained to be helpful.

That quite strong "no one" surprised me so I checked and looking through a few blog posts from back then, they did advertise gpt-3.5-turbo as a straight up improvement and, once text-davinci-003 was to be deprecated, the instruct tuned variant as the drop in replacement [0]. If anything, they did not just promise similar performance but actually an improvement ("our best model") when compared to text-davinci-003:

> It’s also our best model for many non-chat use cases—we’ve seen early testers migrate from text-davinci-003 to gpt-3.5-turbo with only a small amount of adjustment needed to their prompts.

That's why I still remember this so well, they claimed one model to be their best and a straight up drop-in during deprecation when in my (back then even more amateurish then today) testing this was plainly not the case. A model cannot be "best" if it's measurably worse in many situations, then what was still available at the time.

[0] https://openai.com/index/gpt-4-api-general-availability/

[1] https://openai.com/index/introducing-chatgpt-and-whisper-api...

Re: Accelerating GPT-5.6 Sol Ultrafast

#115

Whoa. This looks both powerful and expensive. My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference. Within labs, I've heard spend is already far beyond this per developer.

not sure how "top developers" are defined here, but there is huge diminishing return curve starts kicking in after $200/month price point for typical eng work.

If you're talking Opus pricing, it's more like $200/day.

Re: Accelerating GPT-5.6 Sol Ultrafast

#116

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

An irrational gripe of mine is how GPT uses 7× instead of 7x.

I recognize that the former is the multiplication symbol, but I don't think it should be used that way.

Re: Accelerating GPT-5.6 Sol Ultrafast

#117

I don’t know if this is that useful for coding. In some autonomous world, where no one check the code and the agent can just spend 10X more time checking its work and leading to better results, yes maybe it is useful. But if humans need to check its work, then 10X speed doesn’t really matter I guess.

Iterations get much faster, which makes keeping attention much easier, and hence the work is easier to review.

Re: Accelerating GPT-5.6 Sol Ultrafast

#118

Whoa. This looks both powerful and expensive. My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference. Within labs, I've heard spend is already far beyond this per developer.

Isn't TCO lower with Cerebras chips compared to Nvidia? Theoretically, most developers should eventually be running on Ultrafast.

Re: Accelerating GPT-5.6 Sol Ultrafast

#119
post #58

Fast mode is already 1.5 times faster and 2x more expensive in the Codex subscription plan. If this thing is 14 times faster, then I can imagine running out of my quota in one session.

There is zero chance this will be offered to subscription users.

It will eventually. Right now everyone is stuck on the equivalent of dialup.

Re: Accelerating GPT-5.6 Sol Ultrafast

#120

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

I'm finding Luna suprisingly adequate for my work. I slept on it due to the benchmarks, but it's very fast and even on low reasoning I'm finding it more than adequate for "menial" work. (The speed is crucial for "interactive" work -- if a model is fast enough it goes from "async" to "real time", subjectively, which is a huge difference.)

In fact, I'd say it's overqualified for the kind of work I'm doing, because it spends >half the time verifying trivial changes (and the verification isn't as helpful as you'd expect, even with bigger models).

Maybe I can prompt it to be less aggressive about that (the new GPT models do it even without prompting).

Anyway, Ultrafast Luna would be amazing, though I strongly doubt they can offer Cerebras at anything approaching the current prices. Now we wait for Moore's Law? :)

Post reply on HN