Live data from Hacker News

Accelerating GPT-5.6 Sol Ultrafast

cerebras.ai

281–290 of 295 posts

Re: Accelerating GPT-5.6 Sol Ultrafast

#281
post #238

Earlier quoted context omitted.

In 5-10 years nice smartphones will be able to run ChatGPT (~gpt3-4) class models. A memory rich laptop (highend mac/framework) can run GPT-OSS:120b or full Gemma4 at very interactive speeds. High end phones can already run the smaller models at enough speed to be probably useful, especially for background/overnight photo tagging and curation and things like that.

In 5 years the models will probably be so much better and more compact that phones will be running models equivalent at least to Opus 4.6 if not Fable, at least within the areas they are tuned for (which probably won't include coding).

I'm really excited for this future. Both what you said and datadrivenangel.

The fact that Apple shipped a more than capable laptop for most of the population using a last generation iPhone chip is just mind blowing. Silicon advancements are going to allow this, and I think the global majority will catch up and make their own chips that compete or exceed western performance. Especially when the US is scared of science, rapidly divesting and defunding it.

Re: Accelerating GPT-5.6 Sol Ultrafast

#282
post #274
post #259

Earlier quoted context omitted.

> I thought they'd prefer their own solutions, but they frequently preferred the other model's solutions. I suspect that's downstream from their sycophancy.

Ha, maybe. Sample size was low but usually both models ended up preferring the same solution.

I had a similar situation set up where I would have two models come up with their own independent diagnoses and plans to fix problems.

Then I would have them each read the other's plan. But I would tell them each something like this, "I had a friend look at this too. He's smart, but in general you're smarter and more knowledgeable than him, so don't be afraid to say where he's wrong."

Re: Accelerating GPT-5.6 Sol Ultrafast

#283
post #27

Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance. Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfull…

I mean I still regularly use 5.3 spark (the cerebrus model) that comes with my sub to do rapid reviews of 5.6's work and it finds oodles of problems in about a minute.

Re: Accelerating GPT-5.6 Sol Ultrafast

#284
post #169

Earlier quoted context omitted.

I discovered yesterday that the “amazing thing that comes out of OpenAI” is Sol, due to its token efficiency. Dollar for tokens, Sol and Fable are the same price. However, Sol uses (literally: in testing) around 10-100x less output tokens compared to Fable for the same task. We run our frontier models nearly 24/7, so switching to Sol will save us around $500 per day. And, due to less guardrails, Sol also performed be…

We’ve literally saved tens of millions of dollars already (no exaggeration! already 8 digits) by switching to Luna for many workloads at my company. The amount of workloads we can shift with an advisor model pattern continues to grow. It’s seriously amazing.

[dead]

Re: Accelerating GPT-5.6 Sol Ultrafast

#285

Earlier quoted context omitted.

Strong disagree, there’s nothing better about an agent generating a 1000 lines in 1 second or 1 minute. I can’t read it that quickly anyway.

But with the 1 second model you can get to reading the code immediately, and with the 1 minute model you have to sit doing 'nothing' for a minute. Rinse and repeat for each change/iteration

Yea but then you're either spending 15min or 16min review code.

Re: Accelerating GPT-5.6 Sol Ultrafast

#286

Earlier quoted context omitted.

It is a normalisation procedure for inductive and coinductive data types. Basically, the code is a function which takes in an expression where you can use generic data structures as variables and then some specific data structures, and it plugs them in for the variables. It then computes the structure of the resulting data type. So, admittedly not a trivial task – hence the choice of Fable as the model. Also, this wo…

What language?

Agda. Fable and Opus both know it quite well now. They need some guidance to create good-looking code, but man, they know how to grind!

Re: Accelerating GPT-5.6 Sol Ultrafast

#287

Earlier quoted context omitted.

At 14,000t/s that's effectively a motor cortex for an android, you no longer need to train the robot to walk, it has a general idea for how to walk (baked into the 1b model), and then just corrects based on sensor input, in real time.

I still personally think that a heavy lean into MoE will be better for that sort of thing. Our brains are subdivided into large parts but I'm sure (and I'm not a brain scientist) that those parts can be subdivided even further into systems that run at various frequencies and latencies depending on what they're used for. I was thinking about it the other day actually. How our brains evolved structure. I imagine it was…

    like MoE with a billion "experts".
That seems promising to me too, although, the thing I've always read is that you can't make the "experts" too narrow. Even if you had a "coding expert" it has to know a lot more than coding - if you tell it to make an online store it needs to parse your language, understand the internet, what a "store" is in this context, etc.

I am not a primary source, probably not even a secondary or tertiary source, so take this with all the grains of salt.

Re: Accelerating GPT-5.6 Sol Ultrafast

#288

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

I discovered yesterday that the “amazing thing that comes out of OpenAI” is Sol, due to its token efficiency. Dollar for tokens, Sol and Fable are the same price. However, Sol uses (literally: in testing) around 10-100x less output tokens compared to Fable for the same task. We run our frontier models nearly 24/7, so switching to Sol will save us around $500 per day. And, due to less guardrails, Sol also performed be…

[flagged]

Re: Accelerating GPT-5.6 Sol Ultrafast

#289
post #274

Earlier quoted context omitted.

Ha, maybe. Sample size was low but usually both models ended up preferring the same solution.

I had a similar situation set up where I would have two models come up with their own independent diagnoses and plans to fix problems. Then I would have them each read the other's plan. But I would tell them each something like this, "I had a friend look at this too. He's smart, but in general you're smarter and more knowledgeable than him, so don't be afraid to say where he's wrong."

I even tell them to give me an adversarial review and tell me what's wrong with the approaches.

Re: Accelerating GPT-5.6 Sol Ultrafast

#290

The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/ There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding

They're nearly certainly going to use it internally to speed up research that is serially bottlenecked. I would bet this is why they're interested in the Cerebras partnership more than everything else

i would assume internally they have even better tps
Post reply on HN