Live data from Hacker News

Accelerating GPT-5.6 Sol Ultrafast

cerebras.ai

221–230 of 295 posts

Re: Accelerating GPT-5.6 Sol Ultrafast

#221
post #169

Earlier quoted context omitted.

We’ve literally saved tens of millions of dollars already (no exaggeration! already 8 digits) by switching to Luna for many workloads at my company. The amount of workloads we can shift with an advisor model pattern continues to grow. It’s seriously amazing.

Luna came out about a month ago, you're saying that the cost saving from switching to Luna has saved your company $20 000 000+ in 1 months spending on API usage?

OpenAI and Anthropic are believed to be making billions per month in revenue selling tokens to enterprises. It’s not that surprising for enterprise SaaS they would have individual customers making up ~1% of their total sales. Although, perhaps a bit more surprising those few customers are posting about it here.

Another interpretation would this is a counterfactual savings, like they previously paid $1M for y tokens, and now that tokens are cheaper they increased usage and paid $1M for 20*y tokens.

Re: Accelerating GPT-5.6 Sol Ultrafast

#222

Earlier quoted context omitted.

The stake in the side of cerebras has always been that the economics are pretty poor. Who knows if they will subsidizes it to mitigate sticker shock, but it's a safe assumption that it will be scarily expensive. However if you are in a "cost is no obstacle, speed is god" position, it will likely be pure magic.

Can anyone explain why Cerberus needs to be _fast_ instead of _cheap_? I don't think I understand why they aren't leveraging the increased speed to do batching to serve more customers at a "normal" tok/s. Is the limitation, even on cerberus, still that the cache can only serve so many concurrent sessions over time? Is there no scaling advantage? I genuinely do not understand how any of this works.

For us, it would be for SRE stuff. We have agents reviewing traces and logs, and inspecting system behavior daily - for non obvious problems, not surfacing in metrics. When we hit an issue, we use the /fast mode to triage, propose a fix and then build and deploy. It is trivial amount of money all things considered, and I'd happily authorize a 100x spend for when we have a prod outage on a mission critical service.

When you think about it, it would still be dirt cheap compared to normal way of doing things. In the old days, if you had an outage on a serious user facing system, you'd wake up people across various timezones, wake up their managers and scramble to find the root cause, identify a solution, brainstorm on possible side effects of a fix, and then rush to build it and deploy. This cycle would involve, sometimes, dozens of people, for, say, 10 man hours each. So lets make it 120 man hours per serious outage, and lets assume and average of $100 per hour - so, $12,000 per a serious outage fixed under a day, counting conservatively and not including the costs of the actual outage.

I'd guess the pricing for those ultrafast, very energy inefficient and hardware heavy models will be competing with that. Its going to be possible to get a fix out in 30 minutes, 10 of which will be tests, 5 will be the deploy, and the remaining 15 will be some unlucky guy trying to keep up with the super fast model throwing a 50 "load-bearing deferrals earning their keep" per minute :-)

The pricing on those things is competing with costs to run entire departments. I'd, for one, imagine offshore ops teams will be a thing of the past in under a year, since one gets way better initial response to anything from a model, given right setup, esp. on codebases that have been built from the ground up with agentic coding - so with good documentation and effective test coverage baked into repos.

Re: Accelerating GPT-5.6 Sol Ultrafast

#223

The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/ There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding

They're nearly certainly going to use it internally to speed up research that is serially bottlenecked. I would bet this is why they're interested in the Cerebras partnership more than everything else

Yeah, the main value of this for OAI is probaby to speedup RSI

Re: Accelerating GPT-5.6 Sol Ultrafast

#225

Earlier quoted context omitted.

I just finished waiting almost four hours for Fable to write 700 lines of code, based on my three paragraph prompt. Some speed on these harder tasks would definitely be welcome! It also spent almost 800k tokens on these lines…

I'm very curios what the code is doing.

It is a normalisation procedure for inductive and coinductive data types.

Basically, the code is a function which takes in an expression where you can use generic data structures as variables and then some specific data structures, and it plugs them in for the variables. It then computes the structure of the resulting data type.

So, admittedly not a trivial task – hence the choice of Fable as the model. Also, this would have taken me few days to do by hand! So, we are living in the future! But one could always wish for more speed and more intelligence.

Re: Accelerating GPT-5.6 Sol Ultrafast

#226

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

Isn't ultrafast just making hundreds of subagents?

Re: Accelerating GPT-5.6 Sol Ultrafast

#228

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

An irrational gripe of mine is how GPT uses 7× instead of 7x. I recognize that the former is the multiplication symbol, but I don't think it should be used that way.

Interesting I always thought that "x" (whichever) is used informally for multiplication and proper sign is dot: 3 · 4 = 12

It looks like there is a difference between English speaking languages and the rest in that regard.

Re: Accelerating GPT-5.6 Sol Ultrafast

#229
post #120

Earlier quoted context omitted.

I'm finding Luna suprisingly adequate for my work. I slept on it due to the benchmarks, but it's very fast and even on low reasoning I'm finding it more than adequate for "menial" work. (The speed is crucial for "interactive" work -- if a model is fast enough it goes from "async" to "real time", subjectively, which is a huge difference.) In fact, I'd say it's overqualified for the kind of work I'm doing, because it s…

Sol medium/high planner orchestrating -> Luna xhigh subagents doing implementation ...has been REALLY good for me. Even on xhigh, Luna is crazy cheap. Subjectively I'd say it's way better than Sonnet at a fraction of the cost. Luna xhigh can do some decently challenging things on its own, but when orchestrated by a model that is actually good like Sol, I am finding it very very nice.

How are you doing orchestration - using sol for plan mode in codex? Or some other pattern/harness?

Re: Accelerating GPT-5.6 Sol Ultrafast

#230
post #169

Earlier quoted context omitted.

We’ve literally saved tens of millions of dollars already (no exaggeration! already 8 digits) by switching to Luna for many workloads at my company. The amount of workloads we can shift with an advisor model pattern continues to grow. It’s seriously amazing.

what has a single company accomplished with tens of millions of token spend?

A pelican on a bike accurate to a subatomic level
Post reply on HN