Live data from Hacker News

Accelerating GPT-5.6 Sol Ultrafast

cerebras.ai

121–130 of 295 posts

Re: Accelerating GPT-5.6 Sol Ultrafast

#121

Whoa. This looks both powerful and expensive. My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference. Within labs, I've heard spend is already far beyond this per developer.

not sure how "top developers" are defined here, but there is huge diminishing return curve starts kicking in after $200/month price point for typical eng work.

I agree with the diminishing returns on spend, but worth noting that when on an Enterprise seat and paying API rates, I'd say that you can easily spend above 200/mo before seeing the curve begin to flatten

Obviously there are a ton of ways to spend money / tokens and people have different levels of experience that will put this ceiling at very different levels for different people.

Re: Accelerating GPT-5.6 Sol Ultrafast

#123

People underestimate the importance of speed on quality of thought, because people underestimate just how much quality is a result of simple iteration. When an LLM thinks, it typically just makes one pass. It outputs tokens from top to bottom, beginning to end, and then it's done. But when people think, especially strong thinkers, we typically iterate and revise our thoughts on the fly. We do numerous passes. We stop…

Just putting more details in, and giving it ways to check itself brings massive improvements. I often ask model to set up test for the problem before actually trying to solve it and it improves it a lot, both in how much babysitting is required (if it can test it itself quickly it goes faster), and the fact the context now contains more detailed description of the problem that came up when making tests.

Turns out TDD is far better for robots than humans, who knew

Re: Accelerating GPT-5.6 Sol Ultrafast

#124
post #72

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

Output from Cerebras with GPT model is 750 tokens per second. Don’t blink. (Chatjimmy has 14,200 TPS.)

Unfortunately AMD bought them, so I don't think we will get to see another release from them.

Re: Accelerating GPT-5.6 Sol Ultrafast

#125

People underestimate the importance of speed on quality of thought, because people underestimate just how much quality is a result of simple iteration. When an LLM thinks, it typically just makes one pass. It outputs tokens from top to bottom, beginning to end, and then it's done. But when people think, especially strong thinkers, we typically iterate and revise our thoughts on the fly. We do numerous passes. We stop…

This is foundationally similar to a lesson I've found from years of pre-LLM software development: builds that turn around in 500ms instead of 5 minutes fundamentally change the way you can work as a software engineer. I think a lot of the same applies to working with LLMs. I'm not sure, though, what the path from where we are today to some future state of high speed token abundance actually looks like...I think there's plenty of chance that we see the bubble pop in the near term over token costs and complexities of today's infrastructure, then some totally different landscape of LLM use in 5-10 years that looks quite unlike what we have today, similar to how waiting 30 minutes for an MP3 of a single song to download on a 28k modem in 1999 seems quaint today.

Re: Accelerating GPT-5.6 Sol Ultrafast

#126
post #26
post #6

> Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode. Awesome work. I'm personally very excited for faster models/inference. I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast…

What do you need speed for? That's a genuine question, I feel like the limiting factor already is my creativity, attention span and budget. And I'm not even yet optimizing cost by batching things like review to slow local models over night, or schedule tasks to take full advantage of my subscriptions.

There are speed thresholds that open up new use cases.

Imagine speeding up current agents 10x, you switch from directing agents to pair-vibing on the fly.

Speed them up 10x more, and you get a SOTA model capable of analyzing and rethinking your entire file in between your key strokes. That would make for one hell of an autocomplete.

Pivot over application, going from coding to anything else, and this can easily give computers features previously impossible to make. In video games, fully general characters reacting realistically to arbitrary dynamic situations. In "serious" apps, interactive work with a system that understands your goals and adapts to you on the fly. Hell, even an OS that can tell you "hey, the data you're obviously looking for is in the tab over there, now highlighted".

And that's just tip of the iceberg. I'd personally love to explore the possibilities.

Re: Accelerating GPT-5.6 Sol Ultrafast

#127
post #120

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

I'm finding Luna suprisingly adequate for my work. I slept on it due to the benchmarks, but it's very fast and even on low reasoning I'm finding it more than adequate for "menial" work. (The speed is crucial for "interactive" work -- if a model is fast enough it goes from "async" to "real time", subjectively, which is a huge difference.) In fact, I'd say it's overqualified for the kind of work I'm doing, because it s…

I feel odd to use these models, because it feels like a faster model doesn't feel that much faster if it spends reading files, making edits and running checks.

It feels too situational.

Re: Accelerating GPT-5.6 Sol Ultrafast

#129
post #27

Unless I have read over it, besides the animation in the intelligence vs speed graph which only mentions internal data and not whether they truly reran the AA suite, there is no actually solid statement on the important aspect of performance. Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfull…

This is what Cerebras does- take other people's models and run them very very fast.

Re: Accelerating GPT-5.6 Sol Ultrafast

#130

I've been waiting so long for something amazing to come out of the OpenAI and Cerebras collaboration. > In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a s…

I discovered yesterday that the “amazing thing that comes out of OpenAI” is Sol, due to its token efficiency.

Dollar for tokens, Sol and Fable are the same price.

However, Sol uses (literally: in testing) around 10-100x less output tokens compared to Fable for the same task.

We run our frontier models nearly 24/7, so switching to Sol will save us around $500 per day.

And, due to less guardrails, Sol also performed better, and we lost less tokens due to guardrails shutting down sessions (I feel like it’s illegal to take $50 of someone’s token money and then shut down a session with guardrails before they get an answer, and yet Anthropic do it to us constantly… either take our money and commit, or trigger the guardrails immediately)

Post reply on HN