Earlier quoted context omitted.
I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence
Cerebras CS-4
111–120 of 281 posts
Re: Cerebras CS-4
#112Earlier quoted context omitted.
not only that, but I was so happy with their GLM 4.8 that they got rid of yesterday :(
What do they do with old ones? Their hardware physically can't run other models right?
Re: Cerebras CS-4
#113Earlier quoted context omitted.
I read it as it is impressive because smaller models 2.5T are squeezing similar returns as 10T models despite being 1/4th size not that there beyond 2T today the number or parameters do not have much meaning
or the latest qwen3.8 27B doing so well at ~1/100 the size of K3
Re: Cerebras CS-4
#114OpenAI needs to immediately move to acquire Cerebras. Nvidia's extreme margin is the opportunity for OpenAI's cost reduction. Buying Cerebras would pay for itself and they should take all of its future production (after filling required contracts). Right now China's models have no silicon moat. Cerebras as a drastic speed-up / cost-reduction potential, can assist in building a competitive moat. And every time a Cereb…
Do you know about Jalapeno?
OpenAI is partnering with Cerebras while simultaneously investing in their own silicon play. Hedged bets.
After sitting thru their keynote today, it makes sense. The main throughput speedups they tout are an obvious evolution of the GPU that all companies will be building in the next year. Wafer-scale interconnected memory and compute is just going to beat out mountains of network cabling any day on both cost and performance metrics.
Re: Cerebras CS-4
#115Earlier quoted context omitted.
or the latest qwen3.8 27B doing so well at ~1/100 the size of K3
What about general knowledge you can get out of it before hallucinations start?
Re: Cerebras CS-4
#116Re: Cerebras CS-4
#117Earlier quoted context omitted.
or the latest qwen3.8 27B doing so well at ~1/100 the size of K3
What about general knowledge you can get out of it before hallucinations start?
I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data
Re: Cerebras CS-4
#118Earlier quoted context omitted.
save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence
I wonder why they removed DeepSWE from their incorporates evaluations
It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
Re: Cerebras CS-4
#119Earlier quoted context omitted.
Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
You just proved that AI cannot currently do that
Re: Cerebras CS-4
#120Earlier quoted context omitted.
Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…