Live data from Hacker News

Cerebras CS-4

cerebras.ai

151–160 of 281 posts

Re: Cerebras CS-4

#151
post #134

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too. Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months.

....

:T

Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter.

https://taalas.com/products/

Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s.

https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250

You're looking at

1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or

2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).

Could happen, but it's a long shot for a market that could be satiated by specialized accelerators.

Re: Cerebras CS-4

#152

Earlier quoted context omitted.

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

The 500x efficiency gain makes their approach a no brainer. Just make a new chip every 6 months, you still win.

Re: Cerebras CS-4

#153

Earlier quoted context omitted.

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?

GPU cost is about half of datacenter cost. Other half is cooling, power and networking.

Re: Cerebras CS-4

#154
post #80

Earlier quoted context omitted.

Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the soft…

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

You still need experts to know what is good and what is not.

Re: Cerebras CS-4

#155

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.

Re: Cerebras CS-4

#156

Earlier quoted context omitted.

TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.

Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.

So far, at least by Openrouter's weekly numbers, there doesn't seem to be a satisfied limit.

https://openrouter.ai/rankings#top-models

And their market share sits at around 16-20%.

At 75.3 trillion tokens for the week ending 10 Aug 2026, that means that up to 450 trillion tokens were plausibly demanded by the whole market for that week.

My take: At max saturation, each person on earth could have their demands satiated by an average of 16 agents running concurrently. Sometimes more, often times less, but the average would likely be at 16.

At 200 tokens/second for each agent, that would mean 15.48288 quintillion tokens per week.

We're currently at about 0.00290643601% of the calculated demand ceiling.

Even if the demand limit per person is just 1 agent at 50 tokens/second, the current demand's still 0.186011905% of the theoretical ceiling.

Re: Cerebras CS-4

#157

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

I thought your numbers must be wrong.

So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.

So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

Re: Cerebras CS-4

#158

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

Press doubt. Single GPU? Maybe. MultiGPU behemoths like NVL144 and NVL576? I don't think so.

NVLink is at gen9. they had a lot of teething problems and can codesign the hardware and software.

in the name of openness (AMD's only """weapon"""), the UALink spec is a hodgepodge of corporate opinions with very different implementations (looking at you, Broadcom). at spec version 1 (in hardware).

I wish them good luck as I really like AMD, but they compete no more on this than Lambo vs Bugatti.

Re: Cerebras CS-4

#159

Earlier quoted context omitted.

It's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day. Yes, please!

Most people probably don't care about nginx performance. It shouldn't be your bottleneck unless you serve massive amounts of static data.

Ok, how about postgres?
Post reply on HN