Live data from Hacker News

Cerebras CS-4

cerebras.ai

251–260 of 281 posts

Re: Cerebras CS-4

#251

Earlier quoted context omitted.

1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations. 2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation. > If Amazon doesn't buy but Microsoft…

1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer. 2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens. 3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying…

1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices.

2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption.

3. Source? Gemini for example is 70% on TPU. I have yet to see data on Maia-300 beyond Microsoft PR. Remember it doesn’t have to be better it has to be more cost efficient. The overwhelming majority of inference spend does not care if token output is 20% slower if it is 50% cheaper.

> Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.

What Nvidia needs. Whether neoclouds can stay competitive vs hyperscalers paying Nvidia tax is far from clear, particularly when inference margins compress.

Re: Cerebras CS-4

#252
post #34

Earlier quoted context omitted.

God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.

So the answer is yes? Putting in a few of those racks for special tasks shouldn't break the power assumptions of a data center.

Distribution's always the problem - how much power actually gets delivered to each rack. 162 is ~an order of magnitude higher than normal, which means nothing in a standard data center is going to be built to deliver that kind of power to one rack.

Re: Cerebras CS-4

#254

Earlier quoted context omitted.

> GPT-OSS 120B which is nigh useless nowadays: I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent. I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence. Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-throug…

>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs

Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.

Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.

Re: Cerebras CS-4

#255

Earlier quoted context omitted.

1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer. 2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens. 3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying…

1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices. 2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption. 3. Source? Gemi…

1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.

2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be?

3. Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can. The biggest tell on how Nvidia is doing is that their share in inference has increased despite the increase in competition: https://archive.md/CKP0N. So while competition is getting bigger and bigger because the overall pie is getting exponentially bigger, Nvidia's growth is still higher than average.

Some other sources:

https://www.businessinsider.com/amazon-nvidia-aws-ai-chip-do...

https://www.businessinsider.com/startups-amazon-ai-chips-les...

Re: Cerebras CS-4

#257

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

A design that bakes the architecture into silicon would be 10x faster, and imagine a version that does all the multiplication ops using single log-amp addition versus dozens of transistors to cut down the amount of silicon used by 50x. The ceiling for AI optimized hardware is extremely high.

Stack on top of that the fact that diffusion based models like the ones made by Inception Labs are far faster and more efficient than autoregressive LLMs and have an even higher ceiling of optimization (single step path prediction via model distillation versus 50 step denoise is currently an active area for image diffusion)

The human brain is soon neither going to be more powerful nor energy efficient than the stuff we use to run AI.

Re: Cerebras CS-4

#258
post #236

Earlier quoted context omitted.

> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers tha…

Coding subs are good when they promote usage and adoption of your models in enterprises at API rates. Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact. Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.

> Should NVIDIA do a coding subscription too?

Yes, obviously! Well maybe not a subscription but definitely an inference service.

https://build.nvidia.com/

https://resources.nvidia.com/en-us-inference-infrastructure/...

https://www.nvidia.com/en-us/data-center/dgx-cloud-lepton/

In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of their own hardware (across a bunch of 3rd party models) on a practical level and gain whatever insights or details might be relevant to pass on to other hardware and software teams.

Re: Cerebras CS-4

#259
post #14

Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 p…

Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling. There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but…

Is this article not exactly about ASICs built for LLM inference and training acceleration?

Re: Cerebras CS-4

#260

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

The 500x efficiency gain makes their approach a no brainer. Just make a new chip every 6 months, you still win.

Yes, power efficiency may be the largest benefit of these chips actually - especially if the projections are true that the US and other countries simply aren't able to ramp up power generation to meet forecasted datacenter demand.

There's a bit of a ticking time bomb there, something that a Taalas-like architecture can clearly resolve.

Post reply on HN