Live data from Hacker News

Cerebras CS-4

cerebras.ai

261–270 of 281 posts

Re: Cerebras CS-4

#261

Earlier quoted context omitted.

I do not rely on any LLM of any size for general knowledge baked into the weights, they all hallucinate and that is the wrong way to hold them imo I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data

> I do not rely on any LLM of any size for general knowledge baked into the weights You have to rely on it to a certain level for agentic/coding work, presuming that's the general subject we're talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn't already know "what is" xterm.js and a bunch of its associated npm-related/node related software. If it…

for sure, there is a minimum size and knowledge base that is required to be useful

at the same time, search may find newer or better alternatives, and you can always specify specific technologies you want to use, I typically do this when starting a new project

Re: Cerebras CS-4

#262

Earlier quoted context omitted.

1. Not really, current valuations are priced for persistent 80%+ margins based on spot. If auxiliary hardware lasts longer (I.e. next gen GPU reusing the same shell) then that reduces supply pressure and spot prices. 2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption. 3. Source? Gemi…

1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument. 2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be? 3. Google, Amazon, Microsoft, Meta are all buyin…

> 1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument.

Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.

> Total consumption drives more demand for the already supply constrained hardware.

Demand is the wrong metric.

Only number that matters is whether AI attributable revenue will be sufficient to pay back enough of each successive capex vintage (e.g. 750B this year, 1T next year, 1.2T in 2028) so that hyperscalers and neoclouds can either self-fund or continue to issue debt as bond markets are already straining and tax-payer backed sovereign debt is providing a high baseline. Otherwise they downgrade capex projections and the bubble pops.

Expensive compute needs expensive inference to justify 30-40B/year/GW of compute. There are many reasons why frontier API pricing which is what the industry is based on may not persist. It is also almost certainly the case that 2026 is the worst year of supply and demand mismatch to allow for 80%+ margins. HBF next year has the potential to single handedly pop the DRAM spot bubble.

> Can AI hardware market go bust? Sure it can.

This is the bear thesis. It is not that AI will crash or be useless.

> But being early is the same as being wrong in the investment market. When do you predict the bust to be?

Q4 27-Q2 28 is when the bill becomes due at the latest. There are sufficient financial levers left to buy time without returns until then.

> Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can.

All of these companies have rock solid revenue streams and can easily swallow 500B of capex devaluation over time. Their buying of Nvidia today is not necessarily the indicator you are implying as there are strong competitive reasons to make the game more expensive for everyone else.

Re: Cerebras CS-4

#264
post #236

Earlier quoted context omitted.

> Why would they offer a coding subscription and start competing with some of their biggest customers; when they are capacity-bound and companies like OpenAI will take however many wafers that Cerebras sells to them? Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers tha…

Coding subs are good when they promote usage and adoption of your models in enterprises at API rates. Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact. Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.

> Should NVIDIA do a coding subscription too?

They sort of do? They offer free access to various versions of nemotron via multiple routing services.

Re: Cerebras CS-4

#265
post #103

Earlier quoted context omitted.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

If Fable is seriously around 10T and Kimi K3 sidles up to it at 2.4T That would be extremely surprising and a massive blunder by Anthropic in model design architecture ... which I highly doubt to be the case.

In real world comparisons K3 is somewhere between Sonnet and Opus. Fable is just a completely different (higher) level.

The long tail of tasks and queries is where you see the difference.

Re: Cerebras CS-4

#266

Earlier quoted context omitted.

1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument. 2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be? 3. Google, Amazon, Microsoft, Meta are all buyin…

> 1. The whole Burry argument is that AI hardware becomes obsolete faster. If aux hardware can be reused, that works against the argument. Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused. > Total consumption drives more demand for the already supply constrained hardware. Demand is t…

  Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
And why does he think depreciation is understated? It is because he thinks newer Nvidia GPUs will make older ones obsolete faster. Hence, my entire post.

The rest of your argument centers around whether AI growth will meet the cap ex expenses. I don't see anything new in it.

  HBF next year has the potential to single handedly pop the DRAM spot bubble.
I'll believe it when I see it. Jevons paradox will apply here again in my opinion. HBF does not replace HBM.

Re: Cerebras CS-4

#267

Earlier quoted context omitted.

Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.

I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.

Because comparing it to an average car is wrong. It isn't about fast/expensive. It is about fast/capacity.

A Ferrari can seat up to 4 people[0], which is about the same as the average car. Capacity doesn't change much.

Meanwhile, a bus/subway system is meant to support millions of people. Tokyo's metro has to support up to 37 million people. You can't do that with Ferraris.

[0]https://www.ferrari.com/en-EN/auto/ferrari-purosangue

Re: Cerebras CS-4

#269

Earlier quoted context omitted.

> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. .... :T Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter. https://taalas.com/products/ Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s. https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250 You're looking at…

In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.

Chiplet-style layouts could cut the etching quality requirements per chip down, but it still can't avoid the base cost for manufacturing silicon.

Re: Cerebras CS-4

#270
post #4

Earlier quoted context omitted.

On the plus side, lots of cheap servers to swoop up :)

Look at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(

No but dedicated server prices will likely drop or at least you'll get more for your money
Post reply on HN