Live data from Hacker News

Cerebras CS-4

cerebras.ai

201–210 of 281 posts

Re: Cerebras CS-4

#201
post #120

Earlier quoted context omitted.

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

I just want to but hardware so I can run a model at home that is fast. I don't see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...

At 250w when it's working (I understand), and it works for tiny amounts of time per query...

Re: Cerebras CS-4

#202
Interestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.

Re: Cerebras CS-4

#203

Earlier quoted context omitted.

Hence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).

Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web…

> It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.

I distinctly remember 32-bit/33 MHz PCI accelerator cards for SSL being a real thing (for use on OpenBSD or FreeBSD), in an era when something like a single core 700 MHz Pentium 3 1U system was a relatively powerful individual bare metal httpd box.

http://www.aster.si/partnerji/compaq/atalla/axl200.html

The CPU load of doing a lot of SSL purely in software was a problem in terms of scaling things up, so this was one attempt at a (very short lived) solution. Note that this predated TLS1.0.

Re: Cerebras CS-4

#204

Earlier quoted context omitted.

It's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day. Yes, please!

Most people probably don't care about nginx performance. It shouldn't be your bottleneck unless you serve massive amounts of static data.

In the case of needs to process natural language, instead, massive efficiency (esp. time) can be a game changer. It's like "you have two years to complete the project" vs "you have two hours to complete the project": if you can squeeze that "two years worth" into a negligible delay, it's a game changer.

Re: Cerebras CS-4

#205
post #80

Earlier quoted context omitted.

Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the soft…

Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.

i haven't seen it. see the recent browser attempts.

Re: Cerebras CS-4

#206

AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

Like NVIDIA bought Groq, AMD might do well buying Cerebras.

AMD did enter into an agreement to buy Taalas, which is speculated [1] will be used to augment their Helios offering.

1. https://www.youtube.com/watch?v=3MKRjt59hh4&pp=0gcJCRMMAYcqI...

Re: Cerebras CS-4

#207
post #87

Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.

Nvidia is at this time a pretty well run company tech wise. They are going to keep iterating on the inferencing hardware stack over the next five years too.

The only way I see in which Nvidia can catch up is by buying Cerebras.

Re: Cerebras CS-4

#209
post #103

Earlier quoted context omitted.

Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T. Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai... That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx Both Grok and Bytedance are training 10T models.

The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac. Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.

The open models don't really match Opus.

For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention.

I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.

Re: Cerebras CS-4

#210

Earlier quoted context omitted.

By the way, this is the same argument that Michael Burry used to short Nvidia. He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0] The logic is fundamentally flawed in my opinion. Let'…

1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations. 2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation. > If Amazon doesn't buy but Microsoft…

1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.

2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.

3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.

Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.

Post reply on HN