Live data from Hacker News

Cerebras Inference: AI at Instant Speed

cerebras.ai

1–10 of 75 posts

Re: Cerebras Inference: AI at Instant Speed

#4
post #2

This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).

"wafer scale" inference using 44GB of straight up SRAM per chip does not exactly sound "smaller and cheaper" to me. Just optimized.

Re: Cerebras Inference: AI at Instant Speed

#5
post #2

This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).

Cerebras' Wafer Scale Engine is the opposite of small and cheap.

https://cerebras.ai/product-chip/

Re: Cerebras Inference: AI at Instant Speed

#6
post #4
post #2

This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).

"wafer scale" inference using 44GB of straight up SRAM per chip does not exactly sound "smaller and cheaper" to me. Just optimized.

Cheaper on a per token basis.

Re: Cerebras Inference: AI at Instant Speed

#8
post #2

This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).

Cerebras' Wafer Scale Engine is the opposite of small and cheap. https://cerebras.ai/product-chip/

I updated my comment to be clearer. I meant smaller and/or cheaper per token. In this case it's cheaper per token.

Re: Cerebras Inference: AI at Instant Speed

#9
post #6
post #4

Earlier quoted context omitted.

"wafer scale" inference using 44GB of straight up SRAM per chip does not exactly sound "smaller and cheaper" to me. Just optimized.

Cheaper on a per token basis.

Doubtful. SRAM is not cheap, and this is entirely about SRAM vs HBM.
Post reply on HN