Cerebras Inference: AI at Instant Speed
cerebras.ai
Cerebras Inference: AI at Instant Speed
1–10 of 75 posts
Re: Cerebras Inference: AI at Instant Speed
#2Re: Cerebras Inference: AI at Instant Speed
#3Re: Cerebras Inference: AI at Instant Speed
#4This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).
Re: Cerebras Inference: AI at Instant Speed
#5This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).
Re: Cerebras Inference: AI at Instant Speed
#6This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).
"wafer scale" inference using 44GB of straight up SRAM per chip does not exactly sound "smaller and cheaper" to me. Just optimized.
Re: Cerebras Inference: AI at Instant Speed
#7Re: Cerebras Inference: AI at Instant Speed
#8This is where we always assumed the industry was going. Expensive GPUs are great for training, but inference is getting so optimized that it will run on smaller and/or cheaper processors (per token).
Cerebras' Wafer Scale Engine is the opposite of small and cheap. https://cerebras.ai/product-chip/