Live data from Hacker News

GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

cerebras.ai

21–30 of 31 posts

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#21
post #8

It’s an absolute beast. I run it via OpenRouter, where I have Groq and Cerebras as the providers. Cheap enough as to be almost free, strong performance, and lightning fast.

Cheap enough for now, but of all the companies selling inference at a loss, Cerebras and Groq are probably losing the most per token. Their hardware is ungodly expensive and its reliance on huge amounts of SRAM bottlenecks how much cheaper it can get, since SRAM density is improving at a snails pace at this point.

You're pointing out a bunch of high capex costs (hardware, SRAM), but then concluding that their opEx is greater than their revenue on a per unit basis. Are they really losing money on every token? It seems that using hardware acceleration would decrease inference costs and they could make it up on unit economics over time.

But I'm just reasoning from first principles. I don't have any specific data about them.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#22

Earlier quoted context omitted.

Well if they give it out for free (aka they pay for it), asking you to register is a reasonable ask. It's not a public service funded by taxpayers.

> Well if they give it out for free (aka they pay for it), asking you to register is a reasonable ask They have other options... rate limiting, serving (more) quantized to non-registered etc. etc.

Those options are still not free. And giving a degraded version of your product to free users is a bad way to acquire clients.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#23
post #8

It’s an absolute beast. I run it via OpenRouter, where I have Groq and Cerebras as the providers. Cheap enough as to be almost free, strong performance, and lightning fast.

Cheap enough for now, but of all the companies selling inference at a loss, Cerebras and Groq are probably losing the most per token. Their hardware is ungodly expensive and its reliance on huge amounts of SRAM bottlenecks how much cheaper it can get, since SRAM density is improving at a snails pace at this point.

Switching costs are low, so if that happens we’ll just switch.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#25
post #8

Earlier quoted context omitted.

Cheap enough for now, but of all the companies selling inference at a loss, Cerebras and Groq are probably losing the most per token. Their hardware is ungodly expensive and its reliance on huge amounts of SRAM bottlenecks how much cheaper it can get, since SRAM density is improving at a snails pace at this point.

You're pointing out a bunch of high capex costs (hardware, SRAM), but then concluding that their opEx is greater than their revenue on a per unit basis. Are they really losing money on every token? It seems that using hardware acceleration would decrease inference costs and they could make it up on unit economics over time. But I'm just reasoning from first principles. I don't have any specific data about them.

  It seems that using hardware acceleration would decrease inference costs and they could make it up on unit economics over time.
Nvidia GPUs are accelerators too. The reason they can do this so fast is because they're storing entire models in SRAM.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#27
post #2

I absolutely hate it, when a website says "try this" and after you went through the trouble of weiting something comes up with a sign up link first. Makes me leave instantly to never come back.

Headline at the top of the Cerebras page linked to by the OP "Cerebras Raises $1.1B Series G at $8.1B Valuation" . If you're going after the AI money gravy train then you need to wave the "we have $n registered users" carrot on your PPT slides for the investors because registered user == monetization opportunity . I'm not defending it. I hate being forced to register for shit when I just want to try it or use the fre…

Right, being proud of your money making is not something I consider a consumer focused product unless that customer is other moneyseeking orga, which like cancer, often ends up in a bubble.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#28

Earlier quoted context omitted.

Not doubting you but anything to back that up? Happy enough to burn VC money until someone shows up who can run it without losing money, either way.

They’ve filed a S1 [1] last year when attempting to go public. It showed something like a $60M+ loss for the first 6 months of 2024. The IPO didn’t happen because the CEO’s past included some financial missteps and the banks didn’t want to deal with this. At the time the majority of their revenue came from a single source in Abu Dhabi, as well [1] https://www.sec.gov/Archives/edgar/data/2021728/000162828024...

> the majority of their revenue came from a single source in Abu Dhabi, as well

I live in UAE, whose continuing enthusiasm in AI investment stretches well beyond short-term profit, so having AD on-board seems like a plus not a minus. I'm sure there are specific exceptions, but generally Emirati money has seemed like smart money.

Re: GPT-OSS 120B Runs at 3000 tokens/sec on Cerebras

#30

Earlier quoted context omitted.

You're pointing out a bunch of high capex costs (hardware, SRAM), but then concluding that their opEx is greater than their revenue on a per unit basis. Are they really losing money on every token? It seems that using hardware acceleration would decrease inference costs and they could make it up on unit economics over time. But I'm just reasoning from first principles. I don't have any specific data about them.

It seems that using hardware acceleration would decrease inference costs and they could make it up on unit economics over time. Nvidia GPUs are accelerators too. The reason they can do this so fast is because they're storing entire models in SRAM.

There are degrees of acceleration. My understanding, limited as it is, is that groq and cerebras are using highly optimized acceleration to achieve their token generation rates, far beyond that in a regular GPU, and this leads to lower costs per token.

Is this incorrect?

Post reply on HN