Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

31–40 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#32

OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)

Came here to point this out, though not as pithily :D

Really, really bad mark on whoever is in charge of their web marketing. Images should never look that bad, not even in support, but definitely not in marketing.

edit: so this post is more useful, 4k res using Edge browser

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#33

OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)

Last time I viewed it, I believe it wasn't blurry. Perhaps to scale the traffic the images are now displayed in lower quality?

But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models.

Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#34

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Serious question: is it typical to describe client-side computing as "on the edge"?

I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#35
post #24

FYI: Cerebras's nodes are very different than your typical Nvidia training nodes: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-... Each individual "chip" has 40GB of SRAM vs ~76MB for the Nvidia H100, and networked pools of external RAM, SSDs and such. Thats why the training architecture is so different.

https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.

Yeah, and that doesn't even do the nutty IO on these things justice.

A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#36

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Serious question: is it typical to describe client-side computing as "on the edge"? I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.

Yes, "edge computing" can refer to both computing done as close to the user as possible geographically, or even on the device itself. If someone says "I wanna do edge computing" it's not clear enough to know if they just want to have servers they control as close to the user as possible, or do the computing on the device itself. I think Apple would say "edge computing" is on the actual device while CloudFlare would say "edge computing" is on their infrastructure, but distributed to be physically closer to the end user.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#37
post #24

Earlier quoted context omitted.

https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.

Yeah, and that doesn't even do the nutty IO on these things justice. A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.

It’s unbelievable stuff. Does anyone know how much a single box costs? They are selling them it looks like.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#38
Here are the zero-shot accuracy numbers posted in the Huggingface evaluations for Cerebras-GPT 13B vs. the results of LLaMa 13B in their paper:

    Model              BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA
    LLaMa 13B          78.1  80.1 50.4 79.2      73         74.8  52.7  56.4
    Cerebras-GPT 13B   -     76.6 -    51.3      64.6       71.4  36.7  28.6

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#39
post #30

I wonder what led to such a gap between llama 7b and Cerebras 13b. I hope they discuss it in the paper.

Is there a benchmark comparing the two that I missed? Edit: The huggingface page has 0-shot benchmarks which you can compare against the llama paper https://huggingface.co/cerebras/Cerebras-GPT-13B https://arxiv.org/pdf/2302.13971.pdf

I'm on mobile and struggled to compare these two tables properly. Would you mind posting a summary of your findings?

Here are some values but I don't know what they mean. LLama 60B on the left, Cerebras 13B on the right.

PiQA: 82.8 / 76.6 WinoGrade: 77.0 / 64.6 ARC-e: 78.9 / 71.4

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#40

This “AI spring” is really snowballing with the crazy nouns and terminology. Alpaca, llama and now chinchilla??

As mentioned, chinchilla is not part of this trend, and chinchillas are rodents. Alpacas and llamas are South American camelids (animals related to camels). So if additional names are needed, I would expect them to be vicuña and guanaco, as they are also in the group.

I think the relevant category is "Adorable Fuzzy Critters of the Andes". See also https://en.wikipedia.org/wiki/Spectacled_bear
Post reply on HN