Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
31–40 of 231 posts
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#32OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)
Really, really bad mark on whoever is in charge of their web marketing. Images should never look that bad, not even in support, but definitely not in marketing.
edit: so this post is more useful, 4k res using Edge browser
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#33OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)
But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models.
Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#34This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…
I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#35FYI: Cerebras's nodes are very different than your typical Nvidia training nodes: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-... Each individual "chip" has 40GB of SRAM vs ~76MB for the Nvidia H100, and networked pools of external RAM, SSDs and such. Thats why the training architecture is so different.
https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.
A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#36This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…
Serious question: is it typical to describe client-side computing as "on the edge"? I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#37Earlier quoted context omitted.
https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.
Yeah, and that doesn't even do the nutty IO on these things justice. A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#38 Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA
LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4
Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#39I wonder what led to such a gap between llama 7b and Cerebras 13b. I hope they discuss it in the paper.
Is there a benchmark comparing the two that I missed? Edit: The huggingface page has 0-shot benchmarks which you can compare against the llama paper https://huggingface.co/cerebras/Cerebras-GPT-13B https://arxiv.org/pdf/2302.13971.pdf
Here are some values but I don't know what they mean. LLama 60B on the left, Cerebras 13B on the right.
PiQA: 82.8 / 76.6 WinoGrade: 77.0 / 64.6 ARC-e: 78.9 / 71.4
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#40This “AI spring” is really snowballing with the crazy nouns and terminology. Alpaca, llama and now chinchilla??
As mentioned, chinchilla is not part of this trend, and chinchillas are rodents. Alpacas and llamas are South American camelids (animals related to camels). So if additional names are needed, I would expect them to be vicuña and guanaco, as they are also in the group.