Yay there will be a paper let's gooooooo!
Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
41–50 of 231 posts
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#42This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…
Serious question: is it typical to describe client-side computing as "on the edge"? I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.
What I find more interesting is that in the classic "close network proximity", some parts of the world may not have benefited as much from that trend since the closest nodes of a global delivery network could be several countries away.
[0] https://github.com/ggerganov/llama.cpp/discussions/205
[1] https://medium.com/sort-of-like-a-tech-diary/consumer-ai-is-...
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#43OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)
https://www.cerebras.net/wp-content/uploads/2023/03/Downstre...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
EDIT: Looks like it scores better with less training - up until it matches GPT-J/Pythia/OPT and doesn't appear to have much benefit. It maybe scores slightly better then GPT-J which is pretty "eh", I'm not sure if GPT-J level performance is really useful for anything? NeoX 20B outperforms it in everything if you don't care about the amount of training needed.
Does the better performance for less training matter if that benefit only applies when it's only performing a lot worse then GPT-J? It appears to lose it's scaling benefits before the performance is interesting enough to matter?
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#44Here are the zero-shot accuracy numbers posted in the Huggingface evaluations for Cerebras-GPT 13B vs. the results of LLaMa 13B in their paper: Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4 Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#45Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#46Here are the zero-shot accuracy numbers posted in the Huggingface evaluations for Cerebras-GPT 13B vs. the results of LLaMa 13B in their paper: Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4 Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6
I guess it's something. It still goes to show how far open models are behind the proprietary SOTA.
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#47Summary: This is a company that makes AI accelerator ICs. They reimplemented Chinchilla and released the model weights under a permissive license.
Which makes me wonder if Nvidia is doing anything with LLMs too?
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#48Earlier quoted context omitted.
Yeah, and that doesn't even do the nutty IO on these things justice. A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.
It’s unbelievable stuff. Does anyone know how much a single box costs? They are selling them it looks like.
Mostly teasing but my guess would be $500k+ since they'll likely price it so that it is the same $ as the equivalent NVIDIA cluster (or very close to it).
Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#49Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
#50This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…
I imagine this "cat out of the bag" situation, the democratization and commodification of powerful technology accessible and affordable to the public, is similar to what's happening with single-board computers and microcontrollers like Raspberry Pi, Arduino, ESP32.
It might be similar to what happened with mobile phones, but there the power was quite restricted. The (mostly) duopoly of iOS and Android, with devices and apps locked down in various ways. Sure we can "jail break" and "root" our phone, but that's not for the general public.
Maybe solar energy production is going through a similar process, with panels and batteries becoming more efficient and affordable every year.
Certainly, it reminds one of the history of personal computers, the way such a powerful general-purpose tool became ubiquitous and local.