Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

131–140 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#131
post #72

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Google: "confetti has left the cannon" > No results found for "confetti has left the cannon". I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.

Not exactly the same, but from:

https://www.instagram.com/p/CQdBiVyh5C2/?hl=en

Now that the cat is out of the bag, or, should I say the confetti is out of the… can?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#132

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

> And smart people will package them as "one click installers" for people who aren't tech-savvy. This is already happening. Any projects I can follow? Because I haven't seen any one click installers yet that didn't begin with "first install a package manager on the command line"

Not an llm but this 1 click installer for stable diffusion is literally a 1 click installer. Impressively works.

https://github.com/cmdr2/stable-diffusion-ui

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#133
post #24

Earlier quoted context omitted.

https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.

Yeah, and that doesn't even do the nutty IO on these things justice. A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.

Cerebras makes impressive hardware, but Nvidia still performs better in every regard. The most telling factor is that Cerebras claims they're too busy to run common benchmarking (e.g. MLPerf) to compare against Nvidia.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#134

Slightly off-topic: I remember seeing news about the enormous chip Cerebras was/is selling (pdf https://f.hubspotusercontent30.net/hubfs/8968533/WSE-2%20Dat... ). Has there been any indication that the LLMs released in the last few months use exotic hardware like this, or is it all "standard" hardware?

Everyone except Google uses Nvidia for training. Cerebras, Gaudi, and other custom AI accelerators have unable to surpass Nvidia in performance/$ and performance/watt yet.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#135
post #95

Earlier quoted context omitted.

That name aged like milk. First of all, what you said. And second - a parrot can make more parrots without human help, language models can't make GPU chips. Insulting for both LLMs and parrots.

Language models exist not in the physical world but digital. A more apt comparison would be if LLMs can compile new (or slightly modified) LLMs, and I believe the answer will soon be (if not already), a resounding yes.

I agree, LLMs can generate the source code for their network, can generate the training data, supervise a model being trained and explain step by step what they do. That's why I specifically said they can't make GPUs, if they could, they would be self replicators. We'd have to consider them alive.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#136

I've been following open source LLMs for a while and at first glance this doesn't seem too powerful compared to other open models, Flan-Alpaca[0] is licensed under Apache 2.0, and it seems to perform much better. Although I'm not sure about the legalities about that licensing, since it's basically Flan-T5 fine-tuned using the Alpaca dataset (which is under a Non-Commercial license). Nonetheless, it's exciting to see…

Sounds like you might be the right person to ask the “big” question. For a small organization or individual who is technically competent and wants to try and do self-hosted inference. What open model is showing the most promise and how does it’s results compare to the various openAI GPTs? A simple example problem would be asking for a summary of code. I’ve found openAI’s GPT 3.5 and 4 to give pretty impressive englis…

Google's Flan-T5, Flan-UL2 and derivatives, are so far the most promising open (including commercial use) models that I have tried, however they are very "general purpose" and don't perform well in specific tasks like code understanding or generation. You could fine-tune Flan-T5 with a dataset that suits your specific task and get much better results, as shown by Flan-Alpaca.

Sadly, there's no open model yet that acts like a Swiss knife and gets good-enough results for multiple use cases.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#137

I've been following open source LLMs for a while and at first glance this doesn't seem too powerful compared to other open models, Flan-Alpaca[0] is licensed under Apache 2.0, and it seems to perform much better. Although I'm not sure about the legalities about that licensing, since it's basically Flan-T5 fine-tuned using the Alpaca dataset (which is under a Non-Commercial license). Nonetheless, it's exciting to see…

Their goal isn't to make a powerful model. It's to show how well compute-optimal models do on test-loss as a function of increasing model size. This function can be used with some caveats to forecast the test-loss of larger models for which compute-optimality becomes more important.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#138

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Yes, yes, and yes. I'm waiting for an actually open AI that can run on the edge, purely on commodity hardware like our laptops and phones - it's inevitable. I imagine this "cat out of the bag" situation, the democratization and commodification of powerful technology accessible and affordable to the public, is similar to what's happening with single-board computers and microcontrollers like Raspberry Pi, Arduino, ESP3…

If you thought crypto miners drove up the prices of GPUs, just wait until a gaming card can host a person who can get a job.

Integer-quantized LLaMA can fit on a 4090, and isn't quite human-level. If a fully aligned, reliably non-hallucinary turingrade LLM can run on the next generation 5 series GPUs then expect their market price to quadruple or more.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#139
post #131
post #72

Earlier quoted context omitted.

Google: "confetti has left the cannon" > No results found for "confetti has left the cannon". I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.

Not exactly the same, but from: https://www.instagram.com/p/CQdBiVyh5C2/?hl=en Now that the cat is out of the bag, or, should I say the confetti is out of the… can?

Is it strange that I think it's more likely that LLMs generalize and synthesize new things out of old patterns instead of just copying a singular example.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#140

Earlier quoted context omitted.

I guess it's something. It still goes to show how far open models are behind the proprietary SOTA.

Indeed but this is zero-shot performance. Fine-tuning for a task should get you pretty good results. I'm interested in seeing the results of an Alpaca method against this Cerebras 13B model.

I agree fine-tuning for task will give better results. Cerebras actually showed some research recently on this front. Sparse pre-training and dense fine-tuning (https://arxiv.org/abs/2303.10464). You can recover the accuracy of sparse pre-trained models with dense fine-tuning and reduce FLOPs of the end-to-end pipeline by 2.5x compared to dense.
Post reply on HN