Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

121–130 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#121

Looking at their charts it seems like their 6.7B model is considerably worse than GPT-J which is an existing open 6B model from several years ago. I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.

Someone posted this repost from the Cerebras Discord earlier, but sharing for visibility - "We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for ou…

> Cerebras Discord

This is really weird to hear out loud.

I still think of Discord as a niche gaming chatroom, even though I know that (for instance) a wafer scale IC design company is hosting a Discord now.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#123

Earlier quoted context omitted.

Someone posted this repost from the Cerebras Discord earlier, but sharing for visibility - "We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for ou…

Sounds like we should crowd-fund the cost to train and open source one of these models with LLaMa-like quality. I'd chip in!

TBH that seems like a good job for Cerebras.

There are plenty of such efforts, but the organizer needs some kind of significance to attract a critical mass, and a AI ASIC chip designer seems like a good candidate.

Then again, maybe they prefer a bunch of privately trained models over an open one since that sells more ASIC time?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#124

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

An LLM running locally providing type-ahead completions seems inevitable.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#125
post #64

Comparing the 13B model here https://huggingface.co/cerebras/Cerebras-GPT-13B to LLaMA-13B https://github.com/facebookresearch/llama/blob/main/MODEL_CA... you can see that in all of the reasoning tasks Cerebras-GPT lags behind. Any reason to use Cerebras instead of LLaMA? Doesn't seem like it.

Can the LLaMA weights be used for commercial products?

There are two aspects to it.

The first one is whether they would actually sue. The optics would be terrible. A similar situation occurred in the 90s when the RC4 cipher’s code was leaked. Everyone used the leaked code pretending that it was a new cipher called arc4random, even though they had confirmation from people that licensed the cipher that its output was identical. Nobody was sued, and the RSA company never acknowledged it.

The second one is related to the terms. The LLaMA weights themselves are licensed under terms that exclude commercial use:[0]

> You will not […] use […] the Software Products (or any derivative works thereof, works incorporating the Software Products, or any data produced by the Software), […] for […] any commercial or production purposes.

But the definition of derivative works is gray. AFAIK, if LLaMA is distilled, there is an unsettled argument to be had that the end result is not a LLaMA derivative, and cannot be considered copyright or license infringement, similar to how models trained on blog articles and tweets are not infringing on those authors’ copyright or licensing. The people that make the new model may be in breach of the license if they agreed to it, but maybe not the people that use that new model. Otherwise, ad absurdum, a model trained on the Internet will have content that was generated by LLaMA in its training set, so all models trained on the Internet after Feb 2023 will break the license.

IANAL, but ultimately, Meta wins more by benefiting from what the community contributes on top of their work (similar to what happened with React), than by suing developers that use derivatives of their open models.

[0]: https://docs.google.com/forms/d/e/1FAIpQLSfqNECQnMkycAp2jP4Z...

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#126

Earlier quoted context omitted.

I don’t understand why people aren’t more impressed with it clearly understanding and then even inventing idioms. That shows some real intelligence.

It’s because they’re confused in thinking human intelligence isn’t learned stochastic expectation.

That's how the brain evolved, after all: take incoming input, process according to previous patterns, output a reaction to peripheral nervous system that will enhance survival. The evolutionary purpose of a cortex is to store a statistical model of an animal's environment based on past stimulus. Human intelligence is just a complicated permutation of these initial systemic parameters.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#127
post #88
post #64

Comparing the 13B model here https://huggingface.co/cerebras/Cerebras-GPT-13B to LLaMA-13B https://github.com/facebookresearch/llama/blob/main/MODEL_CA... you can see that in all of the reasoning tasks Cerebras-GPT lags behind. Any reason to use Cerebras instead of LLaMA? Doesn't seem like it.

it lags behind because according to their blogpost it was trained on <300B tokens. LLaMAs as far as I know were trained on more than trillion

[deleted]

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#128

Slightly off-topic: I remember seeing news about the enormous chip Cerebras was/is selling (pdf https://f.hubspotusercontent30.net/hubfs/8968533/WSE-2%20Dat... ). Has there been any indication that the LLMs released in the last few months use exotic hardware like this, or is it all "standard" hardware?

You might see more training on Intel XPUs when they come out, since they have such enormous RAM pools. Maybe AMD MI300s and Intel Ponte Vecchio (both 128GB) in the shorter term, though I think they will mostly be in HPC supercomputers instead of cloud instances.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#129

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

> And smart people will package them as "one click installers" for people who aren't tech-savvy. This is already happening.

Any projects I can follow? Because I haven't seen any one click installers yet that didn't begin with "first install a package manager on the command line"

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#130
post #33

OT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)

Last time I viewed it, I believe it wasn't blurry. Perhaps to scale the traffic the images are now displayed in lower quality? But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models. Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.

I can confirm, it does change. As of now, it displays one of higher quality: https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
Post reply on HN