Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

101–110 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#101

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

I for one dream of a future without maps. I want to walk through a distant forest to find an ancient, unconnected ESP-32 in the bark of a tree containing a tiny specialized AI that can only tell me about things relevant to the area, how far to walk upstream to the nearest town. And only if I can find it and scan an RFID tag to wake it up.

A beautiful dream!

> I like to think (right now please!) of a cybernetic forest filled with pines and electronics where deer stroll peacefully past computers as if they were flowers with spinning blossoms.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#102

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Yes, yes, and yes. I'm waiting for an actually open AI that can run on the edge, purely on commodity hardware like our laptops and phones - it's inevitable. I imagine this "cat out of the bag" situation, the democratization and commodification of powerful technology accessible and affordable to the public, is similar to what's happening with single-board computers and microcontrollers like Raspberry Pi, Arduino, ESP3…

> the democratization and commodification of powerful technology accessible and affordable to the public

Empowering everyone equally, very low barrier of entry. The spread of AI will make some skills become less important and other skills more important. There will be painful social changes. Memorising stuff, being exact, being careful, diligent - all these things can be done by LLM. But also creative recombination of ideas. For now we still hold supremacy in long form content, complex software projects and detecting inconsistencies, but they could be emergent qualities in GPT-5.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#103

Earlier quoted context omitted.

I guess I've been out of the loop for a bit and didn't realize that "edge computing" became a term since cloud computing took off. It is kind of cyclical then is not? By that I mean computers used to be shared and to log into it through a terminal. Then the PC came around. Then about 15 years ago Cloud computing became the rage (really an extension or more sophisticated system than the first time shared computers) No…

> It is kind of cyclical then is not? > By that I mean computers used to be shared and to log into it through a terminal. > Then the PC came around. > Then about 15 years ago Cloud computing became the rage (really an extension or more sophisticated system than the first time shared computers) There's a really neat article called "The Eternal Mainframe"[1] that you might be interested. It explores this idea in greate…

Thanks, that was an interesting read!

I wonder if the author's perspective has changed with regards to freedom to compute.

Social Media is often used as an example of privacy invasion though I've failed to see why concerns over Facebook handling your private data is worrying when they don't have a product you need to have.

Email on the other hand, is pretty much a necessity today so privacy concerns are vital there imo. Of course you can host your own server whereas you can't host your own Facebook.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#104

Earlier quoted context omitted.

I guess it's something. It still goes to show how far open models are behind the proprietary SOTA.

Indeed but this is zero-shot performance. Fine-tuning for a task should get you pretty good results. I'm interested in seeing the results of an Alpaca method against this Cerebras 13B model.

Base model performance is what's most important and also impacts fine-tuning quality. Practically, a model that's good out of the box with minimal fine-tuning is also useful to more people. Since they focused on being training compute optimal for some budget, expect their models to lag behind Llama overall. Their 6.7B version should lag behind GPT-J, assuming 20 tokens per parameter.

The Pythia models are also worth checking out, they might be better than or matched to CerebrasGPTs at each size (although they warn it is not intended for deployment).

Conclusion: the landscape of top open models remains unchanged.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#105
post #72

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Google: "confetti has left the cannon" > No results found for "confetti has left the cannon". I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.

For the curious, here was the conversation I had:

ME: What's another saying similar to: "The cat is out of the bag" or "The genie is out of the bottle"?

chatgpt: Another similar saying is "the toothpaste is out of the tube." Like the other two expressions, this phrase conveys the idea that something has been revealed or unleashed, and it cannot be undone or put back into its original state.

ME: Can you invent a new phrase that has the same meaning, which is completely new?

chatgpt: Certainly! Here's a new phrase with a similar meaning: "The confetti has left the cannon." This phrase suggests that once something is out in the open or has happened, it's impossible to reverse or contain it, much like confetti that has been shot out of a cannon during a celebration.

ME: What's the name for this kind of expression? Aphorism? Fable?

chatgpt: This kind of expression is known as an idiom. Idioms are phrases or expressions with a figurative meaning that is different from their literal meaning. They are commonly used in everyday language to convey a specific message or sentiment in a more colorful or memorable way.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#107
post #88
post #64

Comparing the 13B model here https://huggingface.co/cerebras/Cerebras-GPT-13B to LLaMA-13B https://github.com/facebookresearch/llama/blob/main/MODEL_CA... you can see that in all of the reasoning tasks Cerebras-GPT lags behind. Any reason to use Cerebras instead of LLaMA? Doesn't seem like it.

it lags behind because according to their blogpost it was trained on <300B tokens. LLaMAs as far as I know were trained on more than trillion

The LLaMa paper says 1 trillion for the smaller models (7B, 13B) and 1.4 trillion for the larger models (30B, 65B)

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#108
post #95
post #72

Earlier quoted context omitted.

Google: "confetti has left the cannon" > No results found for "confetti has left the cannon". I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.

That name aged like milk. First of all, what you said. And second - a parrot can make more parrots without human help, language models can't make GPU chips. Insulting for both LLMs and parrots.

Language models exist not in the physical world but digital. A more apt comparison would be if LLMs can compile new (or slightly modified) LLMs, and I believe the answer will soon be (if not already), a resounding yes.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#109

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

Yes, this is true. But, I worry about how long it will take for the utility of "GPT-4" on my phone to be close enough to whatever is only possible through models running on large cloud platforms to make that choice relatively drawback free.

Is the curve of what this class of algorithms can provide sigmoid? If so, then yeah, eventually researchers should be able to democratize it sufficiently that the choice to use versions that can run on private hardware rational. But if the utility increases linearly or better over time/scale, the future will belong to whoever owns the biggest datacenters.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#110

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

I'd go one step further if it is not happening yet: smaller companies should really pool their resources to train open LLMs. Say, form a consortium and work with the open source community to build ChatGPT-equivalent. Companies will be crazy to assume that they can hand their future to the APIs offered by a handful of companies during this monumental technological paradigm shift in history.

That is, a real OpenAI with a open government body.

Post reply on HN