Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

161–170 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#161

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

>So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket.

I want to believe you, but I'm ignorant of the hardware requirements for these things. How soon do you think we'd be able to run something reasonably gpt4-like on, say, a 4090?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#162
I wonder how decrotive our world will become as a consequence of how cheap it will become to make art using AI.

I kind of want 3d marble statues and baroque art of a future reinasance everywhere. But wonder if we will turn minimalistic as a response.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#163
post #58

Earlier quoted context omitted.

The point of those smaller models is for the "Cerebras Scaling Law for Compute-Optimal Training" which is the straight line plot in the image at the top of their webpage when you click the link. They want you to think it's reasonable that because the line is so straight (on a flops log scale) for so long, it could be tempting to extrapolate the pile-loss consequences of continuing compute-optimal training for larger…

> the extrapolation can't continue linearly much further if for no other reason than the test loss isn't going to go below zero Isn’t the test loss logarithmic? If so it sure can go below zero.

According to https://pile.eleuther.ai/paper.pdf the test loss on the pile is the log of the perplexity, and the perplexity is 2^H where H is an entropy which is non-negative. So the perplexity is always at least one, so its log is always at least zero.

So yes the test loss can be seen as a log, but no it's not allowed to go below zero.

The intuition is that the test loss is the number of bits that the model would need on average to encode each next token in the test part of the pile, given that you have seen the preceding parts.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#164

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

>So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket. I want to believe you, but I'm ignorant of the hardware requirements for these things. How soon do you think we'd be able to run something reasonably gpt4-like on, say, a 4090?

I feel like no less than 10 years if the singularity doesn't kick in before that. Hardware and energy isn't progressing as fast as we'd like, and that is the main bottleneck. As in, imagine a world where we actually had the same computing power required to train (not run) GPT-4 in 1s in a phone? That kind of world is way beyond AGI and the cure of cancer IMO. Which is great, because it gives us a very objective goal to achieve these things. Sadly, I don't think we're nowhere near that. What was even the total energy consumption of GPT-4 training? Very hard to imagine computers will get that much better anytime soon. IIRC we have some kind of data that the a smartphone today has the power of the best supercomputer... of 30 years ago, right? Don't remember the source, sadly.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#165
post #153
post #151

Earlier quoted context omitted.

And it takes ~20 years to train a new brain so it can coherently answer questions about a wide variety of topics. Even worse, you can't even copy-paste it once you're done!

It arguably needs much less training data though.

What we shouldn't is anthropomorphise it too much. While LLMs can express themselves and interact with us in natural language, their minds are very different from ours - they never learned by having an embodied self, and they can't continuously learn and adapt the way we do - once the conversation is over, it's like it never existed unless it's captured for a future training cycle.

Right now, their ability to learn is severely limited. And, yet, they outcompete us quite easily in a lot of different tasks.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#166

Earlier quoted context omitted.

I don’t understand why people aren’t more impressed with it clearly understanding and then even inventing idioms. That shows some real intelligence.

It’s because they’re confused in thinking human intelligence isn’t learned stochastic expectation.

This seems backward to me. Wouldn’t you be less impressed by ChatGPT if you thought that thought that human intelligence worked the same way as LLMs?

If humans have some special sauce different from the computer, then it’s crazy that ChatGPT can emulate human writing so well. If humans are also just statistical models, then of course you can throw a big training set at some GPUs and it’ll do the same thing. Why should we be surprised or impressed by idioms?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#167
A tangential question: I wonder what, as chiplets become increasingly more common, will Cerebras do to keep their technological advantage of wafer-scale integration. What is the bandwidth and latency of the connections between the tiles? Is there such a thing as bandwidth per frontier length?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#168
post #153

Earlier quoted context omitted.

It arguably needs much less training data though.

What we shouldn't is anthropomorphise it too much. While LLMs can express themselves and interact with us in natural language, their minds are very different from ours - they never learned by having an embodied self, and they can't continuously learn and adapt the way we do - once the conversation is over, it's like it never existed unless it's captured for a future training cycle. Right now, their ability to learn i…

Agreed. There are a hundred different kinds of information processing that go into a human-like mind, and we've kinda-sorta built one piece. And there are a lot of pieces that it would neither be sane nor useful to build (eg. internalized emotions), so we might not see an AI with all the pieces for a very long time ("never" is probably too much to hope for).

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#169
post #116

Earlier quoted context omitted.

Not exactly sure why it would be surprising that it can come up with a convincing idiom when it can produce remarkably good _poetry_

I think the thrilling part is that it's a somewhat atomic concept that can somewhat convincingly be proven to not exist in the training data. While poetry is more impressive if it's as original it's harder to show that it's not just stitched together from the training data.

I just asked GPT-4 to come up with more such "provably original idioms":

""" Here are a few more examples of idioms with meaningful and provable atomic originality:

"The kite has touched the stars" - This phrase could mean that someone has achieved a seemingly impossible goal or reached a level of success that was thought to be unattainable.

"The paint has mingled on the canvas" - This idiom might convey the idea that once certain decisions are made or actions taken, the resulting outcome can't be easily separated or undone, similar to colors of paint that have blended together on a canvas.

"The clock has chimed in reverse" - This expression could be used to describe a situation where something unexpected and unusual has occurred, akin to the unlikely event of a clock chiming in reverse order.

"The flower has danced in the wind" - This phrase could signify that someone or something has gracefully and nimbly adapted to changing circumstances, just as a flower might sway and move in response to the wind. """

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#170

Earlier quoted context omitted.

>So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket. I want to believe you, but I'm ignorant of the hardware requirements for these things. How soon do you think we'd be able to run something reasonably gpt4-like on, say, a 4090?

I feel like no less than 10 years if the singularity doesn't kick in before that. Hardware and energy isn't progressing as fast as we'd like, and that is the main bottleneck. As in, imagine a world where we actually had the same computing power required to train (not run) GPT-4 in 1s in a phone? That kind of world is way beyond AGI and the cure of cancer IMO. Which is great, because it gives us a very objective goal…

I don't think 10 years for training such large models on phones is entirely feasible, if only because phones are mainly concerned with power drain and ergonomics before anything else. Nobody is going to buy an iPhone 15 if it's the size of a brick and lasts 2 hours just because you're training&running models on it.

The focus instead should be on running expansive models locally on your desktop system at home. This is a better focus in two ways: one the power issue is a non-concern (and AMD has already stated power usage will reach 500-700W average by 2025), two once you have it running locally the avenues of use open up such as being able to access your local model through other devices without the heavy burden on those devices.

Post reply on HN