Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

191–200 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#191
post #153
post #151

Earlier quoted context omitted.

And it takes ~20 years to train a new brain so it can coherently answer questions about a wide variety of topics. Even worse, you can't even copy-paste it once you're done!

It arguably needs much less training data though.

How? Organisms with brains process every second of their life, is that not training data on a level comparable with current AI models?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#192
post #72

Earlier quoted context omitted.

Google: "confetti has left the cannon" > No results found for "confetti has left the cannon". I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.

For the curious, here was the conversation I had: ME: What's another saying similar to: "The cat is out of the bag" or "The genie is out of the bottle"? chatgpt: Another similar saying is "the toothpaste is out of the tube." Like the other two expressions, this phrase conveys the idea that something has been revealed or unleashed, and it cannot be undone or put back into its original state. ME: Can you invent a new p…

That’s brilliant. Though personally I prefer "The SARS-COV-2 virus has left the lab"

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#193
post #153

Earlier quoted context omitted.

It arguably needs much less training data though.

How? Organisms with brains process every second of their life, is that not training data on a level comparable with current AI models?

From a pure data amount point of view yes, but relatively little of that would seem to be relevant for our intellectual capacities. If GPT was a robot moving autonomously around in the world with full visual, auditory and tactile apparatus, it may be a bit different.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#194

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

> And smart people will package them as "one click installers" for people who aren't tech-savvy. This is already happening. Any projects I can follow? Because I haven't seen any one click installers yet that didn't begin with "first install a package manager on the command line"

Here's alpaca running in electron. Not exactly one click but close.

https://github.com/ItsPi3141/alpaca-electron

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#195

Earlier quoted context omitted.

I'm definitely impressed when I see things like this. This is much more impressive than writing a 5 paragraph nonsense about a 2 sentence prompt.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

"The AI effect occurs when onlookers discount the behavior of an artificial intelligence program by arguing that it is not real intelligence."

https://en.wikipedia.org/wiki/AI_effect

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#196
post #195

Earlier quoted context omitted.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

"The AI effect occurs when onlookers discount the behavior of an artificial intelligence program by arguing that it is not real intelligence." https://en.wikipedia.org/wiki/AI_effect

Better that than the opposite effect, to assume that because a system solves a single problem very well, it is intelligent.

Is Stockfish intelligent?

Is a system with A* pathfinding intelligent?

I would define intelligence as the ability to solve a wide variety of novel problems. A system built to be excellent at a single task may be better than humans at that task but still lack "intelligence".

We still don't know what that even is exactly, but historically people consistently underestimate how difficult it is.

Not even knowing how to approach it, researchers work on solving single specialized problems instead and make little progress on whatever "intelligence" is.

(If you'd prefer a different definition of intelligence under which Stockfish and GPT are intelligent, then what would you call the ability to solve a wide variety of novel problems? Feel free to substitute that word for "intelligence" above if you'd like to understand what I'm saying.)

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#197

Earlier quoted context omitted.

I'm definitely impressed when I see things like this. This is much more impressive than writing a 5 paragraph nonsense about a 2 sentence prompt.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

To continue with some human brain powered pedantry for fun I’d argue that “the snow is out of the globe” is not a valid idiom to suit the “X out of Y” pattern.

We have multiple example pairings of X and Y but the common components are that putting X back in Y would be impossible or extremely difficult, and also that X is in some way meant to no longer be contained inside Y as part of either a desired outcome, it’s normal function or the natural expected and thus inevitable result. Cats want to escape, helium leaks, confetti is expelled to have the desired effect, and toothpaste is squeezed out to use it…

For the snow to come out of a snow globe you have to smash it which is not normal usage, not normally meant to happen, and shouldn’t happen by itself. Making the idiom “The snow is out of the globe” not a proper member of this “family” of idioms. (Also I’m not sure if there’s an agreed upon collective noun for idioms)

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#198

Earlier quoted context omitted.

I'm definitely impressed when I see things like this. This is much more impressive than writing a 5 paragraph nonsense about a 2 sentence prompt.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

The issue I see here is you are doing a worse job at this than ChatGTP. Creating idioms is hard, that is why we left most of them to Shakespeare.

- I regularly return cereal to its box.

- "helium" and "balloon" have a more awkward rhythm than "confetti" and "cannon". It also loses the connotations of sudden, explosive and exciting change.

- Snow & globe I'm not even sure what that means in practice. It has poor prospects as an idiom. Is the snow even known for leaving globes?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#199
post #175

Earlier quoted context omitted.

>So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket. I want to believe you, but I'm ignorant of the hardware requirements for these things. How soon do you think we'd be able to run something reasonably gpt4-like on, say, a 4090?

It's literally already here: https://twitter.com/thiteanish/status/1635678053853536256

llama-7b is nowhere even close to GPT-4. Not even llama-65b is.

We'll see if fine tuning can improve this, but e.g. alpaca-30b is still inferior to ChatGPT-3.5.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#200

Earlier quoted context omitted.

It's a pretty mad architecture tbh. Compile times must be absolutely insane. Also Tesla's Dojo also uses a manufacturing technique that has basically obsoleted their WSI design already.

Compile times are not a whole lot different than any other large model build. It's a kernel based compilation pipeline and the kernels are simply tiled over a 'core' area in the weight streaming architecture.

I seriously doubt that. What's your source?

I used to work for a competitor with a more flexible architecture and even our compile times were bad (significant fractions of a day in some cases). And we didn't have to do place and route!

I just googled it and it's apparently bad enough that they had to implement incremental place and route.

Post reply on HN