Live data from Hacker News

Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

cerebras.net

211–220 of 231 posts

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#211

Earlier quoted context omitted.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

Like you, I thought these pieces of software and data were little more than statistics-based text generators. But it turns out that this is a Category Mistake. There was an argument made by Raphaël Millière in a recent Mindscape Podcast [1] with Sean Carroll that finally landed for me. He used the example that human beings are driven to eat and reproduce, so by that argument all humans are just eating and reproducing…

That emergence is precisely what I'm looking for evidence of.

Human beings evolved to eat and reproduce and yet here we are, building computers and inventing complex mathematical models of language and debating whether they're intelligent.

We're so far from the environment we evolved to solve that we've clearly demonstrated the ability to adapt.

ChatGPT doing well at a language task isn't demonstrating that same ability to adapt because that's the task it was designed and trained to do. ChatGPT doing something completely different would be the impressive example.

In short: I don't categorically reject the possibility that LLMs might become capable of more than being "statistics-based text generators", I simply require evidence.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#212

Earlier quoted context omitted.

Better that than the opposite effect, to assume that because a system solves a single problem very well, it is intelligent. Is Stockfish intelligent? Is a system with A* pathfinding intelligent? I would define intelligence as the ability to solve a wide variety of novel problems. A system built to be excellent at a single task may be better than humans at that task but still lack "intelligence". We still don't know w…

> Is Stockfish intelligent? > Is a system with A pathfinding intelligent?* I'm not sure if we should get stuck on definitions of intelligence. The fact is that these tools are useful, as are the currently existing AI's. The latter can also pass for humans, in many ways, while the algorithms you mentioned can only pass for humans in very narrow domains. Both can exceed human performance in some ways. Eventually, AI's…

We'll have to "get stuck on" definitions of intelligence if we want to talk precisely about what LLMs are capable of.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#213
post #198

Earlier quoted context omitted.

The issue I see here is you are doing a worse job at this than ChatGTP. Creating idioms is hard, that is why we left most of them to Shakespeare. - I regularly return cereal to its box. - "helium" and "balloon" have a more awkward rhythm than "confetti" and "cannon". It also loses the connotations of sudden, explosive and exciting change. - Snow & globe I'm not even sure what that means in practice. It has poor prosp…

> "helium" and "balloon" have a more awkward rhythm than "confetti" and "cannon". It also loses the connotations of sudden, explosive and exciting change. Not only that, but "the confetti has left the cannon" is an alliteration, which makes the phrase even more poetic.

I actually agree, it's a very good phrase.

But I do think it's cherry picking the most impressive example. I repeated the dialog (and some variations), each time asking for a completely new idiom, and ChatGPT responded with several phrases that aren't new at all:

"The toothpaste is out of the tube"

"The lid has been lifted"

"The secret has been spilled"

"The arrow has left the bow"

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#214
post #203

Earlier quoted context omitted.

Is it? There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of. It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe. And it's exactly the one thi…

In a way, this comment perfectly encapsulates why the argument "machines will never replicate human behavior" is so ridiculous. Instead of engaging with the discussion and topic, you chose a position, and then tried to justify it without really thinking about why one example works and the other one doesn't. In doing so you're literally showing that for certain topics, machines are already more capable than some human…

I didn't say "machines will never replicate human behavior" so I don't think you're engaging with what I said.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#215

Earlier quoted context omitted.

Like you, I thought these pieces of software and data were little more than statistics-based text generators. But it turns out that this is a Category Mistake. There was an argument made by Raphaël Millière in a recent Mindscape Podcast [1] with Sean Carroll that finally landed for me. He used the example that human beings are driven to eat and reproduce, so by that argument all humans are just eating and reproducing…

That emergence is precisely what I'm looking for evidence of. Human beings evolved to eat and reproduce and yet here we are, building computers and inventing complex mathematical models of language and debating whether they're intelligent. We're so far from the environment we evolved to solve that we've clearly demonstrated the ability to adapt. ChatGPT doing well at a language task isn't demonstrating that same abil…

Another poster shared a link to this paper last week for Theory of Mind: https://arxiv.org/abs/2302.02083

We're seeing those other capabilities emerge; like being able to play chess though it's not been trained to do so. That is, these LLMs are displaying emergent abilities associated with reasoning.

These LLMs aren't R. Daneel Olivaw or R2D2 (which is what I think of when I think of the original term for AI, and what we took to calling AGI). We're closer to seeing the just-the-facts AIs we encounter in Blindsight. Intelligence without awareness.

Funny that we still have to use science fiction to make our comparisons because our philosophy of intelligence, mind, and consciousness are insufficient to speak on the matter clearly.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#216
post #193

Earlier quoted context omitted.

How? Organisms with brains process every second of their life, is that not training data on a level comparable with current AI models?

From a pure data amount point of view yes, but relatively little of that would seem to be relevant for our intellectual capacities. If GPT was a robot moving autonomously around in the world with full visual, auditory and tactile apparatus, it may be a bit different.

Hm, not sure how most of that data would be irrelevant, could you clarify? I think all of that data as well as interacting with the environment creates the level of knowledge and intelligence we have today.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#217

Earlier quoted context omitted.

> Is Stockfish intelligent? > Is a system with A pathfinding intelligent?* I'm not sure if we should get stuck on definitions of intelligence. The fact is that these tools are useful, as are the currently existing AI's. The latter can also pass for humans, in many ways, while the algorithms you mentioned can only pass for humans in very narrow domains. Both can exceed human performance in some ways. Eventually, AI's…

We'll have to "get stuck on" definitions of intelligence if we want to talk precisely about what LLMs are capable of.

Not necessarily, as we can just evaluate them on their performance as we give them ever greater challenges.

To do this we do not need to consider whether they're intelligent at all.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#218

This type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the e…

This is a shocking turn of events given there's no edge equivalent of the previous most powerful information tools (web-scale search). It does seem like it will still be a challenge to continuously collect, validate, and train on fresh information. Large orgs like Google/YouTube/TikTok/Microsoft still seem to have a huge advantage there.

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#219

Earlier quoted context omitted.

Sounds like you might be the right person to ask the “big” question. For a small organization or individual who is technically competent and wants to try and do self-hosted inference. What open model is showing the most promise and how does it’s results compare to the various openAI GPTs? A simple example problem would be asking for a summary of code. I’ve found openAI’s GPT 3.5 and 4 to give pretty impressive englis…

Google's Flan-T5, Flan-UL2 and derivatives, are so far the most promising open (including commercial use) models that I have tried, however they are very "general purpose" and don't perform well in specific tasks like code understanding or generation. You could fine-tune Flan-T5 with a dataset that suits your specific task and get much better results, as shown by Flan-Alpaca. Sadly, there's no open model yet that act…

Iterating on the question, what model/weights would be the most appropriate for the specific use case of code generation right now?

Re: Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models

#220

Earlier quoted context omitted.

> "helium" and "balloon" have a more awkward rhythm than "confetti" and "cannon". It also loses the connotations of sudden, explosive and exciting change. Not only that, but "the confetti has left the cannon" is an alliteration, which makes the phrase even more poetic.

I actually agree, it's a very good phrase. But I do think it's cherry picking the most impressive example. I repeated the dialog (and some variations), each time asking for a completely new idiom, and ChatGPT responded with several phrases that aren't new at all: "The toothpaste is out of the tube" "The lid has been lifted" "The secret has been spilled" "The arrow has left the bow"

> But I do think it's cherry picking the most impressive example.

Yes, but don't we cherry pick from what humans have said, too? I'm sure there have been many dumb and obvious proverbs that didn't survive.

> ChatGPT responded with several phrases that aren't new at all

Are you using the GPT-4 version of ChatGPT? That's what GP used.

Post reply on HN