Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

421–430 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#421
post #60

Earlier quoted context omitted.

> I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers. Isn’t that already the case? Sure, it costs $60K, but that is accessible to a surprisingly large minority, considering the potency of this software.

...what? 60 thousand dollars for a dedicated computer that you can't use is not everyone, not on their own computers, and is also a crazy large amount of money for nearly everyone. Sure there are some that could, but that's not what I said.

Eh, 60k is just a bit more expensive than your average car, and lots of people have cars, and that's just how things are today. I imagine capabilities will be skyrocketing and prices will fall drastically at the same time.

Re: YaLM-100B: Pretrained language model with 100B parameters

#422
post #21

Seeing those gigantic models it makes me sad that even the 4090 is supposed to stay at 24GB of RAM max. I really would like to be able to run/experiment on larger models at home.

Take a look at Apple's M1 Max, a lot of fast unified memory. No idea how useful though

Apple is selling M1's with > 200gb ram? Have a link so I can buy one?

Re: YaLM-100B: Pretrained language model with 100B parameters

#423
post #264

Earlier quoted context omitted.

Perhaps on quantity. Substantially slower though around ~3x from what I can tell…substantial roadblock if you’re training models that take weeks.

I meant for inference, not training. People just want to run the magic genies locally and post funny AI content.

ah right - gotcha

Re: YaLM-100B: Pretrained language model with 100B parameters

#424

I am one of the people who worked on Google's PaLM model. Having skimmed the GitHub readme and medium article, this announcement seems to be very focused on the number of parameters and engineering challenges scaling the model, but it does not contain any details about the model, training (learning rate schedules, etc.), or data composition. It is great that more models are getting released publicly, but I would not…

> this announcement seems to very focused on number of parameters And yet your own project headline is "Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrough Performance"[0]. 0- https://ai.googleblog.com/2022/04/pathways-language-model-pa...

1. The OP did not criticize the headline; they criticized the content. If you read the article that you linked, you would find that they do, in fact, evaluate the performance of the model.

2. 540 billion parameters is notable for its size, which is likely why they lead with that particular headline.

Re: YaLM-100B: Pretrained language model with 100B parameters

#426
post #150

Earlier quoted context omitted.

The bots/machine vs human reminds me of that famous experiment from the 30s in which Winthrop Kellogg[0], a comparative psychologist, and his wife decided to raise their human baby (Donald) simultaneously with a chimpanzee baby (Gua) in an effort to "humanize the ape". It was set out to last 5 years but was relatively quickly abrupted after only 9 months. The explicit reason wasn't stated only that it successfully pr…

A tangentially related thought: Actors attempt to imitate humans. “Good acting” is convincing; the audience believes the actor is giving a reasonable response to the portrayed situation. But the audience is also trying to imitate the actors to some degree. Like you point out, humans imitate. For some subset of the population, I’d imagine the majority of social situations they are exposed to, and the responses to situ…

This is a common phenomena where the fake is more believable than the real thing due to over exposure of the imitation.

Famously the bald eagle sounds nothing like it does in tv and the movies and explosions are rarely massive fireballs. For human interaction it’s much harder to pin down cause and effect but if it happens in other cases it would be very surprising to not happen there.

Re: YaLM-100B: Pretrained language model with 100B parameters

#429
post #6

I have huge respect for developers at Yandex. It's kind of sad that achievements like these are tainted by the fact that they come from Russia (and I speak as a Ukrainian). I wonder if the permissive license is able to mitigate that.

What percentage of American inventions and scientific developments post WW2 were led or influenced by former Nazi scientists?

Probably about same as share as of USSR's. They just were open about it. Also, crucial word here is "former". Like, there is a big difference between being of a former fascist state, and carrying on ongoing genocide.
Post reply on HN