Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

561–570 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#561
post #65

Earlier quoted context omitted.

It's also a power issue. The 4090 sounds like you're going to need a much, MUCH higher PSU than you currently use.. or it'll suddenly turn off as it uses 2-3x the power. You'll need your own wiring to run your PC soon :-)

I think it is a stupid question, but does the power consumption needed by processors to infer compared to human brains demonstrate that there is something fundamentally wrong for the AI approach or is it more physics related? I am not a physicist or biologist or anything like that so my intuition is probably completely wrong but it seems to me that for more basic inference operations (lets say add two numbers) power…

The AI is much faster than the brain, if you batch requests the cost goes down.

Re: YaLM-100B: Pretrained language model with 100B parameters

#562

Earlier quoted context omitted.

Is there a reason why it is required to fill the swapfile with zeroes here? Normally you'd see something like "dd of=/swapfile bs=1G seek=3 count=0", creating a file of size 3G but with no space allocated (yet). It's much quicker to complete the setup this way.

I assume if you force the file system to allocate inodes you are likely to have a less fragmented file than if you create a sparse file that gets inodes assigned over time when each part is used.

Interesting guess but wrong I'm afraid :)

It's simply because it's an easy way to create a file of a certain size that most Linux users would be familiar with.

The quicker way (and possibly more "proper" way) is to use fallocate, but who has even heard of that vs dd ?

Re: YaLM-100B: Pretrained language model with 100B parameters

#563

Earlier quoted context omitted.

Search Google and Yandex for "2020 election fraud." The results are VERY different. The Zach Vorhies leak shows that Google regularly does blatant censorship for political purposes.[1] [1] https://www.breitbart.com/tech/2021/08/19/google-whistleblow...

I don't know man, "thegatewaypundit.com" as a top reputable source? seems to me like it's not "honest two-sided results" but just, well, a rather random mix of result of widely varying quality. Mad Altavista vibes! What I'm trying to say is that even if you believe that "was the 2020 US election stolen?" is worth debating, which it isn't, the yandex results are shit.

If you get all your information through mainstream channels, and you don't want to see anything contradicting those channels then you should continue to use Google because they explicitly implement the algorithms on controversial topics to prefer mainstream news sources[1]. What I mean by "better" in terms of controversial searches is that on controversial matters, it will rank the searches the same way it does for all other searches. I mean yeah, I don't have access to the internal code base of Yandex, but it certainly feels more organic.

[1]https://www.breitbart.com/tech/2019/05/12/study-the-cnn-sear...

Re: YaLM-100B: Pretrained language model with 100B parameters

#564

I am one of the people who worked on Google's PaLM model. Having skimmed the GitHub readme and medium article, this announcement seems to be very focused on the number of parameters and engineering challenges scaling the model, but it does not contain any details about the model, training (learning rate schedules, etc.), or data composition. It is great that more models are getting released publicly, but I would not…

> this announcement seems to very focused on number of parameters And yet your own project headline is "Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrough Performance"[0]. 0- https://ai.googleblog.com/2022/04/pathways-language-model-pa...

The difference is PaLM was extensively benchmarked and it performed as well as it should, which is to say, amazingly well. The irony here is that you should instead be invoking that other ~500b model, Nvidia's Megatron-530b, which was undertrained, only cursorily evaluated (no interest in any new capabilities or even examining old ones like inner monologues) and promptly forgotten by everyone after the headlines about being the largest dense model: https://arxiv.org/abs/2201.11990#microsoftnvidia

Re: YaLM-100B: Pretrained language model with 100B parameters

#565

Earlier quoted context omitted.

No, it is more like generating a conversations, translating text, summarization texts, writing code, etc.

If I wanted to use it for summarization, what would I have to do?

Postfix "tldr:" to the text being summarized. (Even GPT-2 could do that.)

Re: YaLM-100B: Pretrained language model with 100B parameters

#566
post #560

Earlier quoted context omitted.

> It's not just money, it's people, it's culture, it's all the great projects the company does. What about people killed by the Russian army, sponsored by Yandex? I guess those matter less than the company culture, right?

The Russian army is not sponsored by Yandex. The money comes from selling natural resources... and mostly to Europe, surprise. It's about $1 billion per day. Tax money from private companies is nothing compared to that. So Europe is sponsoring the war way more than Yandex. Let's then shut down the Europe, right? You can say — look, they're trying hard to get rid of Russian resources. But Yandex is also trying hard to…

> The Russian army is not sponsored by Yandex.

It is.

> So Europe is sponsoring the war way more than Yandex.

That's unfortunately true. The dependency is real, and it will take a long time to get rid of it.

> And all that "canceling" of Yandex really doesn't help (it does the opposite in fact).

Cancelling Yandex completely, as in forcing it to collapse, would help a lot. Yandex services (together with VK) are extremely important in the Russian society and economy, and their collapse would weaken Russia and its ability to wage (military/economic) war a lot. As such, this would be the best course of action (as mentioned before, burn the equipment, delete the code).

Re: YaLM-100B: Pretrained language model with 100B parameters

#567
post #13

I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers. Have to wonder what the knock-on effects of that will be, especially if the models improve drastically. With so much of our social lives being moved online, if we have the easy ability to create fake lives of fake people one has to wonder what's real and what isn't. Maybe the dead internet theory will rea…

>> I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers.

Do you mean train or run? My assumption was all these models could be run on most computers, probably with a simple docker container, as long as there is sufficient RAM to hold the network, which should be most laptops > 16gb ram.

Speaking of which, anyone have recommendations on pre-trained docker containers with weights included?

Re: YaLM-100B: Pretrained language model with 100B parameters

#568
post #560

Earlier quoted context omitted.

The Russian army is not sponsored by Yandex. The money comes from selling natural resources... and mostly to Europe, surprise. It's about $1 billion per day. Tax money from private companies is nothing compared to that. So Europe is sponsoring the war way more than Yandex. Let's then shut down the Europe, right? You can say — look, they're trying hard to get rid of Russian resources. But Yandex is also trying hard to…

> The Russian army is not sponsored by Yandex. It is. > So Europe is sponsoring the war way more than Yandex. That's unfortunately true. The dependency is real, and it will take a long time to get rid of it. > And all that "canceling" of Yandex really doesn't help (it does the opposite in fact). Cancelling Yandex completely, as in forcing it to collapse, would help a lot. Yandex services (together with VK) are extrem…

> Cancelling Yandex completely, as in forcing it to collapse, would help a lot.

It's just a wishful thinking. It wont "collapse", it would just become controlled by government, and then it truly becomes the instrument of the evil, so that not only News, but every service Yandex provides will serve the government needs. They will recruit soldiers through Yandex services, they make Yandex develop AI-controlled tanks and whatnot. Every thing that Yandex doesn't do now (because they do not actually support the war) — they will make it to do.

> their collapse would weaken Russia and its ability to wage (military/economic) war a lot

Of course not, because the Russian army and the military industrial complex is in no way dependent on the search engine and the food delivery service Yandex provides. You can destroy those, sure. People lifes get slightly worse, and then competitors catch up (there is a lot of competition to Yandex in Russia and they are not going to fade away).

Re: YaLM-100B: Pretrained language model with 100B parameters

#569
post #568

Earlier quoted context omitted.

> The Russian army is not sponsored by Yandex. It is. > So Europe is sponsoring the war way more than Yandex. That's unfortunately true. The dependency is real, and it will take a long time to get rid of it. > And all that "canceling" of Yandex really doesn't help (it does the opposite in fact). Cancelling Yandex completely, as in forcing it to collapse, would help a lot. Yandex services (together with VK) are extrem…

> Cancelling Yandex completely, as in forcing it to collapse, would help a lot. It's just a wishful thinking. It wont "collapse", it would just become controlled by government, and then it truly becomes the instrument of the evil, so that not only News, but every service Yandex provides will serve the government needs. They will recruit soldiers through Yandex services, they make Yandex develop AI-controlled tanks an…

> It's just a wishful thinking. It wont "collapse", it would just become controlled by government

That's why part of my suggestion is to burn the equipment/infrastructure and delete the code.

> They will recruit soldiers through Yandex services, they make Yandex develop AI-controlled tanks and whatnot

And the only thing stopping them now from doing that is that Yandex is not nationalized. Yeah, sure.

> Of course not, because the Russian army and the military industrial complex is in no way dependent on the search engine and the food delivery service Yandex provides.

Yandex provides many services, it's much like google - maps, translation, drive, mail etc. etc. Bringing it down would cripple many private and economic activities. Russia can't sustain waging wars if they don't have an economy and disgruntled population.

With the exception of VK, there isn't really any step-in competition to Yandex. Even if there was, losing all your data in e.g. mail/drive will have significant consequences.

Re: YaLM-100B: Pretrained language model with 100B parameters

#570
post #456

To add a voice of skepticism. The recent rush to open source these models may be indicative that the tens of millions that’s spent training these things has relatively poor roi. There may be a hope that someone else figures out how to make these commercially useful.

HuggingFace will soon release their BigScience model: https://twitter.com/BigScienceLLM/status/1539941348656168961

"a 176 billion parameter transformer model that will be trained on roughly 300 billion words in 46 languages"

So anything smaller than that will become worthless. May be a factor, companies have a last chance to make a PR splash before it happens.

Read more about it: https://bigscience.huggingface.co/blog/model-training-launch...

Post reply on HN