Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

571–580 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#571

I love Yandex. They are the best search engine by far for politically controversial topics. They also release a language model to benefit everyone even if it says politically incorrect stuff. They also name their projects "cocaine" probably to perhaps to prevent western competitors from using them. You look at OpenAI and how they don't release their models mainly because they fear "bad people" will use them for "bad…

>> They are the best search engine by far for politically controversial topics

FYI, they are Russian subject that follows ALL their censorship laws (and oh boy do they have a lot of it).

>> probably to perhaps to prevent western competitors from using them The irony here. All yandex products are exact copies of western, adjusted to local market.

Re: YaLM-100B: Pretrained language model with 100B parameters

#572

In the download script, it skips parts of the model (02 and 83); any ML people have ideas why you'd do that?

It appears the indexing for the model parts is deliberately not contiguous; the 03-82 range represents the main 80 transformer layers. https://github.com/yandex/YaLM-100B/blob/main/megatron_lm/me...

That makes sense, thanks for clearing it up!

Re: YaLM-100B: Pretrained language model with 100B parameters

#573

Earlier quoted context omitted.

I don't know man, "thegatewaypundit.com" as a top reputable source? seems to me like it's not "honest two-sided results" but just, well, a rather random mix of result of widely varying quality. Mad Altavista vibes! What I'm trying to say is that even if you believe that "was the 2020 US election stolen?" is worth debating, which it isn't, the yandex results are shit.

If you get all your information through mainstream channels, and you don't want to see anything contradicting those channels then you should continue to use Google because they explicitly implement the algorithms on controversial topics to prefer mainstream news sources[1]. What I mean by "better" in terms of controversial searches is that on controversial matters, it will rank the searches the same way it does for a…

Why link to Breitbart of all places instead of the original source?

https://www.cjr.org/tow_center/google-news-algorithm.php

Btw Wikipedia’s first few sentences on Breitbart are not inspiring

> Its journalists are widely considered to be ideologically driven, and much of its content has been called misogynistic, xenophobic, and racist by liberals and traditional conservatives alike.[10] The site has published a number of conspiracy theories[11][12] and intentionally misleading stories.[13][14]

Re: YaLM-100B: Pretrained language model with 100B parameters

#574

Earlier quoted context omitted.

Take a look at Apple's M1 Max, a lot of fast unified memory. No idea how useful though

What's the difference between Apple's unified memory and the shared memory pool Intel and AMD integrated GPUs have had for years? In theory you could probably assign a powerful enough iGPU a few hundred gigabytes of memory already, but just like Apple Silicon the integrated GPU isn't exactly very powerful. The difference between the M1 iGPU and the AMD 5700G is less than 10% and a loaded out system should theoretical…

Apple gets significant latency and frequency benefits from placing their LPDDR4 on the SoC itself.

Re: YaLM-100B: Pretrained language model with 100B parameters

#575
post #456

To add a voice of skepticism. The recent rush to open source these models may be indicative that the tens of millions that’s spent training these things has relatively poor roi. There may be a hope that someone else figures out how to make these commercially useful.

They did not publish benchmarks about quality of the models, which is very suspicious. I personally squinted hard when they said removing dropout improves training speed (which is in iterations per second), but said nothing about how it affects the performance (rate of mistakes in inference) of the trained model.

I agree that the lack of benchmarks makes it hard to determine how valuable this model is. But on the topic of dropout, dropout has been dropped for the pretraining stage of several other large models. Off the top of my head: GPT-J-6B, GPT-NeoX-20B, and T5-1.1/LM.

Re: YaLM-100B: Pretrained language model with 100B parameters

#576
This is one of the funniest threads I’ve ever seen on this website. People are yelling at eachother about the CIA and the legitimacy of Israel and Assange and the definition of fascism and… anything that pisses anybody off about international politics in general. In a thread about a piece of software that’s (to me and likely many others) prohibitively expensive to play around with.

Anyway I hope somebody creates a playground with this so I can make a computer write a fan fiction about Kirby and Solid Snake trying to raise a human baby on a yacht in the Caspian Sea or whatever other thing people will actually use this for.

Re: YaLM-100B: Pretrained language model with 100B parameters

#577

Earlier quoted context omitted.

If you get all your information through mainstream channels, and you don't want to see anything contradicting those channels then you should continue to use Google because they explicitly implement the algorithms on controversial topics to prefer mainstream news sources[1]. What I mean by "better" in terms of controversial searches is that on controversial matters, it will rank the searches the same way it does for a…

Why link to Breitbart of all places instead of the original source? https://www.cjr.org/tow_center/google-news-algorithm.php Btw Wikipedia’s first few sentences on Breitbart are not inspiring > Its journalists are widely considered to be ideologically driven, and much of its content has been called misogynistic, xenophobic, and racist by liberals and traditional conservatives alike.[10] The site has published a numbe…

This is the association fallacy, which is, unfortunately, how most people determine what to believe these days.

An absurd example of this fallacy would be, Wikipedia, which you cite, has articles that indicate tobacco smoking may cause disease. The nazis were also anti-smoking[1]. Therefore Wikipedia is Nazi propaganda and you should not trust anything on there.

[1]https://www.amazon.com/Nazi-War-Cancer-Robert-Proctor/dp/069...

Re: YaLM-100B: Pretrained language model with 100B parameters

#578

Earlier quoted context omitted.

Still no. First case I would not even consider censorship. The third one was temporary until Google stopped operating in Russia altogether. A quote from the second one: "cumulative 45 percent decrease in traffic from Google searches"

There is a difference between "Google does not censor anti-war content" and "Google does censor anti-war content, but usually has an excuse I find acceptable". When a company puts Jon Lennon's Merry Xmas (War is Over) behind age restriction banner[1], the question stops being "Is there censorship?" and becomes about the logic of such censorship. >The third one was temporary until Google stopped operating in Russia al…

Precisely zero of what you mentioned so far is censoring anti-war content.

Even in the translation case (which I assume you mean by your "excuse" remark) the original source is still available as is. I am not even sure from the description what translation team it was talking about and what does it have to do with Google exactly. "translate company text for the Russian market" this passage sounds like it talks about translating Google's own interfaces, help pages, press releases, or support articles to Russian. E.g. no external voice is being censored.

Re: YaLM-100B: Pretrained language model with 100B parameters

#579
post #456

To add a voice of skepticism. The recent rush to open source these models may be indicative that the tens of millions that’s spent training these things has relatively poor roi. There may be a hope that someone else figures out how to make these commercially useful.

HuggingFace will soon release their BigScience model: https://twitter.com/BigScienceLLM/status/1539941348656168961 "a 176 billion parameter transformer model that will be trained on roughly 300 billion words in 46 languages" So anything smaller than that will become worthless. May be a factor, companies have a last chance to make a PR splash before it happens. Read more about it: https://bigscience.huggingface.co/blo…

Not necessarily, only ~30% of the database is in English, so it likely won't be as good as a smaller model trained solely or mostly on English words.

https://bigscience.huggingface.co/blog/building-a-tb-scale-m...

Re: YaLM-100B: Pretrained language model with 100B parameters

#580

I love Yandex. They are the best search engine by far for politically controversial topics. They also release a language model to benefit everyone even if it says politically incorrect stuff. They also name their projects "cocaine" probably to perhaps to prevent western competitors from using them. You look at OpenAI and how they don't release their models mainly because they fear "bad people" will use them for "bad…

>> They are the best search engine by far for politically controversial topics FYI, they are Russian subject that follows ALL their censorship laws (and oh boy do they have a lot of it). >> probably to perhaps to prevent western competitors from using them The irony here. All yandex products are exact copies of western, adjusted to local market.

Actually they're not, some of the Yandex products are actually better and pretty innovative (ignoring the political stuff). Maps and Go are especially good. Ditto with Russian banking apps, they put American bank apps to shame.
Post reply on HN