Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

211–220 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#211
post #205
post #178

Earlier quoted context omitted.

It's not that Russia Bad, it's that if you know a search engine will serve you censored, biased results that makes it an unreliable search engine.

And someone upthread claimed their image search was so great in comparison to google... because google also censors their results. They just censor different things.

No post body was provided.

Re: YaLM-100B: Pretrained language model with 100B parameters

#212
post #174

Earlier quoted context omitted.

Quoted post unavailable.

Yeah we definitely shouldn't worry about the political sympathies/vulnerabilities of the web services we use as the foundations of our shared knowledge...

Do you have the same level of concern about the leverage Five-Eyes intelligence agencies have over Facebook and Google?

Re: YaLM-100B: Pretrained language model with 100B parameters

#213

Earlier quoted context omitted.

no it's not. they straight up serve kremlin, promoting kremlin fake news and silencing russian opposition (not much to silence but still). they can have whatever functionality they like, I still won't use it in billion years.

Quoted post unavailable.

> yes russia bad

that is the point

Re: YaLM-100B: Pretrained language model with 100B parameters

#214
post #98

Earlier quoted context omitted.

Google absolutely does the same thing.

Care to elaborate?

There is lots of content Google bans/hides. Copyrighted content, Adult content, child pornography, official secrets, etc.

I don't think thats so different from other countries which also have a (partially overlapping) list of whats not allowed.

Normally, when people think about that they say "well pictures of naked children are morally wrong, whereas talking about LGBTQ stuff is fine". But people in other parts of the world might have different morals and might think the other way around.

Re: YaLM-100B: Pretrained language model with 100B parameters

#215
post #182

Earlier quoted context omitted.

I believe you're confusing the amount of A100 graphics cards used to train the model (the cluster was actually made up of 800 A100s), and the amount you need to run the model : > The model [...] is supposed to run on multiple GPUs with tensor parallelism. > It was tested on 4 (A100 80g) and 8 (V100 32g) GPUs, [but should work] with ≈200GB of GPU memory. I don't know what the price of a V100 is, but given $10k a piece…

The $10k price is for an A100 with 40GB ram, so you need 8 of those. If you can get your hands on the 80GB variant, 4 are enough. Also, if you want to have a machine with eight of these cards, it will need to be a pretty high-spec rack-mounted or large tower. To feed these GPUs, you will want to have a decent amount of PCIe-4 lanes, meaning EPYC are the logical choice. So that's $20k for an AMD EPYC server with at le…

Do you happen to know the cost of the 80GB variant?

Re: YaLM-100B: Pretrained language model with 100B parameters

#216
post #147

Earlier quoted context omitted.

Not for being Russians, but for active participation in censorship by tweaking their news aggregation to show only hand picked government approved sources

Is there a source for this? I'm curious.

Just read about Yandex.News which display censored news sources on Yandex frontpage that millions of people visit. There is really no hidden censorship here - they just follow Russian law that literally whitelist exclusively press controlled by the state.

They show Kremlin propoganda on their front page which makes Yandex part of Kremlin propoganda machine. They could have shut down news agregator, but they choose not to.

Re: YaLM-100B: Pretrained language model with 100B parameters

#218
post #174

Earlier quoted context omitted.

Quoted post unavailable.

Yeah we definitely shouldn't worry about the political sympathies/vulnerabilities of the web services we use as the foundations of our shared knowledge...

How do you feel about western-owned web services?

Re: YaLM-100B: Pretrained language model with 100B parameters

#219
post #163

Earlier quoted context omitted.

What have Israel connections got to do with shadiness? What are you insinuating?

Settling land that was recently taken from Palestinian families by force, often (literally) knocking the existing houses over with a bulldozer. In my opinion, that's a moral red flag to participate in such an atrocity.

[deleted]

Re: YaLM-100B: Pretrained language model with 100B parameters

#220
post #21

Seeing those gigantic models it makes me sad that even the 4090 is supposed to stay at 24GB of RAM max. I really would like to be able to run/experiment on larger models at home.

Take a look at Apple's M1 Max, a lot of fast unified memory. No idea how useful though

What's the difference between Apple's unified memory and the shared memory pool Intel and AMD integrated GPUs have had for years?

In theory you could probably assign a powerful enough iGPU a few hundred gigabytes of memory already, but just like Apple Silicon the integrated GPU isn't exactly very powerful. The difference between the M1 iGPU and the AMD 5700G is less than 10% and a loaded out system should theoretically be tweakable to dedicate hundreds of gigabytes of VRAM to it.

It's just a waste of space. An RTX3090 is 6 to 7 times faster than even the M1, and the promised performance increase of about 35% for the M2 will means nothing when the 4090 will be released this year.

I think there are better solutions for this. Leveraging the high throughput of PCIe 5 and resizable BAR support might be used to quickly swap out banks of GPU memory, for example, at a performance decrease.

One big problem with this is that GPU manufacturers have incentive to not implement ways for consumers GPUs to compete with their datacenter products. If a 3080 with some memory tricks can approach an A800 well enough, Nvidia might let a lot of profit slip through their hands and they can't have that.

Maybe Apple's tensor chip will be able to provide a performance boost here, but it's stuck on working with macOS and the implementations all seem proprietary so I don't think cross platform researchers will really care about using it. You're restricted by Apple's memory limitations anyway, it's not like you can upgrade their hardware.

Post reply on HN