Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

431–440 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#431

Earlier quoted context omitted.

Google doesn't censor antiwar propaganda.

Blatantly incorrect. Google engages in egregious political censorship all the time. Including censorship for Russian government and censorship of US anti-war voices. https://reclaimthenet.org/youtube-responds-to-cpac-censorshi... https://reclaimthenet.org/google-expanded-its-censorship-of-... https://reclaimthenet.org/russia-continues-to-order-google-t... In US they pretend to "decide" to censor things "on their own"…

> This is either astounding ignorance or blatant gaslighting.

Can you please edit name-calling / swipes like that out of your HN comments? It breaks the site guidelines and weakens your point.

https://news.ycombinator.com/newsguidelines.html

Re: YaLM-100B: Pretrained language model with 100B parameters

#432

I am one of the people who worked on Google's PaLM model. Having skimmed the GitHub readme and medium article, this announcement seems to be very focused on the number of parameters and engineering challenges scaling the model, but it does not contain any details about the model, training (learning rate schedules, etc.), or data composition. It is great that more models are getting released publicly, but I would not…

Given that Yandex is a crucial part of Russian propaganda arm, we should consider the whole range of possibilities from:

* Good. This is great researchers helping community by sharing great work. (which is what I'd like to assume before I have any proof of the contrary)

* Bad. This very expensive training has been approved by Ya leadership (which is under Western personal sanctions) because they've secretly built in RU's propaganda talking points into the model. Such as "war in Ukraine is not a war but special operation" etc.

Re: YaLM-100B: Pretrained language model with 100B parameters

#433
post #409

Earlier quoted context omitted.

>I reject the false equivalence of the DHS and FSB. Not gonna both-sides this, sorry. lmao mkay. Not identical, but very similar. It's not even 'Alex Jones'-tier to say this. I think you forget you are if you are under US or (even NATO). YOU WILL hear propaganda from your side, as the Russians do. It's NORMAL. We live under control of a hegemon with self-interests. May I have to remind you of these? And tell me the d…

Yeah, you could go on, but you get paid per post, not per word.

Unmasked as a shill for saying that great powers engage in propaganda, false flagging and dissent crushing.

I have no skin in the game. War, no war, it doesn't matter to me the outcome of this war to be honest.

EDIT: But if you are all moral highground, answer me this:

Why did the US goad Ukraine into taking a hostile stance against a neighbouring (and somewhat rival) great power? Whas this to the interest of Ukranians? Or to the geopolitical interests of US? https://www.youtube.com/watch?v=93eyhO8VTdg

Re: YaLM-100B: Pretrained language model with 100B parameters

#434

Earlier quoted context omitted.

>if you are worried about misuses why is morality into this? is this the same discussion of car manufacturers not selling cars to certain people because they are worried about misuse?

Automotive companies, in fact, have product liability. It's about liability, not morality.

when you release a project into the wild under a permissive license, aren't you essentially washing yourself from any "liability" ?

> MIT " IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE."

don't commercial licenses have same/similar wording so what liability are you talking about?

Re: YaLM-100B: Pretrained language model with 100B parameters

#435

Earlier quoted context omitted.

This is a stupid argument. Ukrainians are being killed and you compare that to fear of being arrested? It's a nice excuse. "Oh yes I don't support my government, but you know these arrests, I'd rather stay in my cosy home and enjoy my tea. Now could you please lift the sanctions? I already said I don't support my government, why normal people like me should suffer?" etc etc

Wait until you hear about the folks your country is killing!

https://en.wikipedia.org/wiki/And_you_are_lynching_Negroes

Re: YaLM-100B: Pretrained language model with 100B parameters

#436
post #425

Did they bias it toward ru propaganda talking points? Edit: I would like to see more details in addition to size and languages (en, ru) about training data. For example, did they use their own Yandex.news (a cesspool of propoganda)?

You've made a version of this comment 3 times in this thread now. It's shallow and flamebaity, and the repetition just adds noise and does no good, so please don't keep doing that. I understand the strong feelings, but the rules still apply—in fact that's when they apply most.

https://news.ycombinator.com/newsguidelines.html

Re: YaLM-100B: Pretrained language model with 100B parameters

#437
post #45

Earlier quoted context omitted.

Well... I'm sorry if I reach for the reductio at Hitlerum, but any achievements Nazi scientists might have reached in concentration camps are definitely tainted. Similarly, achievements in the field of online consumer analysis in a country where consumer-privacy protections are nonexistent, surely should be considered tainted...?

And Yandex's AI work got helped by the Russian invasion of Ukraine how, exactly? Did they train the bots on Ukrainian captives first?

Yandex search, which works on top of their AI is straight out propaganda machine for Russian government. Every time you go to yandex.ru, you're greeted with curated happy news about how Ukrainians are killing themselves, and russians are not fascists at all.

Their government does. They empower it. Just to be clear, it is their army, and their government doing the deed. They elected, they pay salaries to.

Re: YaLM-100B: Pretrained language model with 100B parameters

#438

Earlier quoted context omitted.

This is a stupid argument. Ukrainians are being killed and you compare that to fear of being arrested? It's a nice excuse. "Oh yes I don't support my government, but you know these arrests, I'd rather stay in my cosy home and enjoy my tea. Now could you please lift the sanctions? I already said I don't support my government, why normal people like me should suffer?" etc etc

Yeah, by being arrested you're doing so much more for the Ukrainians

More people having to deal with prisoners, potentially fewer people going to Ukraine. Certainly better than doing nothing.

Re: YaLM-100B: Pretrained language model with 100B parameters

#439
post #150
post #13

I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers. Have to wonder what the knock-on effects of that will be, especially if the models improve drastically. With so much of our social lives being moved online, if we have the easy ability to create fake lives of fake people one has to wonder what's real and what isn't. Maybe the dead internet theory will rea…

The bots/machine vs human reminds me of that famous experiment from the 30s in which Winthrop Kellogg[0], a comparative psychologist, and his wife decided to raise their human baby (Donald) simultaneously with a chimpanzee baby (Gua) in an effort to "humanize the ape". It was set out to last 5 years but was relatively quickly abrupted after only 9 months. The explicit reason wasn't stated only that it successfully pr…

thanks for the awesome analogy, I always had the sinking feeling that the bots are finding it increasingly easy to fit in among the humans because the humans on social media act increasingly like bots.

"monkey see, monkey do"

Re: YaLM-100B: Pretrained language model with 100B parameters

#440
post #254

Earlier quoted context omitted.

Fair point, but you’ll still be bounded by disk read speed on an SSD. The access pattern itself matters less than the read cache being << the parameter set size.

Top SSDs do over 4GB/s so you can infer in 50 seconds if disk bound. You can also infer a few tokens at once, so it will be more than 1 char a minute. Probably more like sentence a minute.

You can read bits at that rate yes, but keep in mind that it’s 250 GiB /parameters/, and matrix-matrix multiplication is typically somewhere between quadratic and cubic in complexity. Then you get to wait for the page out of your intermediate result etc etc.

It’s difficult to estimate how slow it would be, but I’m guessing unusably slow.

Post reply on HN