Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

461–470 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#461
post #150
post #13

I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers. Have to wonder what the knock-on effects of that will be, especially if the models improve drastically. With so much of our social lives being moved online, if we have the easy ability to create fake lives of fake people one has to wonder what's real and what isn't. Maybe the dead internet theory will rea…

The bots/machine vs human reminds me of that famous experiment from the 30s in which Winthrop Kellogg[0], a comparative psychologist, and his wife decided to raise their human baby (Donald) simultaneously with a chimpanzee baby (Gua) in an effort to "humanize the ape". It was set out to last 5 years but was relatively quickly abrupted after only 9 months. The explicit reason wasn't stated only that it successfully pr…

My mind is blown. Thanks for sharing. Especially with the movie analogy. I’m a very movie person and I imitate my personality traits a lot based on characters on movies…

Re: YaLM-100B: Pretrained language model with 100B parameters

#462
post #13

I have to wonder if 10 years down the line, everyone will be able to run models like this on their own computers. Have to wonder what the knock-on effects of that will be, especially if the models improve drastically. With so much of our social lives being moved online, if we have the easy ability to create fake lives of fake people one has to wonder what's real and what isn't. Maybe the dead internet theory will rea…

Unpopular opinion: something will stop egalitarian power for the masses. I had high hopes for multicore computing in the late 90s and early 2000s but it got blocked every step of the way by everyone doubling down on DSP (glorified vertex buffer) approaches on video cards, leaving us with the contrived dichotomy we see today between CPU and GPU. Whatever we think will happen will not happen. A less-inspired known-good…

In what sense is the dichotomy between CPU and GPU contrived? Those are designed around fundamentally different use cases. For low power devices you can get CPU and GPU integrated into a single SOC.

Re: YaLM-100B: Pretrained language model with 100B parameters

#463
post #150

Earlier quoted context omitted.

The bots/machine vs human reminds me of that famous experiment from the 30s in which Winthrop Kellogg[0], a comparative psychologist, and his wife decided to raise their human baby (Donald) simultaneously with a chimpanzee baby (Gua) in an effort to "humanize the ape". It was set out to last 5 years but was relatively quickly abrupted after only 9 months. The explicit reason wasn't stated only that it successfully pr…

Case in point: recently, I've noticed that I'm getting more and more emails with the sign off "Warm regards." This is not a coincidence. It is an autosuggestion from Google. If you start signing off an email, it will automatically suggest "Warm regards." It just appears there -- probably an idea generated from an AI network. There are more and more of these algorithmic "suggestions" appearing every day, in more and m…

Which is why it's important for folks to start applying AI to more interesting (but harder, more nuanced) problems. Instead of making it easier for people to write emails, or targeting ads, it should be used to help doctors, surgeons and scientists.

The problem is that these problems are less profitable. And that the companies with enough compute to train these types of models are concerned about getting more eyeballs, not making the world a better place.

Re: YaLM-100B: Pretrained language model with 100B parameters

#464
post #433

Earlier quoted context omitted.

Yeah, you could go on, but you get paid per post, not per word.

Unmasked as a shill for saying that great powers engage in propaganda, false flagging and dissent crushing. I have no skin in the game. War, no war, it doesn't matter to me the outcome of this war to be honest. EDIT: But if you are all moral highground, answer me this: Why did the US goad Ukraine into taking a hostile stance against a neighbouring (and somewhat rival) great power? Whas this to the interest of Ukrania…

I don't accept that you can make this case:

"Why did you, W, goad X into expressing their sovereignty against Y? Didn't you know that Y would react with violence? That makes W the bad guy"

No, Y is always on the wrong side; you can't use the threat of violence and then claim via realpolitik that the other side was in the wrong. "Moral high ground" means you act out of principle, not political convenience. In this case, Ukraine didn't want to be in the Russian sphere, so we supported them.

And now yeah, the US is paying a lot of money and inconvenience to support Ukraine. Gas will be more expensive, we're spending tens of billions on weapons. But that's because it's the right thing to do; not every decision is a realpolitik game about maximizing revenue from vassal states (which I hope Russia will learn someday).

Re: YaLM-100B: Pretrained language model with 100B parameters

#465

Earlier quoted context omitted.

Aren't eleutherai's model so?

doesnt seem the code is there - pretrained models are there. https://github.com/kingoflolz/mesh-transformer-jax/#gpt-j-6b https://huggingface.co/EleutherAI/gpt-j-6B isnt that so ?

The code is there: https://github.com/EleutherAI/gpt-neox

Re: YaLM-100B: Pretrained language model with 100B parameters

#466

Earlier quoted context omitted.

Case in point: recently, I've noticed that I'm getting more and more emails with the sign off "Warm regards." This is not a coincidence. It is an autosuggestion from Google. If you start signing off an email, it will automatically suggest "Warm regards." It just appears there -- probably an idea generated from an AI network. There are more and more of these algorithmic "suggestions" appearing every day, in more and m…

Which is why it's important for folks to start applying AI to more interesting (but harder, more nuanced) problems. Instead of making it easier for people to write emails, or targeting ads, it should be used to help doctors, surgeons and scientists. The problem is that these problems are less profitable. And that the companies with enough compute to train these types of models are concerned about getting more eyeball…

The problem is not that those problems are less profitable. The problem is a combination of 1. Those problems are much harder 2. The potential harm from getting them wrong is much larger

Re: YaLM-100B: Pretrained language model with 100B parameters

#467
post #21

Seeing those gigantic models it makes me sad that even the 4090 is supposed to stay at 24GB of RAM max. I really would like to be able to run/experiment on larger models at home.

If you don't care about inference speed being in the 1-5sec range, then that should be doable with CPU offloading, with e.g. DeepSpeed.

200+ GiB of RAM still sounds like a pretty steep hardware requirement.

Re: YaLM-100B: Pretrained language model with 100B parameters

#468
post #436
post #425

Did they bias it toward ru propaganda talking points? Edit: I would like to see more details in addition to size and languages (en, ru) about training data. For example, did they use their own Yandex.news (a cesspool of propoganda)?

You've made a version of this comment 3 times in this thread now. It's shallow and flamebaity, and the repetition just adds noise and does no good, so please don't keep doing that. I understand the strong feelings, but the rules still apply—in fact that's when they apply most. https://news.ycombinator.com/newsguidelines.html

Thanks for reminder, I deleted other two comments which were more flamebaity. My overall point still stands - they did not give any details other than size on the training data. This is crucial (I train LLMs for a living)

Re: YaLM-100B: Pretrained language model with 100B parameters

#469

Earlier quoted context omitted.

> Yes, it got challenged AFTER needing to be leaked by a whistleblower that still can't return to his home. Good chance is that whistleblowing would be protected in this specific case.

Yes no, > Snowden was charged with theft, “unauthorized communication of national defense information” and “willful communication of classified communications intelligence information to an unauthorized person,” according to the complaint. The last two charges were brought under the 1917 Espionage Act.

The Espionage Act has no whistleblower protection. If the courts were allowed rule honestly and without political entanglements, there's no way the Espionage Act is constitutional at prima facia.

Re: YaLM-100B: Pretrained language model with 100B parameters

#470

Earlier quoted context omitted.

Google doesn't censor antiwar propaganda.

Blatantly incorrect. Google engages in egregious political censorship all the time. Including censorship for Russian government and censorship of US anti-war voices. https://reclaimthenet.org/youtube-responds-to-cpac-censorshi... https://reclaimthenet.org/google-expanded-its-censorship-of-... https://reclaimthenet.org/russia-continues-to-order-google-t... In US they pretend to "decide" to censor things "on their own"…

None of your links show Google censoring anti-war propaganda.
Post reply on HN