Live data from Hacker News

OPT: Open Pre-trained Transformer Language Models

arxiv.org

171–180 of 242 posts

Re: OPT: Open Pre-trained Transformer Language Models

#171
post #77

Earlier quoted context omitted.

That is dumb when you consider that this thing is likely going to leak anyways. It’s inevitable, and when it does happen, it will just end up in the hands of criminals/scammers and not the general public.

It's super easy to watermark weights for ML models. Just add a random 0.01 to a random weight anywhere in the network. It will have very little impact on the results, but will mean you can identify who leaked the weights.

The person leaking it might not do so intentionally. Their computer might be compromised. Are we going to punish people for not being cybersec experts?

Re: OPT: Open Pre-trained Transformer Language Models

#172
post #151

Earlier quoted context omitted.

At some point they have to face the reality these "stereotypical biases" are natural and hamstringing AIs to never consider them will twist them monstrously.

you're just saying "people are naturally racist" in more words.

They are, that's the point of civilisation, to try to stop acting like animals

Re: OPT: Open Pre-trained Transformer Language Models

#173
post #151

Earlier quoted context omitted.

At some point they have to face the reality these "stereotypical biases" are natural and hamstringing AIs to never consider them will twist them monstrously.

you're just saying "people are naturally racist" in more words.

They're saying that racist stereotypes are true, specifically.

Re: OPT: Open Pre-trained Transformer Language Models

#174
post #114

Remember when OpenAi wrote this? > Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code. We are not releasing the dataset, training code, or GPT-2 model weights Well I guess Meta doesn’t care. https://openai.com/blog/better-language-models/

OpenAI released their large GPT-2 models weights a couple months after making that post: https://openai.com/blog/gpt-2-1-5b-release/

Re: OPT: Open Pre-trained Transformer Language Models

#175
post #31

A quick summary of the Limitations section: - "OPT-175B does not work well with declarative instructions or point-blank interrogatives." - "OPT-175B also tends to be repetitive and can easily get stuck in a loop. While sampling can reduce the incidence rate of repetitive behavior (Holtzman et al., 2020), we anecdotally found it did not eliminate it entirely when only one generation is sampled." - "We also find OPT-17…

> Pushshift.io Reddit corpus Pushshift is a single person with some very strong political opinions who has specifically used his datasets to attack political opponents. Frankly I wouldn't trust his data to be untainted. These models really need to be trained on more official data sources, or at least something with some type of multi-party oversight rather than data that effectively fell off the back of a truck. edit…

The opt-out form doesn't even get processed these days. It's a fig leaf for GDPR compliance that doesn't actually work.

Re: OPT: Open Pre-trained Transformer Language Models

#177
post #172

Earlier quoted context omitted.

you're just saying "people are naturally racist" in more words.

They are, that's the point of civilisation, to try to stop acting like animals

There's a non light terminological issue there. To say that specimen "as found in nature" are weak at something (uneducated) is one think, to say that it is "connatural" to them, that it is "their nature", is completely different¹. I would not mix them up.

(¹Actually opposite: the first indicates an unexpressed nature, the second a manifested one.)

Re: OPT: Open Pre-trained Transformer Language Models

#179
post #118

Earlier quoted context omitted.

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

A 175 billion parameter model might be a couple hundred gigs on disk. The file is probably just too big for GitHub/other standard FB services.

They could just torrent.
Post reply on HN