Live data from Hacker News

OPT: Open Pre-trained Transformer Language Models

arxiv.org

41–50 of 242 posts

Re: OPT: Open Pre-trained Transformer Language Models

#41

Earlier quoted context omitted.

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

Couple of random ideas: - They are concerned about the usage of the largest model, so want to vet people - The 175B parameter model is so large that it doesn't play nice with GitHub or something along those lines

> - The 175B parameter model is so large that it doesn't play nice with GitHub or something along those lines

There is no frickin' way that the difficulty or cost of distributing the model is a factor, even if it was several dozen terabytes in size (and it is probably somewhere around 1.5 terabytes). Not for Meta, and not when CDNs and torrrents are available as options.

If they are gatekeeping access to the model, there is no need to ascribe it to a side effect of something else. Their intent IS to limit access to the full model. I'm not really sure why they are bothering, unless they're assuming that unsavory actors won't be motivated enough to pay some grad student for a copy.

I suppose they may be adding a fingerprint or watermark of some sort to trace illicit copies back to the source if they're serious about limiting redistribution, but those can usually be found and removed if you have copies from two or more different sources.

Re: OPT: Open Pre-trained Transformer Language Models

#42
The big one, OPT-175B, isn't an open model. The word "open" in technology means that everyone has equal access (viz. "open source software" and "open source hardware"). The article says that research access will be provided upon request for "academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories.".

Don't assume any good intent from Facebook. This is obviously the same strategy large proprietary software companies have been using for a long time to reinforce their monopolies/oligopolies. They want to embed themselves in the so-called "public sector" (academia and state institutions), so that they get free advertising for taxpayer money. Ordinary people like most of us here won't be able to use it despite paying taxes.

Some primary mechanisms of this advertising method:

1. Schools and universities frequently use the discounted or gratis access they have to give courses for students, often causing students to be only specialized in the monopolist's proprietary software/services.

2. State institutions will require applicants to be well-versed in monopolist's proprietary software/services because they are using it.

3. Appearance of academic papers that reference this software/services will attract more people to use them.

Some examples of companies utilizing this strategy:

Microsoft - Gives Microsoft Office 365 access for "free" to schools and universities.

Mathworks - Gives discounts to schools and universities.

Autodesk (CAD software) - Gives gratis limited-time "student" (noncommercial) licenses.

Altium (EDA software) - Gives gratis limited-time licenses to university students.

Cadence (EDA software) - Gives a discount for its EDA software to universities.

EDIT: Previously my first sentence stated that the models aren't open - in fact, only OPT-175B is not (but the other ones are much smaller).

Re: OPT: Open Pre-trained Transformer Language Models

#43
post #38
post #32

Earlier quoted context omitted.

What are you talking about?

It's pretty simple. GPT models are essentially information weapons. People are going to get their hands on them, so might as well give them a model where you can identify content generated with them, so you can know who is using them for nefarious purposes. Like how many printers encode hidden patterns on paper that identify the model of the printer and other information[0] 0. https://www.bbc.com/future/article/20170…

How can you identify content generated with them?

Re: OPT: Open Pre-trained Transformer Language Models

#44

Earlier quoted context omitted.

You think GPT-3 generates text that's truthful? Have you used it even once?

I haven't used GPT-3, but I did try out a site that was based on GPT2. I believe it was called "talk to transformer". But I never tried quarrying anything controversial. However, I bet this a concern and certain queries will be filtered or "corrected" to be more politically correct. To give you an example, a few days ago I made a comment one Alex Jones, and wanted to google him. The second link returned on him was fr…

> The second link returned on him was from ADL. No way that's an organic result.

It might be, actually. I understand why you'd think that, but look at the results for other search engines.

Kagi: ADL in 2nd place

Bing: ADL in 3rd place

Yandex: ADL not on the first page, but SPLC[1] is the the 6th result

[1]: https://www.splcenter.org/fighting-hate/extremist-files/indi...

Re: OPT: Open Pre-trained Transformer Language Models

#45
post #43
post #38

Earlier quoted context omitted.

It's pretty simple. GPT models are essentially information weapons. People are going to get their hands on them, so might as well give them a model where you can identify content generated with them, so you can know who is using them for nefarious purposes. Like how many printers encode hidden patterns on paper that identify the model of the printer and other information[0] 0. https://www.bbc.com/future/article/20170…

How can you identify content generated with them?

By training a GAN. A trained GAN will be able to accurately guess whether a block of text was produced by this GPT model, some other GPT model, or is authentic.

Re: OPT: Open Pre-trained Transformer Language Models

#47

Earlier quoted context omitted.

I haven't used GPT-3, but I did try out a site that was based on GPT2. I believe it was called "talk to transformer". But I never tried quarrying anything controversial. However, I bet this a concern and certain queries will be filtered or "corrected" to be more politically correct. To give you an example, a few days ago I made a comment one Alex Jones, and wanted to google him. The second link returned on him was fr…

> The second link returned on him was from ADL. No way that's an organic result. It might be, actually. I understand why you'd think that, but look at the results for other search engines. Kagi: ADL in 2nd place Bing: ADL in 3rd place Yandex: ADL not on the first page, but SPLC[1] is the the 6th result [1]: https://www.splcenter.org/fighting-hate/extremist-files/indi...

This logic kind of fails quickly. I bet you wouldn't use it to show that Tiananmen Square did not happen, by showing all Chinese Search Engine are in apparent agreement on it not happening.

Re: OPT: Open Pre-trained Transformer Language Models

#49

Earlier quoted context omitted.

I haven't used GPT-3, but I did try out a site that was based on GPT2. I believe it was called "talk to transformer". But I never tried quarrying anything controversial. However, I bet this a concern and certain queries will be filtered or "corrected" to be more politically correct. To give you an example, a few days ago I made a comment one Alex Jones, and wanted to google him. The second link returned on him was fr…

> The second link returned on him was from ADL. No way that's an organic result. It might be, actually. I understand why you'd think that, but look at the results for other search engines. Kagi: ADL in 2nd place Bing: ADL in 3rd place Yandex: ADL not on the first page, but SPLC[1] is the the 6th result [1]: https://www.splcenter.org/fighting-hate/extremist-files/indi...

Guess it depends on the "algorithm" but if we were still in the PageRank era there's no way in hell ADL or SLPC would be anywhere near the top results for "Alex Jones", considering how many other news stories, blogs, comments, etc. about him exist.

Re: OPT: Open Pre-trained Transformer Language Models

#50
post #38
post #32

Earlier quoted context omitted.

What are you talking about?

It's pretty simple. GPT models are essentially information weapons. People are going to get their hands on them, so might as well give them a model where you can identify content generated with them, so you can know who is using them for nefarious purposes. Like how many printers encode hidden patterns on paper that identify the model of the printer and other information[0] 0. https://www.bbc.com/future/article/20170…

This is nonsense.
Post reply on HN