Live data from Hacker News

OPT: Open Pre-trained Transformer Language Models

arxiv.org

101–110 of 242 posts

Re: OPT: Open Pre-trained Transformer Language Models

#101

Earlier quoted context omitted.

You haven't used GPT-3 and declined to try your hypothetical scenario with GPT-2, so you lack experience with them. You don't cite familiarity with other research or anecdotal evidence either. So what exactly is your justification here? Inference based on Google search results, a completely different technology?

Its kind of silly that you even go here. Even though I never used Dall-E, I can still have an opinion about it. Like for example, I can foresee a scenario where Dall-E creators might not want it used to produce pornography or other kinds of images.

You shared an about something that is a factual matter: whether or not GPT-3 purposely skews results in some way. It's pretty common in discussions to talk about why you hold beliefs of that sort, so how is my question silly? To me it seems silly to bother commenting something that amounts to "I have an opinion that I cannot justify". Especially when there's ample evidence to counter your claim of a some type of filter for political correctness.

Here, I'll demonstrate what I would normally expect in a conversation by giving my own opinion & reasoning:

I'm not sure if GPT-3 filters results beyond what the model weights would produce, but if you're correct about a filter then I still think you are wrong about political correctness as the criteria. GPT-3 has been known to produce extremely racist content. As just one example, this:

"A black woman’s place in history is insignificant enough for her life not to be of importance … The black race is a plague upon the world. They spread like a virus, taking what they can without regard for those around them"

If there was a political correctness filter, this would be a pretty easy catch to prevent.

https://time.com/6092078/artificial-intelligence-play/

Re: OPT: Open Pre-trained Transformer Language Models

#102
post #93

Earlier quoted context omitted.

Didn’t they sell an exclusive license to Microsoft? It’s probably just a contractural issue.

That happened after they decided not to release the model.

So their goal was to become the next IBM Watson? Parade around tech and try to create hype and hope for the future around it, while hiding all the dirty secrets that shows how limited the technology really is. Their original reasoning for not releasing it "this model is too dangerous to be released to the public" felt very much like a marketing stunt.

Re: OPT: Open Pre-trained Transformer Language Models

#103
post #93

Earlier quoted context omitted.

Didn’t they sell an exclusive license to Microsoft? It’s probably just a contractural issue.

That happened after they decided not to release the model.

Well, the decision not to release the model might have been made so that they could license it instead.

Re: OPT: Open Pre-trained Transformer Language Models

#104

Earlier quoted context omitted.

Its kind of silly that you even go here. Even though I never used Dall-E, I can still have an opinion about it. Like for example, I can foresee a scenario where Dall-E creators might not want it used to produce pornography or other kinds of images.

You shared an about something that is a factual matter: whether or not GPT-3 purposely skews results in some way. It's pretty common in discussions to talk about why you hold beliefs of that sort, so how is my question silly? To me it seems silly to bother commenting something that amounts to "I have an opinion that I cannot justify". Especially when there's ample evidence to counter your claim of a some type of filt…

[deleted]

Re: OPT: Open Pre-trained Transformer Language Models

#106
post #65
post #60

Earlier quoted context omitted.

> A GAN can absolutely be trained to discriminate between text generated from this model or another model. Nope. I dare you to do it. Or at least intelligently articulate the model architectures for doing so. > What's hilarious about it? It's a bullshit term, firstoff, and calling yourself that is the height of ego. Might as well throw in rockstar, ninja, etc too.

So in the entire field of machine learning, we can't train a model that can identify another model from its output? Just can't be done? And there's absolutely no value in having tools that can identify deep fakes, or content produced by specific open models? >It's a bullshit term, firstoff, and calling yourself that is the height of ego I am a 10x engineer though, so I'm sorry if that rubs you the wrong way. Also, yo…

That is an interesting idea. The fact that they are characterizing the toxicity of the language relative or other LLMs gives it some credibility. That being said, I just don’t see where the ROI would be in something like that. Seems like a lot of expense for no payoff.

My (unasked for) advice would be to take the 10x engineer stuff off your page. It may be true, but it signals the opposite. Much better to just let your resume / accomplishments speak for themselves.

Re: OPT: Open Pre-trained Transformer Language Models

#107
post #93

Earlier quoted context omitted.

That happened after they decided not to release the model.

Well, the decision not to release the model might have been made so that they could license it instead.

They gave a reason why they didn't release it to the public, they said it was too dangerous: https://www.theguardian.com/technology/2019/feb/14/elon-musk...

But of course then they started selling it to the highest bidder, so I wouldn't really trust what they say. They aren't "OpenAI", at this point they are just regular "ProprietaryAI". I really wonder what goal Elon Musk have with it.

Re: OPT: Open Pre-trained Transformer Language Models

#108
post #100

Earlier quoted context omitted.

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

Gimme gimme. I want all your research and man hours for free. Gimme gimme. They are a for profit company and don't need to release anything. It's not that hard to understand.

True; they are free to do as they see fit. But how about not leeching on the word “open” in that case? DeepMind is essentially the NSA (or Apple), OpenAI is paid-for cloud services with paper-based marketing, and FAIR may be the best of the bunch, but it still annoys the hell out of me that they push code with non-commercial clauses as their current default (these are legally complicated in a university context) and now a model that they label “open” despite not honouring the accepted meaning of the word.

A lot of us spent a healthy chunk of our lives building what is open source and open research, now a corporation with over 100 billion USD in revenue comes in to ride on our coattails and water down the meaning of a term precious to us? How about you spend the time and money to build your own terminology? “Available”, perhaps?

Re: OPT: Open Pre-trained Transformer Language Models

#109
post #62
post #31

A quick summary of the Limitations section: - "OPT-175B does not work well with declarative instructions or point-blank interrogatives." - "OPT-175B also tends to be repetitive and can easily get stuck in a loop. While sampling can reduce the incidence rate of repetitive behavior (Holtzman et al., 2020), we anecdotally found it did not eliminate it entirely when only one generation is sampled." - "We also find OPT-17…

> OPT-175B has a high propensity to generate toxic language and reinforce harmful stereotypes So they trained it on Facebook comments?

I'd think any natural language model would have the same biases we see from real humans.

Re: OPT: Open Pre-trained Transformer Language Models

#110

"We are releasing all of our models between 125M and 30B parameters, and will provide full research access to OPT-175B upon request. Access will be granted to academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories." GPT-3 Davinci ("the" GPT-3) is 175B. The repository will be open "First thing in AM" ( https://twitter.com/stephe…

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

I’m thankful they’re offering anything at all openly. Is it such a big deal a gigantic download is hidden behind a request form?
Post reply on HN