Live data from Hacker News

OPT: Open Pre-trained Transformer Language Models

arxiv.org

111–120 of 242 posts

Re: OPT: Open Pre-trained Transformer Language Models

#111

Earlier quoted context omitted.

Well, the decision not to release the model might have been made so that they could license it instead.

They gave a reason why they didn't release it to the public, they said it was too dangerous: https://www.theguardian.com/technology/2019/feb/14/elon-musk... But of course then they started selling it to the highest bidder, so I wouldn't really trust what they say. They aren't "OpenAI", at this point they are just regular "ProprietaryAI". I really wonder what goal Elon Musk have with it.

> I really wonder what goal Elon Musk have with it.

You mean Sam Altman? Isn't he the CEO?

Re: OPT: Open Pre-trained Transformer Language Models

#112
post #93

Earlier quoted context omitted.

That happened after they decided not to release the model.

So their goal was to become the next IBM Watson? Parade around tech and try to create hype and hope for the future around it, while hiding all the dirty secrets that shows how limited the technology really is. Their original reasoning for not releasing it "this model is too dangerous to be released to the public" felt very much like a marketing stunt.

It does feel like the Tesla FSD playbook citing "pending regulatory approval"

Re: OPT: Open Pre-trained Transformer Language Models

#113

Earlier quoted context omitted.

They gave a reason why they didn't release it to the public, they said it was too dangerous: https://www.theguardian.com/technology/2019/feb/14/elon-musk... But of course then they started selling it to the highest bidder, so I wouldn't really trust what they say. They aren't "OpenAI", at this point they are just regular "ProprietaryAI". I really wonder what goal Elon Musk have with it.

> I really wonder what goal Elon Musk have with it. You mean Sam Altman? Isn't he the CEO?

Isn't Elon paying for it? I thought the original point was to democratize AI, ie the venture wasn't intended to make money but to help advance humanity, so it was funded by wealthy people who didn't need the money back. But maybe I just fell for their marketing?

Re: OPT: Open Pre-trained Transformer Language Models

#114
Remember when OpenAi wrote this?

> Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT-2 along with sampling code. We are not releasing the dataset, training code, or GPT-2 model weights

Well I guess Meta doesn’t care.

https://openai.com/blog/better-language-models/

Re: OPT: Open Pre-trained Transformer Language Models

#115
post #65

Earlier quoted context omitted.

So in the entire field of machine learning, we can't train a model that can identify another model from its output? Just can't be done? And there's absolutely no value in having tools that can identify deep fakes, or content produced by specific open models? >It's a bullshit term, firstoff, and calling yourself that is the height of ego I am a 10x engineer though, so I'm sorry if that rubs you the wrong way. Also, yo…

That is an interesting idea. The fact that they are characterizing the toxicity of the language relative or other LLMs gives it some credibility. That being said, I just don’t see where the ROI would be in something like that. Seems like a lot of expense for no payoff. My (unasked for) advice would be to take the 10x engineer stuff off your page. It may be true, but it signals the opposite. Much better to just let yo…

>That being said, I just don’t see where the ROI would be in something like that. Seems like a lot of expense for no payoff.

I consider these types of models as information weapons, so I wouldn't be surprised if they have some contract/agreement with the US government that they can only release these things to the internet if they have sufficient confidence in their ability to detect them, when they inevitably get used to attack the interests of the US and our allies. I don't know how (or even if) that translates to a financial ROI for Meta.

Re: OPT: Open Pre-trained Transformer Language Models

#116

Earlier quoted context omitted.

> I really wonder what goal Elon Musk have with it. You mean Sam Altman? Isn't he the CEO?

Isn't Elon paying for it? I thought the original point was to democratize AI, ie the venture wasn't intended to make money but to help advance humanity, so it was funded by wealthy people who didn't need the money back. But maybe I just fell for their marketing?

Elon hasn’t been involved for 2+ years. Didn’t like the direction afaik.

Re: OPT: Open Pre-trained Transformer Language Models

#117

"We are releasing all of our models between 125M and 30B parameters, and will provide full research access to OPT-175B upon request. Access will be granted to academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories." GPT-3 Davinci ("the" GPT-3) is 175B. The repository will be open "First thing in AM" ( https://twitter.com/stephe…

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

> Why do I have to request anything?

I'm guessing it could be one or a mix of these:

They want to build a database of people interested in this and vetted by some other organization as worth hiring. Just more people to feed to their recruiters.

To see the output of the work. While academics will credit their data sources, seeing "XXX from YYY" requested, and then later "YYY releases product that could be based on the model" is probably pretty valuable vs wondering which ML it was based on.

A veneer of responsible use, maybe required by their privacy policy or just to avoid backlash about "giving people's data away".

Re: OPT: Open Pre-trained Transformer Language Models

#118

"We are releasing all of our models between 125M and 30B parameters, and will provide full research access to OPT-175B upon request. Access will be granted to academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories." GPT-3 Davinci ("the" GPT-3) is 175B. The repository will be open "First thing in AM" ( https://twitter.com/stephe…

I don't like "available on request". I just want to download it and see if I can get it to run and mess around with it a bit. Why do I have to request anything? And I'm not an academic or researcher, so will they accept my random request? I'm also curious to know what the minimum requirements are to get this to run in inference mode.

A 175 billion parameter model might be a couple hundred gigs on disk. The file is probably just too big for GitHub/other standard FB services.

Re: OPT: Open Pre-trained Transformer Language Models

#119
post #59
post #17

Earlier quoted context omitted.

Ending up in the wild is an eventuality, whether FB creates it or someone else, why draw it out? Bandwidth concerns is nonsensical these days, fb has nearly unlimited resources in that department. Set it free! It wants to be free.

"It wants to be free" is a ridiculous statement, considering that after full two years (GPT-3 was published in May 2020), there is no public release of anything comparable. In May 2020, was your estimate of time to public release of anything comparable shorter or longer than two years? I bet it was shorter.

> "It wants to be free" is a ridiculous statement

"It wants to be free" is based on the standard line "code/data wants to be free". It doesn't mean this cost nothing to produce or isn't valuable.

Re: OPT: Open Pre-trained Transformer Language Models

#120
post #95

Earlier quoted context omitted.

Please don't start a profile analysis flamewar. It just escalates and makes everyone unhappy. I think it's OK if people notice you work at Facebook. There are people on HN that like to attack anyone nice enough to engage with them just because they work at a big company. I worked at Google for many years, and people were off to blame me personally for every decision that Google made that they didn't like. My approach…

To be clear, I wasn't intending to come across as attacking voz, only pointing out that I don't think anyone "in the know" at Meta/Facebook would admit to it even if they were doing it, so hearing "This is nonsense." doesn't really tell anybody much. They would likely say the same thing whether they thought it was nonsense or not.

No, they would likely not say anything. Explicitly denying it is saying something. But also - just to backup your claim how do you fingerprint a model? It seems logically impossible to me, if you are trying to mimic a certain intelligence, and you specifically "unmimic" it... then you may as well not try.
Post reply on HN