Live data from Hacker News

OPT: Open Pre-trained Transformer Language Models

arxiv.org

81–90 of 242 posts

Re: OPT: Open Pre-trained Transformer Language Models

#81

Earlier quoted context omitted.

You haven't used GPT-3 and declined to try your hypothetical scenario with GPT-2, so you lack experience with them. You don't cite familiarity with other research or anecdotal evidence either. So what exactly is your justification here? Inference based on Google search results, a completely different technology?

Its kind of silly that you even go here. Even though I never used Dall-E, I can still have an opinion about it. Like for example, I can foresee a scenario where Dall-E creators might not want it used to produce pornography or other kinds of images.

[deleted]

Re: OPT: Open Pre-trained Transformer Language Models

#82

Earlier quoted context omitted.

GPT3 will do that right now. There aren’t any controls on its text, it just warns you if it looks offensive. And of course nothing it says is true except coincidentally. If you've seen GPT-3 interviews ( https://twitter.com/minimaxir/status/1513957106868637696 ) it'll happily say some wild stuff. As a mild example I recommend interviewing "a man who is currently beating you up".

Is it true to say they are true coincidentally, because that kind of suggests randomly true. I understand the AI doesn't really comprehend if something is true or false. My understanding is the results are more than random, maybe something closer to like weighted opinion.

Weighted random is still random.

Re: OPT: Open Pre-trained Transformer Language Models

#83
post #17

Earlier quoted context omitted.

Couple of random ideas: - They are concerned about the usage of the largest model, so want to vet people - The 175B parameter model is so large that it doesn't play nice with GitHub or something along those lines

Ending up in the wild is an eventuality, whether FB creates it or someone else, why draw it out? Bandwidth concerns is nonsensical these days, fb has nearly unlimited resources in that department. Set it free! It wants to be free.

In big companies, something as simple as "host it on facebook.com/model.tar.gz" can be mountains of approval and paperwork.

Re: OPT: Open Pre-trained Transformer Language Models

#84
post #75

Earlier quoted context omitted.

That would be interesting if it was true, but I think it can’t be true because LLMs main advantage is they memorize text in their weights and so your discriminator model would need to be the same size as the LLM. That said the smaller GPT3 models break down quite often so they’re probably detectable.

In the same way we can train models that can identify people from their choice of words, phrasing, grammar, etc, we can train models that identify other models.

That's anthropomorphizing them - a large language model doesn't have a bottleneck the same way a human does (in terms of being able to express things), it can get on a path where it just outputs memorized text directly and it won't be consistent with what it usually seems to know at all.

Also, you could break a discriminator model by running a filter over the output that changes a few words around or misspells things, etc. Basically an adversarial attack.

Re: OPT: Open Pre-trained Transformer Language Models

#85

The big one, OPT-175B, isn't an open model. The word "open" in technology means that everyone has equal access (viz. "open source software" and "open source hardware"). The article says that research access will be provided upon request for "academic researchers; those affiliated with organizations in government, civil society, and academia; and those in industry research laboratories.". Don't assume any good intent…

[deleted]

Re: OPT: Open Pre-trained Transformer Language Models

#87
post #43
post #38

Earlier quoted context omitted.

It's pretty simple. GPT models are essentially information weapons. People are going to get their hands on them, so might as well give them a model where you can identify content generated with them, so you can know who is using them for nefarious purposes. Like how many printers encode hidden patterns on paper that identify the model of the printer and other information[0] 0. https://www.bbc.com/future/article/20170…

How can you identify content generated with them?

I'm not saying that Meta did it, but recent research shows that it is possible and hard to detect - https://arxiv.org/abs/2204.06974 - so if they really wanted to, they could.

Re: OPT: Open Pre-trained Transformer Language Models

#88
post #55
post #51

Earlier quoted context omitted.

Would an AI @ FB employee admit it if it was true?

> I will never discuss FB technical details, internals, or anything else on this site, so please do not ask. My claim of nonsense has nothing to do with FB. You cannot fingerprint models like this, that's just not how it works. Also, if we are reading profiles, you call yourself a 10x engineer on your blog, that's hilarious. Maybe 10x the nonsense?

Please don't start a profile analysis flamewar. It just escalates and makes everyone unhappy.

I think it's OK if people notice you work at Facebook. There are people on HN that like to attack anyone nice enough to engage with them just because they work at a big company. I worked at Google for many years, and people were off to blame me personally for every decision that Google made that they didn't like. My approach was to just say, look, the CEO didn't ask me, and if they did I would have said no. If you have concerns with something I actually work on, I'd love to adjust it based on your feedback. (That was network monitoring for Google Fiber, and wasn't very controversial. But, HN loves to lay in to you if you open yourself up for it. I learned a lot about people.)

In this case, I think the best you can do is to say "I don't think it's possible to add fingerprinting, and if it were, I would fight to not add it. I also don't know of any decision to add fingerprinting, and like I said, I would try to make sure we didn't do it." (Or if you're in favor and it's not technically possible, you could say that too!)

Anyway, it is really nice to hear from people "in the trenches". Please don't let people being toxic scare you away or bait you into a flamewar. Comments like yours remind us that even in these big companies whose political decision we may not like, there are still people doing really good engineering, and that's always fun to hear about.

Re: OPT: Open Pre-trained Transformer Language Models

#89
post #37

I often wonder if OpenAIs decision not to open gpt-3 was because it was to expensive to train relative to its real value. They’ve hidden the model behind an api where they can filter out most of the dumb behaviors, while everyone believes they are working on something entirely different.

Didn’t they sell an exclusive license to Microsoft? It’s probably just a contractural issue.

Re: OPT: Open Pre-trained Transformer Language Models

#90

Earlier quoted context omitted.

> The second link returned on him was from ADL. No way that's an organic result. It might be, actually. I understand why you'd think that, but look at the results for other search engines. Kagi: ADL in 2nd place Bing: ADL in 3rd place Yandex: ADL not on the first page, but SPLC[1] is the the 6th result [1]: https://www.splcenter.org/fighting-hate/extremist-files/indi...

Guess it depends on the "algorithm" but if we were still in the PageRank era there's no way in hell ADL or SLPC would be anywhere near the top results for "Alex Jones", considering how many other news stories, blogs, comments, etc. about him exist.

The PageRank era ended almost immediately. Google has had a large editorial team for a long, long time (probably before they were profitable).

It turns out PageRank aways kind of sucked. However, it was competing with sites that did “pay for placement” for the first page or two, so it only had to be better than “maliciously bad”.

Post reply on HN