Live data from Hacker News

The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

thesequence.substack.com

91–100 of 527 posts

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#91
post #88

Earlier quoted context omitted.

It is just a copyright violation. My guess is that it would be fine if you use already scraped data as you haven't accepted TOS, but they have every right to block you or access to your business if you violate this.

I thought the copyright office said that ai generated material isn’t copyrighted?

You’re correct. US law states that intellectual property can be copyrighted only if it was the product of human creativity, and the USCO only acknowledges work authored by humans at present. Machines and generative AI algorithms, therefore, cannot be authors, and their outputs are not copyrightable.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#92
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

IANAL but I really don't see how a case here would go in OpenAI's favor in the long run, except maybe if someone actually agreed to their EULA?

And I really suspect that a lot of AI companies are putting out a lot of bluster about this and are just kind of hoping that nobody challenges them. Maybe LLaMA weights are copyrightable, but I would not take it as a given that they are.

I vaguely suspect (again IANAL) that companies like Facebook/OpenAI might not be willing to even force the issue, because they might be happier leaving it "unsettled" than going into a legal process that they're very likely to lose. I would love to see some challenges from organizations that have the resources to issue them and defend themselves.

Hiding behind the EULA is one thing, but there are a lot of people that have never signed that EULA.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#93

Is this a tactical leak, stemming from a "commoditize your complement" strategy? Open source as a strategic weapon, without having to explain board members/shareholders/whatever that you threw around money on training an open sourced model?

I would assume so. Meta’s ML/AI team is very strong, but they probably don’t have a comparable product offering to ChatGPT ready for public use. So instead, they bought themselves some time by letting the open source community run wild with a lesser model and eat into OpenAI’s moat.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#94

> OpenAI published a detailed blog post outlining some of the principles used to ensure safety in their models. The post emphasize in areas such as privacy, factual accuracy Am I the only one amused by the phrase “factual accuracy”? How many stories have we read like the one where it tries to ghost light the guy that this year is actually last year. “Oh, your phone must be wrong too, because there is no way I could b…

I find the thing incredibly smart and yet utterly useless at times.

I just spent 20 minutes getting the current iteration of ChatGPT to agree with me that a certain sentence is palindromic. Even when you make it print the unaccented characters one by one, spaces excluded, backwards and forwards, it still insists "Élu par cette crapule" isn't palindromic.

I understand how tokenization makes this difficult but come on... this doesn't feel like a difficult task for something that supposedly passes the LSATs and whatnot.

* French for "Elected by this piece of shit"

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#96
post #5

The "leak" is being portrayed as something highly subversive done by the darn 4chan hackers. Before the "leak" Meta was sending the model to pretty much anyone who claimed to be a PhD student or researcher and had a credible college email. Meta has probably been planning to release the model sooner than later. Let's hope they release it under a true open source license.

[dead]

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#97
post #23
post #10

I'm a bit worried the LLaMA leak will make the labs much more cautious about who they distribute models to for future projects, closing down things even more. I've had tons of fun implementing LLaMA, learning and playing around with variations like Vicuna. I learned a lot and probably wouldn't have got so interested in this space if the leak didn't happen.

They clearly expected the leak, they distributed it very widely to researchers. The important thing is the licence, not the access: you are not allowed to use it for commercial purpose.

How could Meta ever find out your private business is using their model without a whistleblower? It's practically impossible.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#98
post #26

Earlier quoted context omitted.

If the copyright office determines model weights are uncopyrightable (huge if), then one might imagine any institutional leak would benefit everyone else in the space. You might see hackers, employees, or contractors leaking models more frequently. And since models are distilled functionality (no microservices and databases to deploy), they're much easier to run than a constellation of cloud infrastructure.

Shouldn't that be the default position? The training methods are certainly patentable, but the actual input to the algorithm is usually public domain, and outputs of algorithms are not generally copyrightable as new works (think of to_lowercase(Harry Potter), which is not a copyrightable work), so the model weights would be a derivative work of public domain materials, and hence also forced into the public domain fro…

I like your legal interpretation, but it's way too early to tell if it is one that accurately represents the reality of the situation.

We won't know until this hits the courts.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#99
post #36

Earlier quoted context omitted.

An alternative interpretation was the LLaMa leak was an effort to shake or curtail the progress of ChatGPT's viral dominance at the time.

"And as long as they’re going to steal it, we want them to steal ours. They’ll get sort of addicted, and then we’ll somehow figure out how to collect sometime in the next decade". That was ironically Bill Gates https://www.latimes.com/archives/la-xpm-2006-apr-09-fi-micro...

It took him a while to come around

https://en.wikipedia.org/wiki/An_Open_Letter_to_Hobbyists

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#100

Is this a tactical leak, stemming from a "commoditize your complement" strategy? Open source as a strategic weapon, without having to explain board members/shareholders/whatever that you threw around money on training an open sourced model?

I would assume so. Meta’s ML/AI team is very strong, but they probably don’t have a comparable product offering to ChatGPT ready for public use. So instead, they bought themselves some time by letting the open source community run wild with a lesser model and eat into OpenAI’s moat.

They didn't leak it. Someone else did.
Post reply on HN