Live data from Hacker News

The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

thesequence.substack.com

71–80 of 527 posts

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#71
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us. Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that? I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have som…

[deleted]

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#72

It appears there is this genre of articles pretending that LLAMA or its RL-HF tuned variants are somehow even close to an alternative to ChatGPT. Spending more than a few moments interacting even with the larger instruct-tuned variants of these models quickly dispels that idea. Why do these takes around open-source AI remain so popular? What is the driving force?

I had the same reaction after seeing lots of "chatgpt on a phone" etc hype around alpaca. Like I knew it wouldn't be close, but was surprised at just how useless it was given the noise around it. Nobody who was talking about it had used it for even five minutes.

This article is almost criminally imprecise around the "leak" and "Open Source model" discussion as well.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#73
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

> If web scraping is legal Source? That LinkedIn case did not resolve how you think it did.

The judgement of the LinkedIn case was that if the scraping bots had 'clicked the button' to accept terms then they should be held to those terms.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#74
post #5

The "leak" is being portrayed as something highly subversive done by the darn 4chan hackers. Before the "leak" Meta was sending the model to pretty much anyone who claimed to be a PhD student or researcher and had a credible college email. Meta has probably been planning to release the model sooner than later. Let's hope they release it under a true open source license.

It sounds like that king that wanted people to overcome their aversion for potatoes. So he put armed guards around the potato fields but instructed them to be very lax and allowed the people to rob it

Tell me more. Real or anecdote?

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#75

Earlier quoted context omitted.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us. Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that? I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have som…

We decided that animals can't create copyrightable works and hence limited the ability to create copyrightable works to humans . I am fine with granting AIs the ability to create copyrightable works provided we grant that right, and human rights, to Orcas and other intelligent species.

Animals seem ok with it. At least they did not indicate otherwise so far.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#76

Earlier quoted context omitted.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us. Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that? I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have som…

Copyright is a practical right, not an inherent right. The only reasons humans get copyright at all is because it's useful for society to give it to them. The onus should be on OpenAI to prove that it will benefit society overall if AIs are given copyright. We've already decided that many non-human processes/entities don't get copyright because there doesn't seem to be any reason to grant those entities copyright. --…

> The comparison to humans is interesting though, because teaching a human how to do something doesn't grant you copyright over their output.

Ehh, in rare cases in can though. If you have someone sign an NDA, they can't go and publish technical details about something confidential that they were trained on. For example, this is fairly common in the tech industry when we send engineers to train on proprietary hardware or software.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#77
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

what's weird to me though, is that we're all trained on both open source and closed source source material. And our output is totally 100% copyrightable by us. Why wouldn't we extend the same muster to computer generated text. If there is a copy-written sentence, go after that? I don't work for openai, but I don't like 1 sided arguments that are just looking for some bottom line. At the end of the day we all have som…

Let's say I were to create an algorithm which generated every possible short story in the English language using Markov chains. Should I be able to copyright all those generated stories, thus legally preventing any other author from ever writing a story again?

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#78
post #42

Earlier quoted context omitted.

The models are a lot of fun to play with, but yeah, every time I've tried to use them for something "serious" they nearly always invent stuff (and are so convincing in how they write about it!). Most recently I've been interested in what's happened with the 4-color theorem since the 1976 computer-assisted proof, and decided to use GPTChat instead of google+wikipedia. GPTChat had me convinced and excited that, apparen…

Before the inevitable idiots come in to say hurr durr but have you tried ChatGPT 4… yes I paid for it, and it is just as prone to hallucinations of factual information. It loves to make up new names for peoples initials.

I found the opposite to be true, i mean sure if youre tricking it. Wait for GPT 5-6 in a year or two and see haha.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#79
[Edited] Isn’t the copyright question a red-Hering? We are talking about models on the verge of generating output not distinguishable from human output. How is then a copyright breach - if it’s only caused by AI, but not by human - enforced long term?

I’m not in favor of the 6 month moratorium- but seriously, we are going to face tough questions very soon - and they will shake a lot of assumptions we have.

We should now really act as society to get standards in place, standards that are enforceable. Otherwise the LeCun’s et al. Will have some pretty bad impact before we start doing something.

We need to work on this globally and fast to not screw it up. I’m nowadays more worried than ever about elections in the near future. Maybe we will have something like real IDs attached to content (First useful use case for crypto) or maybe we will all stop getting information from people we don’t know (yay filter bubble). I hope people smarter than me will find something.

Re: The LLama Effect: Leak Sparked a Series of Open Source Alternatives to ChatGPT

#80
post #53
post #39

Someone needs to legally challenge openAI on using the output of their models to train other commercial models. If web scraping is legal, then this must be legal too , even if openAI tries to curtail it. After all it was all trained on data they don't have rights to.

Website content can be copyrightable, so web scraping for commercial use being legal would be dubious. But even OpenAI can't tell what ChatGPT will output, so I don't see how this can be copyrightable. Should the outputted sentences really be owned by OpenAI?

They are not claiming copyright on the output, but instead make it a part of their terms of use, so it's basically the EULA debate all over again.
Post reply on HN