Live data from Hacker News

Llama and ChatGPT Are Not Open-Source

spectrum.ieee.org

71–80 of 130 posts

Re: Llama and ChatGPT Are Not Open-Source

#71
post #64

Earlier quoted context omitted.

Authors, artists, actors etc whose work have been ripped off by mega corporations like Meta, OpenAI, Google etc would disagree. Copyright may need to be updated to cater for the new world of AI but that doesn't mean as a concept it is evil.

How are Modern Copyright laws helping authors, artists, and actors in those cases?

It's allowing them to sue OpenAI for copyright infringement:

https://www.theguardian.com/books/2023/jul/05/authors-file-a...

Re: Llama and ChatGPT Are Not Open-Source

#72
post #70

Earlier quoted context omitted.

I believe LLMs should be allowed to read/view/consume content and learn from it even if that content has a copyright. We phrase it like somehow the material is being copied into the LLM, but that’s not what it’s doing. It’s building a neural graph from the experience of consuming that content. What would the world be like if humans couldn’t learn, train the weights of the interconnects of their neural tissue, from an…

It’s a form of lossy compression. Can I strip the copyright off an image by JPEG compressing it? At the very least I think LLMs trained on data that the trainer does not own or have rights to use in that manner should not be copyrightable.

An LLM is a lossy compression of the internet and I think it should be treated as such. You can't copyright the internet itself.

Re: Llama and ChatGPT Are Not Open-Source

#73
post #65

Earlier quoted context omitted.

It's shocking to me how many people in tech feel completely entitled to intellectual property that took someone years to master a skill to make. But talk about releasing a proprietary codebase and suddenly they want the lawyers involved because that actually threatens their livelihood.

Is that actually a commonly held position? I've seen IP abolishonosts here, and I've seen people argue the merits of proprietary software, but I don't get the impression that those are generally the same people.

They’re straw-manning, programmers are the best sharers in the world. Open source software has lead the drive for open source learning and information in general.

Re: Llama and ChatGPT Are Not Open-Source

#74

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

There is a lot of mental gymnastic, to submit an artwork publicly on the internet, allowing everyone to copy your arstyle, make derivative art of it, have other learn from it, but if ever a machine 'learn' from it, it's "stealing". It's not because you don't like a derivative work, that the derivative work is "stealing" your content. Saying it's stealing is wrong, it's lying to get your point accross.

You blame "how tech is going to steal everyone copyrighted works", yet, they were already in the tech world: the internet.

Re: Llama and ChatGPT Are Not Open-Source

#75

Earlier quoted context omitted.

Is that actually a commonly held position? I've seen IP abolishonosts here, and I've seen people argue the merits of proprietary software, but I don't get the impression that those are generally the same people.

They’re straw-manning, programmers are the best sharers in the world. Open source software has lead the drive for open source learning and information in general.

https://news.ycombinator.com/item?id=36900844

Right here the comment says they prefers Meta to keep the secret sauce as long as it allows they to (inderectly) access the copyrighted material etc. And the replies below are generally positive. So it's not a straw-man, at least in this thread, at all.

Re: Llama and ChatGPT Are Not Open-Source

#76
post #59

Earlier quoted context omitted.

Pretty much everything nowadays is copyrighted, by omitting such materials, what are you really left with? LLM is a tool much like the internet is a tool. Yes, someone can use it to steal, but stealing is against the law. Instead of encoding a criminal justice system into an LLM by omitting the possibility of stealing an artists work or omitting the knowledge of physics so someone can't learn how to build a bomb, we…

You're not addressing the massive abuse of the commons this still represents. If artists don't have the right to tell you to fuck off for using their work in training data, they're less likely to publicly show that work, which hurts them because they become less visible and hurts the AI because the training data gets worse.

Go on youtube and type "copy arstyle". Now tell me how artists were not stealing from each other.

Re: Llama and ChatGPT Are Not Open-Source

#77

Can we please stop using terminology related to code for something that's not code?

Except it's code. Weight decides how the model works so it's code. Code is just you telling the computer to work.

var a = b * 3 + 2;

You're telling me the * 3 + 2 part isn't code and only var a = b is code?

Re: Llama and ChatGPT Are Not Open-Source

#78

"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…

Does anyone have an idea of what (if anything) is being implied by the last two sentences?

Re: Llama and ChatGPT Are Not Open-Source

#79
post #64

Earlier quoted context omitted.

How are Modern Copyright laws helping authors, artists, and actors in those cases?

It's allowing them to sue OpenAI for copyright infringement: https://www.theguardian.com/books/2023/jul/05/authors-file-a...

It's worth noting you can sue for just about anything, but a case could end up being dismissed or you could simply lose it.

Re: Llama and ChatGPT Are Not Open-Source

#80

Earlier quoted context omitted.

And posted the link to your subscribers only medium page twice in this thread

devils advocate: if it wasn't subscribers only, google's crawler would be stealing it right now to train an ai with, ie, illegally creating derivative works based on their articles

Google's crawler can still crawl it. The regwall is only for humans:

https://archive.is/GmVKl

Also the article seems to be written by an LLM anywway, so it wouldn't matter in the end.

Post reply on HN