Live data from Hacker News

Llama and ChatGPT Are Not Open-Source

spectrum.ieee.org

81–90 of 130 posts

Re: Llama and ChatGPT Are Not Open-Source

#81
post #74

Earlier quoted context omitted.

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

There is a lot of mental gymnastic, to submit an artwork publicly on the internet, allowing everyone to copy your arstyle, make derivative art of it, have other learn from it, but if ever a machine 'learn' from it, it's "stealing". It's not because you don't like a derivative work, that the derivative work is "stealing" your content. Saying it's stealing is wrong, it's lying to get your point accross. You blame "how…

It's a lot of mental gymnastic to think machine learning = human learning. Especially on HN, where people should understand scale matters a lot in real world.

It's generally consider ok to sell your fanart on comic festivals. But do you think it's okay that Disney starts selling fanart of One Piece without the publisher's permission?

A car and a pair of legs both move you from A point to B point. So why do laws treat automobiles and pedestrians so differently?

Re: Llama and ChatGPT Are Not Open-Source

#82

Earlier quoted context omitted.

They’re straw-manning, programmers are the best sharers in the world. Open source software has lead the drive for open source learning and information in general.

https://news.ycombinator.com/item?id=36900844 Right here the comment says they prefers Meta to keep the secret sauce as long as it allows they to (inderectly) access the copyrighted material etc. And the replies below are generally positive. So it's not a straw-man, at least in this thread, at all.

It should be noted, that they're explicitly taking this position because they think it will result in more capabilities. Not because of a respect of intellectual property.

Re: Llama and ChatGPT Are Not Open-Source

#83
post #64

Earlier quoted context omitted.

How are Modern Copyright laws helping authors, artists, and actors in those cases?

It's allowing them to sue OpenAI for copyright infringement: https://www.theguardian.com/books/2023/jul/05/authors-file-a...

Allowing them to block human intellectual progress may not be the long-term win you assume it is.

Re: Llama and ChatGPT Are Not Open-Source

#84
post #50

Earlier quoted context omitted.

Modern copyright is not a good, it's an evil.

That depends on the jurisdiction, but in my view I err on the side of "but it's necessary." A society without copyright would be a much poorer one, and humans figured that out a long time ago.

A society without copyright would be a much poorer one

That was before "Attention Is All You Need" and the subsequent work on LLMs that grew out of it. Things are different now.

Re: Llama and ChatGPT Are Not Open-Source

#85
post #76
post #59

Earlier quoted context omitted.

You're not addressing the massive abuse of the commons this still represents. If artists don't have the right to tell you to fuck off for using their work in training data, they're less likely to publicly show that work, which hurts them because they become less visible and hurts the AI because the training data gets worse.

Go on youtube and type "copy arstyle". Now tell me how artists were not stealing from each other.

Artists are generally pretty encouraging to people entering the field and using their stuff as reference for new artists. That actually contributes to art. It's definitely not the same thing as a massive tech corporation trying to automate their livelihoods, but please keep making this flawed argument analogizing two completely different processes.

Re: Llama and ChatGPT Are Not Open-Source

#86

"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…

Does anyone have an idea of what (if anything) is being implied by the last two sentences?

Bad company scary, I guess. Normally I'm sympathetic to those types, but given how Llama works without a network connection I'm not sure what I have to fear. The "history of this company's choices" has generally been pretty good with AI, so it sounds like speculation on their part.

Re: Llama and ChatGPT Are Not Open-Source

#87
post #12

I wrote a small piece sharing my views on the topic https://medium.com/@anandandbeyond/rethinking-open-source-fo...

And posted the link to your subscribers only medium page twice in this thread

I wanted to delete the other link, but hackernews doesn't allow me to do so. Also, I need to figure out why the link is subscriber-only. It was not intentional.

Re: Llama and ChatGPT Are Not Open-Source

#88
post #65

Yes, we know. Now can we please stop treating them as developer-friendly tools and more as a hostile theft of intellectual property?

It's shocking to me how many people in tech feel completely entitled to intellectual property that took someone years to master a skill to make. But talk about releasing a proprietary codebase and suddenly they want the lawyers involved because that actually threatens their livelihood.

This is exactly right. Almost everything you see on the internet is copyrighted, even if you're allowed to view it for free. Even FOSS and CC-licensed content is copyrighted. It's just as copyrighted a proprietary codebase. But for some reason, you don't see coding LLMs trained on proprietary code, eg. you don't see Microsoft training Copilot on private GitHub repos. It all smells very hypocritical, like these companies feel entitled to harvest free labour from the internet to make their proprietary AI models, but they're not willing to do the same with their own.

Re: Llama and ChatGPT Are Not Open-Source

#90

"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…

Does anyone have an idea of what (if anything) is being implied by the last two sentences?

> He is an associate professor in Language and Communication at the Centre for Language Studies of Radboud University Nijmegen. Dingemanse obtained a MA degree in African Languages and Cultures at Leiden University in 2006, and a PhD degree in arts in 2011 at Radboud University Nijmegen

https://en.wikipedia.org/wiki/Mark_Dingemanse

This co-author? Given the relevance of their academic background, I'd take their opinion on "the dangers of using LLMs released by the untrustworthy Meta" with a grain of salt.

Post reply on HN