Earlier quoted context omitted.
Does anyone have an idea of what (if anything) is being implied by the last two sentences?
Bad company scary, I guess. Normally I'm sympathetic to those types, but given how Llama works without a network connection I'm not sure what I have to fear. The "history of this company's choices" has generally been pretty good with AI, so it sounds like speculation on their part.
Llama and ChatGPT Are Not Open-Source
91–100 of 130 posts
Re: Llama and ChatGPT Are Not Open-Source
#92Earlier quoted context omitted.
Does anyone have an idea of what (if anything) is being implied by the last two sentences?
> He is an associate professor in Language and Communication at the Centre for Language Studies of Radboud University Nijmegen. Dingemanse obtained a MA degree in African Languages and Cultures at Leiden University in 2006, and a PhD degree in arts in 2011 at Radboud University Nijmegen https://en.wikipedia.org/wiki/Mark_Dingemanse This co-author? Given the relevance of their academic background, I'd take their opini…
Re: Llama and ChatGPT Are Not Open-Source
#93Earlier quoted context omitted.
> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…
>Could we stop celebrating how tech is going to steal everyone's copyrighted works Copyright has a fair-use exception for transformative works. It is difficult to look at the LLM's of the day and think "No, they have not taken the copywritten works and transformed them into something completely new." There is no hate of artists. I don't know how you get from "It's not copyright violation" to "I hate artists". This is…
I think an argument could be made that it's fine for an LLM to learn from copyrighted works, but maybe it should have to go to the library to do it. Having your own human-accessible copy of those works (at home, or at Meta), doesn't sound as acceptable to me.
Re: Llama and ChatGPT Are Not Open-Source
#94Earlier quoted context omitted.
When you dig into it, there are very few truths. Who "won" the last US election? Where did covid come from? Humans will argue the right answer until their last days. It's frustrating how on-the-fence chatgpt can be. It's pretty interesting too, because in a professional environment one of the most important things you need to do is have an opinion and take a position otherwise you can't execute.
That's not true though. Every word of this sentence, for instance, has a correct spelling and grammatical rules as to where the words and punctuation go. To English learners, that might be very difficult. I'm learning a new language now and I feel the pain. The textbook and instructor, in this case, is way more correct than the collective opinion of my fellow students. The vast majority of things actually follow this…
As a linguistic descriptivist, hahahahahahahaha.
Re: Llama and ChatGPT Are Not Open-Source
#95Earlier quoted context omitted.
>Could we stop celebrating how tech is going to steal everyone's copyrighted works Copyright has a fair-use exception for transformative works. It is difficult to look at the LLM's of the day and think "No, they have not taken the copywritten works and transformed them into something completely new." There is no hate of artists. I don't know how you get from "It's not copyright violation" to "I hate artists". This is…
So I can have a giant database of all the most recent copyrighted works, on my home computer, as long as I claim I'm using it for training a model? And if I happen to listen to some music or watch some movies too, who's going to know? I think an argument could be made that it's fine for an LLM to learn from copyrighted works, but maybe it should have to go to the library to do it. Having your own human-accessible cop…
Re: Llama and ChatGPT Are Not Open-Source
#96Earlier quoted context omitted.
> He is an associate professor in Language and Communication at the Centre for Language Studies of Radboud University Nijmegen. Dingemanse obtained a MA degree in African Languages and Cultures at Leiden University in 2006, and a PhD degree in arts in 2011 at Radboud University Nijmegen https://en.wikipedia.org/wiki/Mark_Dingemanse This co-author? Given the relevance of their academic background, I'd take their opini…
Since their academic background is so relevant (language studies and cultures being the exact domain of the impact of LLMs on society), I don't see why a grain of salt is needed.
If their academic background were relevant, they would have provided specific negative outcomes that may result from the usage of Llama 2, rather than a vague "the history of this company's choices does not inspire confidence".
What, of the many open source software releases by that company in the ML/AI field "does not inspire confidence"?
Re: Llama and ChatGPT Are Not Open-Source
#97Yes, we know. Now can we please stop treating them as developer-friendly tools and more as a hostile theft of intellectual property?
The entire subculture that this site takes its name from is hostile to IP.
I mean, except for the start-up founders that hang around here, but this site is mostly kept up as a way for them to recruit it seems like (and maybe a way to build good will).
Edit: To be clear, I'm not necessarily saying you're wrong, but you could probably pick a more sympathetic argument from a strategic perspective. What you've done is kinda the equivalent to going to the Vatican and arguing that something is a threat to idol worship. You might be right, but a good chunk of your audience wants that.
Re: Llama and ChatGPT Are Not Open-Source
#98Earlier quoted context omitted.
I believe LLMs should be allowed to read/view/consume content and learn from it even if that content has a copyright. We phrase it like somehow the material is being copied into the LLM, but that’s not what it’s doing. It’s building a neural graph from the experience of consuming that content. What would the world be like if humans couldn’t learn, train the weights of the interconnects of their neural tissue, from an…
It’s a form of lossy compression. Can I strip the copyright off an image by JPEG compressing it? At the very least I think LLMs trained on data that the trainer does not own or have rights to use in that manner should not be copyrightable.
My thinking “the enemy gate is down” when considering the tokens “Ender’s Game” is my recalling a learned association of those tokens to the given token string.
My knowing that doesn’t strip the copyright. My telling someone the meaning and context of the phrase generally doesn’t strip the copyright away from Orson Scott Card. I’m not reproducing his work but my knowledge of it. And it’s dependent on what I do with that knowledge and how if I’ve violated his copyright.
We are prosecuting the LLMs for possessing fragments of knowledge. And we’re assuming that the recall of some of those fragments means a copy of that work is in fact contained within the weights.
Re: Llama and ChatGPT Are Not Open-Source
#99"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…
Title: Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
Authors: Andreas Liesenfeld, Alianda Lopez, Mark Dingemanse
Abstract: Large language models that exhibit instruction-following behaviour represent one of the biggest recent upheavals in conversational interfaces, a trend in large part fuelled by the release of OpenAI's ChatGPT, a proprietary large language model for text generation fine-tuned through reinforcement learning from human feedback (LLM+RLHF). We review the risks of relying on proprietary software and survey the first crop of open-source projects of comparable architecture and functionality. The main contribution of this paper is to show that openness is differentiated, and to offer scientific documentation of degrees of openness in this fast-moving field. We evaluate projects in terms of openness of code, training data, model weights, RLHF data, licensing, scientific documentation, and access methods. We find that while there is a fast-growing list of projects billing themselves as 'open source', many inherit undocumented data of dubious legality, few share the all-important instruction-tuning (a key site where human annotation labour is involved), and careful scientific documentation is exceedingly rare. Degrees of openness are relevant to fairness and accountability at all points, from data collection and curation to model architecture, and from training and fine-tuning to release and deployment.
Re: Llama and ChatGPT Are Not Open-Source
#100"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…
"For basic visitor statistics I use Matomo, an excellent open source alternative for Google Analytics. IP addresses are anonymized and data never leaves the server."