Earlier quoted context omitted.
> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…
I wonder if there's a future title like "AI Model Artist", "AI Model Contributor", "AI Model Author", "AI Model Creator" or something for creators who contribute to the models in use. While I love we can all immediately benefit from previous works, I feel for the countless artists and creators who's work has been integrated into models with zero compensation. There are already outstanding lawsuits seeking reparations…
Llama and ChatGPT Are Not Open-Source
61–70 of 130 posts
Re: Llama and ChatGPT Are Not Open-Source
#62Earlier quoted context omitted.
Right but I feel like in your argument there is an implicit assumption that the only way to get ChatGPT etc. to understand say a global climate system is by omitting false information entirely. Why can't an LLM extract a signal from the noise? It seems as though it does that already. If you're worried that there is more text on the internet saying that the Earth is flat than there is text saying that the Earth is rou…
> Right but I feel like in your argument there is an implicit assumption that the only way to get ChatGPT etc. to understand say a global climate system is by omitting false information entirely. Why can't an LLM extract a signal from the noise? It seems as though it does that already. Because it's not thinking , it's looking at probabilities. If you don't prune branches, you've created a disinformation radiator if t…
Some scams, like MLM, encourage massive volumes of this bullshit to be created and disseminated.
The democracy of information does not lead to an increased accuracy based on whatever information spreads the most. The internet is a multi-decade long global proof of this. There's a reason that classrooms have 1 teacher and say 24 students - only 4% of that room is to be trusted with the material.
To know if something is true, you need to eventually have some kind of external confirmation. Karl Popper wrote a lot about this. These chat systems can make predictions about the existing corpus but it can't like roll around town and do actual science to confirm things.
If all truth must be derived from self-contained systems than you have to make sure your self-contained system isn't bullshit.
A game for instance, can have its own rules and consistencies and things can be rationalized through the game. But how much that maps to a reality cannot be concluded strictly from inside the game.
For instance, if you want expertise in masonry and construction you need to make sure you aren't just feeding your AI Tetris and SimCity.
Re: Llama and ChatGPT Are Not Open-Source
#63Earlier quoted context omitted.
Right but I feel like in your argument there is an implicit assumption that the only way to get ChatGPT etc. to understand say a global climate system is by omitting false information entirely. Why can't an LLM extract a signal from the noise? It seems as though it does that already. If you're worried that there is more text on the internet saying that the Earth is flat than there is text saying that the Earth is rou…
> Right but I feel like in your argument there is an implicit assumption that the only way to get ChatGPT etc. to understand say a global climate system is by omitting false information entirely. Why can't an LLM extract a signal from the noise? It seems as though it does that already. Because it's not thinking , it's looking at probabilities. If you don't prune branches, you've created a disinformation radiator if t…
All that being said, who is to say that humans aren't just looking at probabilities when they think?
Re: Llama and ChatGPT Are Not Open-Source
#64Earlier quoted context omitted.
Modern copyright is not a good, it's an evil.
Authors, artists, actors etc whose work have been ripped off by mega corporations like Meta, OpenAI, Google etc would disagree. Copyright may need to be updated to cater for the new world of AI but that doesn't mean as a concept it is evil.
Re: Llama and ChatGPT Are Not Open-Source
#65Yes, we know. Now can we please stop treating them as developer-friendly tools and more as a hostile theft of intellectual property?
Re: Llama and ChatGPT Are Not Open-Source
#66Earlier quoted context omitted.
Pretty much everything nowadays is copyrighted, by omitting such materials, what are you really left with? LLM is a tool much like the internet is a tool. Yes, someone can use it to steal, but stealing is against the law. Instead of encoding a criminal justice system into an LLM by omitting the possibility of stealing an artists work or omitting the knowledge of physics so someone can't learn how to build a bomb, we…
You're not addressing the massive abuse of the commons this still represents. If artists don't have the right to tell you to fuck off for using their work in training data, they're less likely to publicly show that work, which hurts them because they become less visible and hurts the AI because the training data gets worse.
Even Open Source advocates don't really have the right to stop companies from using Open code. The license discourages it, but everyone from Tesla to Nintendo has been caught violating it's terms. Publishing stuff on the open web has always had consequences, unfortunately.
Re: Llama and ChatGPT Are Not Open-Source
#67I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…
> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…
Copyright has a fair-use exception for transformative works. It is difficult to look at the LLM's of the day and think "No, they have not taken the copywritten works and transformed them into something completely new."
There is no hate of artists. I don't know how you get from "It's not copyright violation" to "I hate artists". This is a matter of existing law. It is of course unsettled, the fair-use interpretation not yet tested in courts. But the same goes for affirming it's a violation of copyright.
These questions will have their day in court, and until then there is no need to make arguments in an inflammatory tone that that borders on personal attack. In the meantime, maybe engage in a conversation about how copyright law would need to be changed to account for this new technology, not condemn people who see a reasonable interpretation of existing law the differs from your own.
Re: Llama and ChatGPT Are Not Open-Source
#68Yes, we know. Now can we please stop treating them as developer-friendly tools and more as a hostile theft of intellectual property?
It's shocking to me how many people in tech feel completely entitled to intellectual property that took someone years to master a skill to make. But talk about releasing a proprietary codebase and suddenly they want the lawyers involved because that actually threatens their livelihood.
Re: Llama and ChatGPT Are Not Open-Source
#69I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…
> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…
We phrase it like somehow the material is being copied into the LLM, but that’s not what it’s doing. It’s building a neural graph from the experience of consuming that content.
What would the world be like if humans couldn’t learn, train the weights of the interconnects of their neural tissue, from any material with a copyright?
Re: Llama and ChatGPT Are Not Open-Source
#70Earlier quoted context omitted.
> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…
I believe LLMs should be allowed to read/view/consume content and learn from it even if that content has a copyright. We phrase it like somehow the material is being copied into the LLM, but that’s not what it’s doing. It’s building a neural graph from the experience of consuming that content. What would the world be like if humans couldn’t learn, train the weights of the interconnects of their neural tissue, from an…
At the very least I think LLMs trained on data that the trainer does not own or have rights to use in that manner should not be copyrightable.