It would be interesting to have a "classical AI model", trained on the contents of the Harvard libraries before 1926 and now out of copyright.
X changes its terms to bar training of AI models using its content
51–60 of 217 posts
Re: X changes its terms to bar training of AI models using its content
#52Or is there stuff in the user agreement that separately prohibits this?
Obviously barring normal copyright law which is still up in the air.
Re: X changes its terms to bar training of AI models using its content
#53I've never signed up for the X developer program, so I'm not bound by these terms. But I did download an archive of my data last week. Do I have implicit permission to use that data (~150k liked tweets) to train AI models? Or is there stuff in the user agreement that separately prohibits this? Obviously barring normal copyright law which is still up in the air.
Re: X changes its terms to bar training of AI models using its content
#54I'm not sure how this will work as crawlers don't read or accept ToS.
Without search engines what the point in posting it on open net if nobody can find.
Re: X changes its terms to bar training of AI models using its content
#55I've never signed up for the X developer program, so I'm not bound by these terms. But I did download an archive of my data last week. Do I have implicit permission to use that data (~150k liked tweets) to train AI models? Or is there stuff in the user agreement that separately prohibits this? Obviously barring normal copyright law which is still up in the air.
If you live in the EU, GDPR dictates that you own your data generally speaking. If you're in the US it varies by state if you have any rights at all.
Re: X changes its terms to bar training of AI models using its content
#56Earlier quoted context omitted.
Why should it be a default? Can you prove that training a model on data you wrote is not fair use? We're already seeing precedent that it might be. https://www.ecjlaw.com/ecj-blog/kadrey-v-meta-the-first-majo... The openness of the internet is a good thing, but it doesn't come without a cost. And the moment we have to pay that cost, we don't get to suddenly go, "well, openness turned out to be a mistake, let's close…
The burden is on the user to show that it is fair use, no? Not everyone else's responsibility to prove that it's _not_ fair use.
Accordingly, anyone on the internet who wants to make comments about how they should be able to prevent others from training models on their data needs to demonstrate competence with respect to copyright by explaining why it's not fair use, as currently it is undecided in law and not something we can just take for granted.
Otherwise, such commenters should probably just let the courts work this one out or campaign for a different set of protection laws, as copyright may not be sufficient for the kind of control they are asking over random developers or organizations who want to train a statistical model on public data.
Re: X changes its terms to bar training of AI models using its content
#57Earlier quoted context omitted.
you can also just ignore the berne convention, and accept whatever consequences there might be
this would void the copyrights of your citizens and companies essentially forever
Re: X changes its terms to bar training of AI models using its content
#58Earlier quoted context omitted.
The burden is on the user to show that it is fair use, no? Not everyone else's responsibility to prove that it's _not_ fair use.
It is definitely the responsibility of anyone suing someone who trained a model on copyrighted data to prove that it isn't fair use, they have to show how it violated law, and while it's in the best interest of those organizations to make things easier for the court by showing why it is fair use, they are technically innocent until proven guilty. Accordingly, anyone on the internet who wants to make comments about ho…
Re: X changes its terms to bar training of AI models using its content
#59If an artist or author can't do this, social media shouldn't be able to do it either. If Xai wants to train on public corpus, it shouldn't be allowed to prevent its own corpus from being used. We need regulations to limit the power grabs. Train all you like, but don't dare try to constrain to your walled gardens. We should also probably nip the "foundation model company / also a social media company" conglomeration i…
Social media should do it to set a legal precedent. > We need regulations to limit the power grabs. Train all you like, but don't dare try to constrain to your walled gardens. No, no one should train, period.
I get that you have your own opinion, but I'm personally tired of living in the butter-churning era and would prefer that this all went a bit faster.
I want my real time super high fidelity holo sim, all of my chores to be automatically done, protein folding, drug discovery. The life extension, P = NP future. No more incrementalism.
If the universe only happens once, and we're only awake for a geological blink of an eye, I'd rather we have an exciting time than just be some paper-pushing animals that pay taxes and vanish in a blip.
I'd be really excited if we found intelligent aliens, had advanced cloning for organ transplants and longevity, developed a colony on Mars, and invented our robotic successor species. Xbox and whatever most normal people look forward to on a day to day basis are boring.
Re: X changes its terms to bar training of AI models using its content
#60Earlier quoted context omitted.
It does surprise me that we haven't seen nations revise their copyright window back to something sensible in a play to seed their own nascent AI industry. The American founding fathers thought 20 years was enough. I'm sure there'd be repercussions in the banking system, but at some point it might be worth the trade.
they can't a 50 year minimum is part of the berne convention, which itself is as close to a universal law as humanity has (even North Korea is a signatory)