Earlier quoted context omitted.
This is not about memory or training. The LLM training process is not being run on books streamed directly off the internet or from real-time footage of a book. What these companies are doing is: 1. Obtain a free copy of a work in some way. 2. Store this copy in a format that's amenable to training. 3. Train their models on the stored copy, months or years after step 1 happened. The illegal part happens in steps 1 an…
But making copies for yourself, without distributing them, is different than making copies for others. Google is downloading copyrighted content from everywhere online, but they don't redistribute their scraped content. Even web browsing implies making copies of copyrighted pages, we can't tell the copyright status of a page without loading it, at which point a copy has been made in memory.
OpenAI asks White House for relief from state AI rules
201–210 of 845 posts
Re: OpenAI asks White House for relief from state AI rules
#202Earlier quoted context omitted.
Can I download a book without paying for it, and print copies of it? Stash copies in my bathroom, the gym, my office, my bedroom etc. to basically have a copy on hand to study from whenever I have some free time? What about movies and music?
owning a copy and learning the information is not the same. you can learn 2+2=4 from a book, but you no longer need that book to get that answer. each year in school, I was issued a book for class, learned from it, returned the book. I did not return the learning. musicians can read the sheet music and memorize how to play it, and no longer need the music. they still have the information.
There's two angles to the lawsuits that are getting confused - the largest one from the book publishers (Sarah Silverman et al) attacked from the angle that the models could reproduce copyrighted information. This was pretty easily quelled / RHLF'd out (used to be that if ChatGPT started producing lyrics a supervisor/censor would just cut off it's response early - tried it now and ChatGPT.com is now more eloquent, "Sorry, I can't provide the full lyrics to "Strawberry Fields Forever" as they are copyrighted. However, I can summarize the song or discuss its themes, meaning, and history if you're interested!")
But there's also the angle of "why does OpenAI have Sarah Silverman's book on their hard drive if they never paid her for it? This is the lawsuit against Meta regarding books3 and torrenting, seems like they're getting away with the "we never redistributed/seeded!" but it's unclear to me why this is a defense against copyright infringement.
Re: OpenAI asks White House for relief from state AI rules
#203Earlier quoted context omitted.
To the best of my knowledge, no individual has ever been sued or prosecuted specifically for downloading books. As long as you're not massively sharing them with others, it's not an issue in practice. Enjoy your reading and learning.
Aaron Swartz, cofounder of Reddit and inventor of RSS and Markdown, was hounded to death by an overzealous prosecutor for downloading articles from JSTOR, with the intent to learn from them. He was charged with over a million dollars in fines and could have faced 35 years in prison. He and Sam Altman were in the same YC class. OpenAI is doing the same thing at a larger scale, and their technology actually reproduces…
> It's shameful that they are making claims that they aren't infringing creator's rights when they have scraped the entire internet.
Scraping the Internet is generally very different from piracy. You are given a limited right to that data when you access it, and you can make local copies. if further use does something sufficiently non-copying, then creator rights aren't being infringed.
Re: OpenAI asks White House for relief from state AI rules
#204Re: OpenAI asks White House for relief from state AI rules
#205Earlier quoted context omitted.
CDs, software, and electronic media, yes. Physical books, no. You can't make archival copies.
sure you can, you could take a physical book, and painstakingly copy each page at a time, that is totally fair use.
You cannot legally photocopy copy an entire book even if you own a physical copy.
Internet people say you can, but there's no actual legal argument or case law to support that.
Re: OpenAI asks White House for relief from state AI rules
#206Earlier quoted context omitted.
> ask for regulation then ask for exempt That's exactly what has been happening: Ask HN: Why is OpenAI pushing for regulation so much - 2023 https://news.ycombinator.com/item?id=36045397
OpenAI lobbied for restrictive rules, and now they want an "out" but only for themselves. Absolute naked regulatory capture.
Re: OpenAI asks White House for relief from state AI rules
#207Earlier quoted context omitted.
Citation needed.
It seems reasonably within the bounds described by fair use, but nobody's ever tested that particular constellation of factors in a lawsuit, so there's no precedent - hand copying a book, that is. 17 U.S.C. § 107 is the fair use carveout. Interestingly, digitizing and copying a book on your own, for your own private use, has also not been brought to court. Major rights holders seem to not want this particular fair us…
What part of fair use pertains to making a physical copy of the complete work?
Re: OpenAI asks White House for relief from state AI rules
#208If they want to avoid paying for the creative effort of authors and other artists then they should also not charge for the use of their models.
Re: OpenAI asks White House for relief from state AI rules
#209I heard the theory that Elon Musk has a significant control over the current US government. They're not best pals with Sam Altman. This seems like it might be a good way to see how much power Elon actually has over the government?
People might think this is a partisan statement, but it's not. It's simply how he is operating. Want power? Want to get things done? Kiss his feet. You saw all the tech boys line up at his inauguration. You saw him tell Zelenskyy "Thank me". Elon might have power, but he is also on a leash.
Re: OpenAI asks White House for relief from state AI rules
#210Earlier quoted context omitted.
https://en.wikipedia.org/wiki/Aaron_Swartz (Please no pedantry about how scientific papers aren't books)
Aaron Swartz downloaded a lot of stuff. Did he publish the stuff too? That would be an infringement. But only downloading the stuff? And never distributing it? Not sure if it’s worth a violation .
A tiny fraction compared to the 80+ terabytes Facebook downloaded.
>Did he publish the stuff too?
No.
> Not sure if it’s worth a violation .
Exactly.