Live data from Hacker News

Microsoft says that it's okay to steal web content because it's 'freeware.'

windowscentral.com

11–20 of 29 posts

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#11

This doesn't deprive the original owner, so they should use "share" or "pirate" instead.

Microsoft is head-quartered in Seattle, which is part of the United States, so it's not unreasonable for harmed parties to bring action under US Copyright law. US Copyright law is explicit about this. Copyright exists upon creation of the work unless specifically disclaimed.

You can reasonably assume that by posting content to the intarwebs, the author provided an implicit license to view the content (for the community that could reasonably be identified as being able to view the work.) But this does not imply an automatic "freeware" or "fair use" exemption.

How can Suleyman say someone who published something on the web in 1995 intended for it to be used to train their LLM in 2020? It's unreasonable to assume that.

There's a reason you see people put explicit licenses on things, it's because they don't exist unless the work's creator releases the work under a specific license. And license for a human to read something on the web is not a license to give your content to a large company for them to remix it into... whatever the heck it is they're remixing it into.

But yeah. It's probably worth it to brush up on robots.txt definitions to exclude spiders from scraping your site if they're going to use it as LLM fodder.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#12
post #8

It's ironic that Microsoft used copyright protection and IP law for years to secure a dominant market position, and now they don't need to play by the same rules because "something something AI".

Truthfully, I agree with this new stance of theirs and believe it's always been the case, but there's no doubt that they're only adopting it now because it's become advantageous.

Sure. But there's a fair amount of US Copyright law that says this has never been the case. To test this hypothesis, upload a video to YouTube with music by Prince playing in the background.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#13

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

Fair use is not granted by a social contract, as he hopes. It's granted by courts. Suleyman's interpretation is the most self-serving possible opinion of the circumstances.

As Joshua Topolski put it, "Only someone utterly divorced from creation and completely insulated in a bubble of entitlement would have such a distorted view of other people's work like this" [0]

[0] https://x.com/joshuatopolsky/status/1806796270402699553?s=19

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#14
Someone's reasons for sharing information are coloured by the situation at the time of sharing it, amongst many other factors.

Two years ago (say) no one predicted the meteoric rise of LLMs and their voracious appetite for data sets for training. These beasties are not simply search engines that are better direction pointers to your stuff (with a frisson of ads) but insist on being the final word and keep you out. To be blunt: It is stealing.

The implied contract for publishing on the web has changed again, just as it has several times in the past. The worst thing here is the use of the term "freeware". Describing original content, displayed for all to see as -ware is outrageous.

They might as well describe the content on Spotify and co as freeware ... bear with me: you could scrape wifi connections through your publicly available APs or even do some more broadband funky spectrum capture analysis and claim that is what an internet search engine does in its spare time and all is fine (lol).

LLMs and GenAI are quite interesting things but I do not think that they are the last word in ... AI. Anyway the latest cool thingie cannot be allowed to break whatever the current unspoken and somewhat undefined social contract is in place.

This bloke from MS seems to have forgotten that there really is a social contract of some sort and that if you say: "fuck you lot, omnomnom ... mmmm data ... ... laters (lol)" there might be some come back.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#15

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

I wish someone would define what scraping "open web" means since physical media has been dead for a while. YouTube content? Doxxing people? Archiving books and old video games? Bank transactions? VoIP calls? Why is it potentially ok for AI to copy these things, but if a human does it they get into trouble? We need an exhaustive list and that's probably impossible to define.

The courts ruled that content can be scraped from the open web legally. From your examples 1. YouTube content: ok because the scraped content becomes derivative works which falls under free use 2. Doxxing people: if the data was already on the web and a human found it then the person was already doxed 3. Bank transactions: are they behind a login wall? Not Ok to scrap 4. VoIP calls: again behind a wall? Not Ok to scrape

If a human can find it online without doing anything beyond accessing a public web link then it is Ok to scrape it. It's not the complicated and the law has a series of frameworks that for existing cases. It's not particularly complicated.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#16

Earlier quoted context omitted.

I wish someone would define what scraping "open web" means since physical media has been dead for a while. YouTube content? Doxxing people? Archiving books and old video games? Bank transactions? VoIP calls? Why is it potentially ok for AI to copy these things, but if a human does it they get into trouble? We need an exhaustive list and that's probably impossible to define.

The courts ruled that content can be scraped from the open web legally. From your examples 1. YouTube content: ok because the scraped content becomes derivative works which falls under free use 2. Doxxing people: if the data was already on the web and a human found it then the person was already doxed 3. Bank transactions: are they behind a login wall? Not Ok to scrap 4. VoIP calls: again behind a wall? Not Ok to scr…

Derivative works are not automatically fair use. Otherwise I could release an album with a Prince song and a Taylor Swift song, claim it's a derivative work and covered under fair use. Again. Try to do this on YouTube (even though YouTube is not the arbiter of what is and isn't fair use, they're just aggressive about removing content.)

Doxxing people can give cause to restraint if it's inciting speech. Consider a situation where you say "Hey. I just found Joe Schmo's Address. He lives at 123 Main St., Anytown, DK 10010." Now compare it to when you say "Hey. I just found Joe Schmo's Address. He lives at 123 Main St., Anytown, DK 10010. Joe's a child molester. Let's go burn down his house." or if you call up the local police department and said "Hey. I saw Joe Schmo with Goody Parker in their house at 123 Main St., Anytown, DK 10010. They're molesting a child right now! Hurry, if you're fast you might stop them!"

Yeah... I don't remember saying bank transactions were copyrighted or in the public domain. Probably someone elses' comment.

VoIP calls? uh... you may want to review the Electronic Communications Privacy Act of 1986, 18USC2510-2522 before assuming VoIP calls can be intercepted for the purpose of training an AI.

You are free to believe anything you can access is fair use. (I mean, maybe in Iceland or Sweden.) But give me a call if you get caught doing this, I can refer you to a decent attorney. (Again... this all assumes you're in the states. International copyright is sort of a wild west. But you would expect MSFT to open up a fully owned offshore subsidiary in a country like Vanuatu or Iceland to evade continental copyright (what is it with islands and copyright???)

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#17

This doesn't deprive the original owner, so they should use "share" or "pirate" instead.

Microsoft is head-quartered in Seattle, which is part of the United States, so it's not unreasonable for harmed parties to bring action under US Copyright law. US Copyright law is explicit about this. Copyright exists upon creation of the work unless specifically disclaimed. You can reasonably assume that by posting content to the intarwebs, the author provided an implicit license to view the content (for the communi…

Unfortunately, most web scrapers ignore robots.txt.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#18
Of course it's okay.

I make an http _REQUEST_, the server voluntarily fulfills the request.

Why is it okay for a person to view your content, memorize it, and use it as a base for new content while it's not okay for an AI? at the end of the day it is the same thing.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#19

Earlier quoted context omitted.

I wish someone would define what scraping "open web" means since physical media has been dead for a while. YouTube content? Doxxing people? Archiving books and old video games? Bank transactions? VoIP calls? Why is it potentially ok for AI to copy these things, but if a human does it they get into trouble? We need an exhaustive list and that's probably impossible to define.

The courts ruled that content can be scraped from the open web legally. From your examples 1. YouTube content: ok because the scraped content becomes derivative works which falls under free use 2. Doxxing people: if the data was already on the web and a human found it then the person was already doxed 3. Bank transactions: are they behind a login wall? Not Ok to scrap 4. VoIP calls: again behind a wall? Not Ok to scr…

It’s very complicated, and if a lawyer tells you it’s not you need to get a new lawyer. Whether something is fair use is determined based on a number of different factors that all are weighed. There’s not some simple if-then algorithm to tell you the answer. There is case law which sets up some known points in the space, but if you walk off that manifold, no one knows how things will be weighed.

LLMs are off the manifold in a lot of ways. We’ve never had a form of information digestion/encoding quite like them.

Politically, both content creators and AI companies see this decision as existential. Ruling one way or the other is actually almost unthinkable. The only solution is going to be some compromise that comes from Congress, because the sorts of things that would be in such a compromise can’t be imposed by the courts with existing legislation.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#20
post #18

Of course it's okay. I make an http _REQUEST_, the server voluntarily fulfills the request. Why is it okay for a person to view your content, memorize it, and use it as a base for new content while it's not okay for an AI? at the end of the day it is the same thing.

Because it's never been okay to reproduce something as the base of "new" content that competes with the original work. It's pretty much why there is copyright. I merely remembered your painting with my chemical film memory machine and then sold prints of my memories.
Post reply on HN