Live data from Hacker News

Microsoft says that it's okay to steal web content because it's 'freeware.'

windowscentral.com

21–29 of 29 posts

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#21
post #18

Of course it's okay. I make an http _REQUEST_, the server voluntarily fulfills the request. Why is it okay for a person to view your content, memorize it, and use it as a base for new content while it's not okay for an AI? at the end of the day it is the same thing.

> Why is it okay for a person to view your content, memorize it, and use it as a base for new content while it's not okay for an AI? at the end of the day it is the same thing.

Because it is not really AI. We can have this comparison when it actually can think. Until then this 'training' is just big data processing and that is not what person can do by reading/watching.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#22

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

What nonsense. You can download it, view it, create a search engine index with it. You cannot "recreate with it", "reproduce with it".

Interesting that Microsoft openly rejects the DMCA. So we now can all reverse engineer and make available for free all Windows versions, since they have been on the web at some point. I guess Suleyman's statement can be used in court against Microsoft in the next IP trial.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#24
post #18

Of course it's okay. I make an http _REQUEST_, the server voluntarily fulfills the request. Why is it okay for a person to view your content, memorize it, and use it as a base for new content while it's not okay for an AI? at the end of the day it is the same thing.

Is this a serious question?

It's never been ok for a person to do that.

You don't even have to make a request to hear a radio broadcast or see a tv broadcast, they are actively broadcast and bathe you whether you wanted it or not.

Now imagine setting up a service, even a free one where you don't make any commercial gain from it, where people ask you for stuff, and you give them bits of the stuff you've collected that was just literally falling from the sky. You give them the stuff not in the form of quotes for discussion, but presented as content itself, devoid of any prior context.

The user asks for some sci fi, and you give them 15 minutes of Star Wars which you got from a tv broadcast. But you don't tell the user "This is Star Wars, written by George Lucas" You just give them the content and the user thinks you created it.

Or worse, the user thinks THEY created it, merely using your tool, just a fancier version of using a spell checker, they are still the author since they directed the tool.

The user asks for a deep insight into human nature, you give them something that some famous author said, but not the authors name. The user then publishes a book containing this gem of deep insight, and other users think they are deeply insightful and buy their book.

This is not what happens when a person learns from life, including reading famous books, and then starts producing their own content.

The fact that it's possible to construct some scenarios where it seems like all the essential elements are the same does not actually make the two things equivalent.

Here's my own little insight into human nature I'll offer: It's pretty common for some people who pride themselves on being smart, for some mysterious reason to only be selectively smart. They will construct good valid logic that serves one purpose readily, and yet somehow fail to construct or even percieve equally valid logic that serves some other purpose.

If one is smart enough to try to present this logical argument that it's ok to steal other peiple's work, I submit that they are then also smart enough to be able to work out how and why that is in fact stealing.

We don't even have to get into the obvious, how MS themselves "broadcast" Windows to everyone by getting manufacturers to preinstall it on every machine. That wasn't even an http request. You just wanted to buy a piece of hardware, and it just came with Windows whether you wanted it or not. Even if you remove it, it was already given to you before you opened the box, and even after you wipe it, it's licence key is still there literally burned into the hardware. Therefor, since MS put it out there themselves, put it into your hands themselves, you are free to reproduce Windows. Oh, just pieces? Pieces are fair use? So, just a handy dll then. That must be ok.

There is no end to the ways this argument doesn't hold water.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#25
post #8

Earlier quoted context omitted.

Truthfully, I agree with this new stance of theirs and believe it's always been the case, but there's no doubt that they're only adopting it now because it's become advantageous.

Sure. But there's a fair amount of US Copyright law that says this has never been the case. To test this hypothesis, upload a video to YouTube with music by Prince playing in the background.

Copyright law is more of an ideal, and if it had "never been the case", there wouldn't have been any content to DMCA off of YouTube in the first place. US copyright law is fighting the tide, not conveying a natural order.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#26

Earlier quoted context omitted.

I wish someone would define what scraping "open web" means since physical media has been dead for a while. YouTube content? Doxxing people? Archiving books and old video games? Bank transactions? VoIP calls? Why is it potentially ok for AI to copy these things, but if a human does it they get into trouble? We need an exhaustive list and that's probably impossible to define.

The courts ruled that content can be scraped from the open web legally. From your examples 1. YouTube content: ok because the scraped content becomes derivative works which falls under free use 2. Doxxing people: if the data was already on the web and a human found it then the person was already doxed 3. Bank transactions: are they behind a login wall? Not Ok to scrap 4. VoIP calls: again behind a wall? Not Ok to scr…

No, the courts ruled that in the case of a search engine, scraping was ok because the results are displayed to the user who then clicks on a link that brings them back to the original source. Seems like many in the web development world like to ignore that second part.

AI responses do not bring the user back to the original source/site, the therefore the data use isn’t in line with that reasoning, and thus is a copyright violation.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#28
post #25

Earlier quoted context omitted.

Sure. But there's a fair amount of US Copyright law that says this has never been the case. To test this hypothesis, upload a video to YouTube with music by Prince playing in the background.

Copyright law is more of an ideal, and if it had "never been the case", there wouldn't have been any content to DMCA off of YouTube in the first place. US copyright law is fighting the tide, not conveying a natural order.

I think you might be the only one in this conversation setting up the strawman that copyright law is not natural law.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#29
post #25

Earlier quoted context omitted.

Copyright law is more of an ideal, and if it had "never been the case", there wouldn't have been any content to DMCA off of YouTube in the first place. US copyright law is fighting the tide, not conveying a natural order.

I think you might be the only one in this conversation setting up the strawman that copyright law is not natural law.

You characterizing it as a strawman with no real rebuttal is pretty strawman-ish yourself.
Post reply on HN