Live data from Hacker News

Microsoft says that it's okay to steal web content because it's 'freeware.'

windowscentral.com

1–10 of 29 posts

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#2
Before we get too upset... can we verify this is MSFT's official position? I suspect this may be hyperbole. It could be Sulyman was constructing a hypothetical point that didn't survive translation into click-bait. That being said... MSFT has a history of chicanery. I'm off to try to find original sources. If anyone else has any, please provide a link.

FWIW... I found a few videos related to Endicott's story:

* This is a quick 5 minute video where Suleyman talks about how indeterminacy is good. So... you know... it's a good think that Co-Pilot can't tell you why it thinks it needs to dump 800 line of java code into your hello world program. At around 3:44, he confuses LLMs (with a surface understanding of syntax married with a markov chain on steroids) with people (who as best we can tell have a different understanding of the thing represented.) Corporate management confusing the the map with the territory? Who could have forseen such a thing: https://youtu.be/GsGFYoIx1YM

* This one seems to be the longer version, but I'm still looking for where Endicott's quote comes from, but around the 14minute mark is where the conversation turns towards "who owns the ip" used to train LLMs and the terms "Fair Use" and "Freeware" are used around the 14m50s mark: https://youtu.be/lPvqvt55l3A

[EDIT: So... yes... get out the pitch-forks... Microsoft is saying anything on the web is inherently freeware or subject to fair use even if you think you remember putting a copyright notice on it (or, as is mentioned in US copyright law, the creator automatically receives copyright protections upon creation of the work.)]

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#3
It’s a pro-AI position but not really controversial?

My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training.

> In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically.

> "With respect to content that is already on the open web, the social contract of that content since the 90s has been that it is fair use. Anyone can copy it, recreate with it, reproduce with it. That has been freeware, if you like. That's been the understanding," said Suleyman.

> "There's a separate category where a website or a publisher or a news organization had explicitly said, 'do not scrape or crawl me for any other reason than indexing me so that other people can find that content.' That's a gray area and I think that's going to work its way through the courts."

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#4

Before we get too upset... can we verify this is MSFT's official position? I suspect this may be hyperbole. It could be Sulyman was constructing a hypothetical point that didn't survive translation into click-bait. That being said... MSFT has a history of chicanery. I'm off to try to find original sources. If anyone else has any, please provide a link. FWIW... I found a few videos related to Endicott's story: * This…

Suleyman is MSFT's AI CEO. If this is his opinion, that's how the division is operating.

Obviously Microsoft is a huge company so he doesn't represent everyone, that would be impossible. But it's a massive red flag.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#6

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

Sure. But this is not what US copyright law says.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#8

It's ironic that Microsoft used copyright protection and IP law for years to secure a dominant market position, and now they don't need to play by the same rules because "something something AI".

Truthfully, I agree with this new stance of theirs and believe it's always been the case, but there's no doubt that they're only adopting it now because it's become advantageous.

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#9

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

> My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training.

In the absence of an explicit copyright, the default copyright is "all rights reserved". Putting something in a public space is not license to reproduce it, however much I might disagree with this position

Re: Microsoft says that it's okay to steal web content because it's 'freeware.'

#10

It’s a pro-AI position but not really controversial? My reading is he is saying content that is not under an explicit license for usage, that is made available publicly and freely, is fair game for training. > In his remarks, Suleyman claimed that all content shared on the web is available to be used for AI training unless a content producer says otherwise specifically. > "With respect to content that is already on t…

I wish someone would define what scraping "open web" means since physical media has been dead for a while. YouTube content? Doxxing people? Archiving books and old video games? Bank transactions? VoIP calls?

Why is it potentially ok for AI to copy these things, but if a human does it they get into trouble?

We need an exhaustive list and that's probably impossible to define.

Post reply on HN