Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

31–40 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#31

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

opera offers a free vpn and builtin adblocker

i would use it daily if the UI/UX was better, or more similar to firefox

Re: The shady world of Brave selling copyrighted data for AI training

#32

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

That’s not at all clear to me. IANAL but first of all it’s a balancing test, not a bright-line test. The judge could focus on any one factor and make an argument for either side quite easily.

Second, “use” here could mean one of two things: training or inference. It’s publishing the results of inference that can lead to actual effects on the market, not the training.

At the end of the day, someone has to prove tangible harm.

Re: The shady world of Brave selling copyrighted data for AI training

#33

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

Microsoft is gambling on the hope that model training will be ruled fair use. This makes it seem that outcome is unlikely.

Do you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?

Re: The shady world of Brave selling copyrighted data for AI training

#34

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

Re: The shady world of Brave selling copyrighted data for AI training

#35

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

I would look at #1 here. Crawling the Internet to collect information is one thing. (And people putting text on the web without requiring authentication seem to be granting at least some kind of license to anyone who sends a GET request.). But crawling the Internet (via centralized robots or users’ browsers), then storing that data and charging money to others for rights to that data (as Brave seems to be doing, quite explicitly) seems like it deserves a very different evaluation under factor #1.

Re: The shady world of Brave selling copyrighted data for AI training

#36
post #33

Earlier quoted context omitted.

Microsoft is gambling on the hope that model training will be ruled fair use. This makes it seem that outcome is unlikely.

Do you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?

The point being raised is quite specific. Not sure if you’re willingly ignoring it or what?

The answer is no, because you reading the article didn’t dramatically degrade its market value.

An AI ingesting all content on the internet and then being ultra-effective at frontrunning that content for a large number of future readers does degrade its market value (and subsumes it into the model’s value).

Re: The shady world of Brave selling copyrighted data for AI training

#37

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

Unpopular opinion time:

A ML model is clearly a derivative work of its input.

Here's what I think would be fair:

Anyone who holds copyright in something used as part of a training corpus is owed a proportional share of the cash flow resulting from use of the resulting models. (Cash flow, not profits, because it's too easy to use accounting tricks to make profits disappear).

In the case of intermediaries (e.g., social media like reddit & twitter) those intermediaries could take a cut before passing it on to the original authors.

Obviously hellishly difficult to administer so it's unlikely to happen but I don't see a better answer.

Re: The shady world of Brave selling copyrighted data for AI training

#38

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Have you worked in any project that required a forked browser?

I did. When we folded less than two years later, one of the CTOs biggest stated regrets was that he went with Firefox instead of Chromium. The extension story in Firefox was easily 10x harder. Interfacing with the OS as well. Getting dbus services to work was a fool's errand.

Re: The shady world of Brave selling copyrighted data for AI training

#39
post #33

Earlier quoted context omitted.

Do you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?

The point being raised is quite specific. Not sure if you’re willingly ignoring it or what? The answer is no, because you reading the article didn’t dramatically degrade its market value. An AI ingesting all content on the internet and then being ultra-effective at frontrunning that content for a large number of future readers does degrade its market value (and subsumes it into the model’s value).

I disagree. People learning how to draw does degrade the future value of copyrighted work. Imagine the future where nobody was allowed to learn to draw, existing copyright value would skyrocket!

Re: The shady world of Brave selling copyrighted data for AI training

#40

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true.

If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false.

For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Plenty can be found on HN. It doesn't generate a new valid windows serial key on the fly, it's memorized them.

If one wants to be cynical, it's not hard to see OpenAI/etc patching in filters to remove copyrighted content from the output precisely because it's legally catastrophic for their "fair use" claim to have the model spit out copyrighted content. As this is both copyright infringement by itself, and evidence that no matter how the internals of these models work, they store some of the training data anyway.

Post reply on HN