Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

41–50 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#41

Earlier quoted context omitted.

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Have you worked in any project that required a forked browser? I did. When we folded less than two years later, one of the CTOs biggest stated regrets was that he went with Firefox instead of Chromium. The extension story in Firefox was easily 10x harder. Interfacing with the OS as well. Getting dbus services to work was a fool's errand.

Cool so your company folded but as I said, palemoon and waterfox seem to be running just fine.

Thunderbird also works.

I happen to own a brwoser extension and have both chromium and Firefox extensions. I kinda know myself.

Re: The shady world of Brave selling copyrighted data for AI training

#42
Unpopular opinion: the next iteration of privacy laws needs to factor in AI. If AI is allowed to slurp up PII or derogative works and the people defending it defend it with the zeal of cryptobros then we're in for a decade of real pain in terms of both copyright law, PII, and IP exposure.

Re: The shady world of Brave selling copyrighted data for AI training

#43
post #25

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

If you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user. If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Altho…

How do they detect if someone poisons their data, if they not at least associate IP addresses to the data?

Re: The shady world of Brave selling copyrighted data for AI training

#44

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

Unpopular opinion time: A ML model is clearly a derivative work of its input. Here's what I think would be fair: Anyone who holds copyright in something used as part of a training corpus is owed a proportional share of the cash flow resulting from use of the resulting models. (Cash flow, not profits, because it's too easy to use accounting tricks to make profits disappear). In the case of intermediaries (e.g., social…

I don't know what a fair settlement would be but I'm looking forward to a copyright-holder suing OpenAI to obtain one. These companies have no value if copyright can be enforced on their training data.

Re: The shady world of Brave selling copyrighted data for AI training

#45
post #42

Unpopular opinion: the next iteration of privacy laws needs to factor in AI. If AI is allowed to slurp up PII or derogative works and the people defending it defend it with the zeal of cryptobros then we're in for a decade of real pain in terms of both copyright law, PII, and IP exposure.

AI is going to do that irregardless- the debate is essentially going to revolve around how and what people can make new commercial works from that data.

Re: The shady world of Brave selling copyrighted data for AI training

#46
post #18

Why use brave if my info is already being leaked by third parties? E.g. experian. Is it worth the inconvenience and their repeated tricky attempts at monetizing their security conscious niche? Not being facetious, just a real question from a non security conscious person.

It’s built in Ad blocker and other features are heads and tails above anything else I’ve used before, personally.

Re: The shady world of Brave selling copyrighted data for AI training

#47
Brave continues to be shady. They claim to respect robots.txt but don't identify their crawler if you want to block it.

> They don't mention their crawler anywhere in their docs, either. So, if you wanted to block Brave from crawling and indexing and ultimately selling your content to third parties, your only option for the time being would be to block all crawlers, which is how Brave would be able to "respect robots.txt".

Re: The shady world of Brave selling copyrighted data for AI training

#48
post #39

Earlier quoted context omitted.

The point being raised is quite specific. Not sure if you’re willingly ignoring it or what? The answer is no, because you reading the article didn’t dramatically degrade its market value. An AI ingesting all content on the internet and then being ultra-effective at frontrunning that content for a large number of future readers does degrade its market value (and subsumes it into the model’s value).

I disagree. People learning how to draw does degrade the future value of copyrighted work. Imagine the future where nobody was allowed to learn to draw, existing copyright value would skyrocket!

Arguments like this are great for getting your side to go "rah rah got 'em" and really, really bad for convincing anyone else.

Legal judgments generally focus on actual impacts rather than quirks that might exist in hypothetical universes.

Re: The shady world of Brave selling copyrighted data for AI training

#49
post #42

Unpopular opinion: the next iteration of privacy laws needs to factor in AI. If AI is allowed to slurp up PII or derogative works and the people defending it defend it with the zeal of cryptobros then we're in for a decade of real pain in terms of both copyright law, PII, and IP exposure.

The fun part is that the GDPR already does. The answer is you're not allowed to use personal data for AI. (And "personal data" here covers things like all public social media posts)

Facebook recently got told by the CJEU that, no, they can't use people's posts to target advertisements. Even if those ads are what's paying for the platform. That you can't claim such processing as "part of the contract" unless it is absolutely necessary in the same way the post office needs an address to send a parcel.

If Facebook can't even do that, there is no way LLMs will be allowed. (And remember. The GDPR does not care if your system doesn't distribute personal data. Any kind of processing at all falls under the GDPR's requirements)

OpenAI is already being chased by the EU's privacy agencies. Right now they're in the process of asking pointed questions, things will heat up after that.

Post reply on HN