Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

101–110 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#101

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Mullvad is actually a Firefox fork and it directly uses Tor's privacy enhancements[0] to Firefox for a private web browsing experience. As a matter of fact, it really looks like Tor Browser but with a VPN baked in instead of Tor.

[0] https://mullvad.net/en/browser

Re: The shady world of Brave selling copyrighted data for AI training

#102
post #91

Earlier quoted context omitted.

How does that align with Google Books scanning libraries full of copyrighted text, offering full reproductions of sections of the work, and then having the supreme court declare it all to be Fair Use? I think that is a far more relevant precedent here: https://en.m.wikipedia.org/wiki/Authors_Guild,_Inc._v._Googl... .

The Supreme Court declined to hear the case on appeal, which is a shade different from endorsing the decision after a hearing. That being said, it doesn’t take a lot of effort to differentiate these cases. Google was indexing copyrighted works and providing access to limited extracts. They weren’t transforming them into new works and then selling access to those new works over APIs.

OpenAI is also providing access to limited extracts. Google wasn't selling this over an API, they were providing "free" access to it while displaying ads to the user. Would the courts see this manner of monetization to be different enough that settled case law wouldn't apply?

Re: The shady world of Brave selling copyrighted data for AI training

#103
post #91

Earlier quoted context omitted.

It actually doesn’t even matter if LLMs reproduce copyrighted data from their training. The issue is that a human copied the data from its source into memory for use in training, and this copy was likely not fair use under cases like MAI Systems . The Supreme Court hasn’t ruled on a software case like this, as far as I know. But given the recent 7-2 decision against Andy Warhol’s estate for his copying of photographs…

How does that align with Google Books scanning libraries full of copyrighted text, offering full reproductions of sections of the work, and then having the supreme court declare it all to be Fair Use? I think that is a far more relevant precedent here: https://en.m.wikipedia.org/wiki/Authors_Guild,_Inc._v._Googl... .

Google also bought copies of each book, I believe, which makes it another step removed from standard ML practice.

Re: The shady world of Brave selling copyrighted data for AI training

#104
post #40

Earlier quoted context omitted.

The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true. If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false. For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Pl…

OpenAI's bias research on DALL-E revealed that most examples of regurgitation come from repeated copies of the same image in the training set. When they filtered out duplicates, DALL-E stopped drawing training examples. The problem is that filtering the training set is naively O(n^2) and n is already extremely large for DALL-E. For LLMs, it's comically huge, plus now you have to do substring search. I've yet to hear…

> The problem is that filtering the training set is naively O(n^2)

There are standard ways to do it that are O(n), FYI.

Re: The shady world of Brave selling copyrighted data for AI training

#105
post #50

The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!

I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…

[deleted]

Re: The shady world of Brave selling copyrighted data for AI training

#106
post #87

Earlier quoted context omitted.

> The answer is no, because you reading the article didn’t dramatically degrade its market value. How about if you read a news article to write a competing one rewording and possibly citing it (one of the most common practices in news)?

How about it? Do you not think it incurs a lot of negative effects?

[deleted]

Re: The shady world of Brave selling copyrighted data for AI training

#107
post #50

The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!

I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…

.es?

Re: The shady world of Brave selling copyrighted data for AI training

#109
post #50

The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!

I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…

It's not a clever approach, it's basically scraping Google results because that's where your users are searching. You follow the bread crumbs from Google searches.

Cliqz entire history was based on this kind of thing, milking off other search engines by just deducting their ranking methods, it's parasitic. There's no cleverness about it.

Re: The shady world of Brave selling copyrighted data for AI training

#110
post #52

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

This is (importantly) opt-in. "Brave doesn’t follow the sneaky practices of other big tech search engines. The Web Discovery Project is opt-in, and the data collected under the Web Discovery Project has specific protections to ensure anonymity." per https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

I think mentioning your affiliation with Brave might go a long way in contextualizing why you are defending/rationalizing this (even if opt-in).

Editing to add that I don't mean to imply ill will on your part, but that I think being affiliated with Brave might have you taking this type of practice a little more lightly than it probably should be taken.

Post reply on HN