It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…
We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…
The shady world of Brave selling copyrighted data for AI training
101–110 of 127 posts
Re: The shady world of Brave selling copyrighted data for AI training
#102Earlier quoted context omitted.
How does that align with Google Books scanning libraries full of copyrighted text, offering full reproductions of sections of the work, and then having the supreme court declare it all to be Fair Use? I think that is a far more relevant precedent here: https://en.m.wikipedia.org/wiki/Authors_Guild,_Inc._v._Googl... .
The Supreme Court declined to hear the case on appeal, which is a shade different from endorsing the decision after a hearing. That being said, it doesn’t take a lot of effort to differentiate these cases. Google was indexing copyrighted works and providing access to limited extracts. They weren’t transforming them into new works and then selling access to those new works over APIs.
Re: The shady world of Brave selling copyrighted data for AI training
#103Earlier quoted context omitted.
It actually doesn’t even matter if LLMs reproduce copyrighted data from their training. The issue is that a human copied the data from its source into memory for use in training, and this copy was likely not fair use under cases like MAI Systems . The Supreme Court hasn’t ruled on a software case like this, as far as I know. But given the recent 7-2 decision against Andy Warhol’s estate for his copying of photographs…
How does that align with Google Books scanning libraries full of copyrighted text, offering full reproductions of sections of the work, and then having the supreme court declare it all to be Fair Use? I think that is a far more relevant precedent here: https://en.m.wikipedia.org/wiki/Authors_Guild,_Inc._v._Googl... .
Re: The shady world of Brave selling copyrighted data for AI training
#104Earlier quoted context omitted.
The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true. If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false. For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Pl…
OpenAI's bias research on DALL-E revealed that most examples of regurgitation come from repeated copies of the same image in the training set. When they filtered out duplicates, DALL-E stopped drawing training examples. The problem is that filtering the training set is naively O(n^2) and n is already extremely large for DALL-E. For LLMs, it's comically huge, plus now you have to do substring search. I've yet to hear…
There are standard ways to do it that are O(n), FYI.
Re: The shady world of Brave selling copyrighted data for AI training
#105The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!
I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…
Re: The shady world of Brave selling copyrighted data for AI training
#106Earlier quoted context omitted.
> The answer is no, because you reading the article didn’t dramatically degrade its market value. How about if you read a news article to write a competing one rewording and possibly citing it (one of the most common practices in news)?
How about it? Do you not think it incurs a lot of negative effects?
Re: The shady world of Brave selling copyrighted data for AI training
#107The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!
I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…
Re: The shady world of Brave selling copyrighted data for AI training
#108Re: The shady world of Brave selling copyrighted data for AI training
#109The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!
I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…
Cliqz entire history was based on this kind of thing, milking off other search engines by just deducting their ranking methods, it's parasitic. There's no cleverness about it.
Re: The shady world of Brave selling copyrighted data for AI training
#110> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…
This is (importantly) opt-in. "Brave doesn’t follow the sneaky practices of other big tech search engines. The Web Discovery Project is opt-in, and the data collected under the Web Discovery Project has specific protections to ensure anonymity." per https://support.brave.com/hc/en-us/articles/4409406835469-Wh...
Editing to add that I don't mean to imply ill will on your part, but that I think being affiliated with Brave might have you taking this type of practice a little more lightly than it probably should be taken.