Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

21–30 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#21
post #12
post #9

Earlier quoted context omitted.

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

Is there something wrong with accessing information that someone has posted for public access?

Yes. Legalities aside, stripping attribution (author names) from contents which specifically requires keeping it, it a really shitty thing to do.

(The fact that they include original URL does not change much, given that they explicitly market it as "Data for AI" and those systems never have attribution)

Re: The shady world of Brave selling copyrighted data for AI training

#22
> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers.

Wait. Brave browser sends back to Brave Search engine about your browsing? Other search engines usage, but also crawl pages on your computer to help build their search index?

Ref: https://github.com/brave/web-discovery-project/blob/main/mod...

Re: The shady world of Brave selling copyrighted data for AI training

#23
It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage.

It's also a for-profit company and you're not the customer, as you're not paying them money.

I'd be way more worried how they're using the data they're collecting on you vs Google or MS

Re: The shady world of Brave selling copyrighted data for AI training

#24

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

People still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.

Re: The shady world of Brave selling copyrighted data for AI training

#25

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

If you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user.

If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Although that would also mean the quote is kind of meaningless…

Re: The shady world of Brave selling copyrighted data for AI training

#26
> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test:

> 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes

> 2) The nature of the copyrighted work

> 3) The amount and substantiality of the portion used in relation to the copyrighted work as a whole

> 4) The effect of the use upon the potential market for or value of the copyrighted work

[emphasis from TFA]

HN always talks about derivative work and transformativeness, but never about these. The fourth one especially seems clear in its implications for models.

Regardless, it makes it seem much less clear cut than people here often say.

Re: The shady world of Brave selling copyrighted data for AI training

#27

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants.

Mullvad

Brave

Opera

Vivaldi

Microsoft

Heck zoho is in on a browser now

What net gain does each of these companies provide over skinning chromium that isn't in Firefox?

Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with firefox, so could someone else but no

Re: The shady world of Brave selling copyrighted data for AI training

#28
post #25

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

If you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user. If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Altho…

The rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.

Re: The shady world of Brave selling copyrighted data for AI training

#29

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

Microsoft is gambling on the hope that model training will be ruled fair use. This makes it seem that outcome is unlikely.

Re: The shady world of Brave selling copyrighted data for AI training

#30

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

> I'd be way more worried how they're using the data they're collecting on you vs Google or MS

Why? They don't even have access to my emails and texts like those other companies do. I also don't see the names of their top executives and founders showing up in articles about connections to Jeffrey Epstein every few months.

Post reply on HN