Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

61–70 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#61

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

Seems you're wrong. Here's an archive link to Braves page in '16 were their planned add replacement is explained:

https://archive.md/W0k4j

Re: The shady world of Brave selling copyrighted data for AI training

#62
post #24

Earlier quoted context omitted.

People still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.

It's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.

Ah that makes sense why brave fans post a bit like cryptobros.

Re: The shady world of Brave selling copyrighted data for AI training

#63

This discussion on fair use are always quite anglocentric. Atricle 3 and 4 of the EU 'Copyright in the Digital Single Market' give data miners quite extensive rights. Move operation to the EU, train a foundational model, than train a constitutional model based on that. As much as I hate the upcoming AI regulation, the CDSM is solid. https://academic.oup.com/grurint/article/71/8/685/6650009 https://eur-lex.europa.eu/e…

It's not clear that "data mining" covers this use. These models are huge, big enough that they can just contain direct copies of copyrighted works. They've been shown to reproduce them relatively easily. The argument is that they've actually generalized enough or learned enough that they're now no longer the sum of the dataset. I can definitely see that being possible but the way the technology works it's really hard to know if that has happened or if what's happening instead is a bunch of copyright washing.

There are some things that would make for good faith displays by the players in the space. For example, Microsoft has been investing a lot and yet their code offering is not trained on their internal code base. Same for Google. Start by doing that and I'll entertain the argument that your tools are fair use or data mining.

Re: The shady world of Brave selling copyrighted data for AI training

#64

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Brave literally started on Gecko.

Re: The shady world of Brave selling copyrighted data for AI training

#65
post #53

Earlier quoted context omitted.

The rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.

This is opt-in only. https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

[deleted]

Re: The shady world of Brave selling copyrighted data for AI training

#66

Earlier quoted context omitted.

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Brave literally started on Gecko.

Is brave currently forking chromium or not?

By your logic, opera was having their own engine till 2013. So what?

Re: The shady world of Brave selling copyrighted data for AI training

#67
post #61

Earlier quoted context omitted.

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

Seems you're wrong. Here's an archive link to Braves page in '16 were their planned add replacement is explained: https://archive.md/W0k4j

One minor note: they made a slight change in plans from that initial design.

How it works now is that when Brave replaces an ad, they put the new ad in a popup, not in-page

Re: The shady world of Brave selling copyrighted data for AI training

#68

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

This specific feature is already opt-in, but historically the answer has always been "yes" for dozens of 'features' like this that fly under the radar until users start complaining, and then eventually get converted to opt-in or removed in order to save face.

Re: The shady world of Brave selling copyrighted data for AI training

#69
post #63

This discussion on fair use are always quite anglocentric. Atricle 3 and 4 of the EU 'Copyright in the Digital Single Market' give data miners quite extensive rights. Move operation to the EU, train a foundational model, than train a constitutional model based on that. As much as I hate the upcoming AI regulation, the CDSM is solid. https://academic.oup.com/grurint/article/71/8/685/6650009 https://eur-lex.europa.eu/e…

It's not clear that "data mining" covers this use. These models are huge, big enough that they can just contain direct copies of copyrighted works. They've been shown to reproduce them relatively easily. The argument is that they've actually generalized enough or learned enough that they're now no longer the sum of the dataset. I can definitely see that being possible but the way the technology works it's really hard…

My reading of the relevant laws would actually lead me to believe that this is not a problem, as long as those reproductions are not returned and the eights holder did not opt out. But courts might decide differently.

Regarding the copyright of returned material here is a good discussion:

https://copyrightblog.kluweriplaw.com/2023/05/09/generative-...

Re: The shady world of Brave selling copyrighted data for AI training

#70
post #18

Why use brave if my info is already being leaked by third parties? E.g. experian. Is it worth the inconvenience and their repeated tricky attempts at monetizing their security conscious niche? Not being facetious, just a real question from a non security conscious person.

You get degoogled Chromium with e2ee bookmarks etc. sync and a lot of nice convenience features like vertical tabs and mobile background video playback.

And if it's your cup of tea, they let you straight up pay money for the search engine.

Post reply on HN