Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

71–80 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#71
post #44

Earlier quoted context omitted.

Unpopular opinion time: A ML model is clearly a derivative work of its input. Here's what I think would be fair: Anyone who holds copyright in something used as part of a training corpus is owed a proportional share of the cash flow resulting from use of the resulting models. (Cash flow, not profits, because it's too easy to use accounting tricks to make profits disappear). In the case of intermediaries (e.g., social…

I don't know what a fair settlement would be but I'm looking forward to a copyright-holder suing OpenAI to obtain one. These companies have no value if copyright can be enforced on their training data.

I think there are ways around it. The simplest would be to generate replacement data, for example by paraphrasing the original, or summarising, or turning it into question-answer pairs. In this new format it can serve as training data for a clean LLM. Of course the public domain data would be used directly, no need to go synthetic there.

An important direction would be to train copyright attribution models, and diff-models to detect when a work is infringing on another, by direct comparison. They would be useful to filter both the training set and the model outputs.

Re: The shady world of Brave selling copyrighted data for AI training

#72

Earlier quoted context omitted.

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

They don't replace ads in-page, but they do something very very similar. They block the in-page ads and instead provide their own ads through popup notifications. So they are replacing advertisements on websites.

Every time this comes up, the argument is the same. People always forget:

- The ad blocker works separately from their own ad service.

- Their own ads are opt-in.

- People receive 70% of the revenue from the ads they see.

- The ads from Brave do not track you and whatever personalisation happens in-device, no data is mined.

So, no. They are not "replacing" anything. They are not stealing anyone's revenue (and no matter how much Linus from LTT argues, he is not entitled to any revenue just because I watched any of his videos) and Brave's own ads are from deals that they closed themselves and a essentially fraud-proof compared with whatever payouts are given by largest ad networks.

In other words, they are just offering something that happens to be infinitely more user-focused than the status quo. Every attempt at framing this as unethical came from an uninformed or biased source.

Re: The shady world of Brave selling copyrighted data for AI training

#73
post #25

Earlier quoted context omitted.

If you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user. If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Altho…

The rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.

> Even Google doesn't do that

At least Bing did, though. https://news.ycombinator.com/item?id=2169793

Re: The shady world of Brave selling copyrighted data for AI training

#74
post #40

Earlier quoted context omitted.

The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true. If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false. For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Pl…

It actually doesn’t even matter if LLMs reproduce copyrighted data from their training. The issue is that a human copied the data from its source into memory for use in training, and this copy was likely not fair use under cases like MAI Systems . The Supreme Court hasn’t ruled on a software case like this, as far as I know. But given the recent 7-2 decision against Andy Warhol’s estate for his copying of photographs…

So how is that supposed to work with people sending it legally obtained copyrighted materials for an analyze?

Re: The shady world of Brave selling copyrighted data for AI training

#75
post #63

Earlier quoted context omitted.

It's not clear that "data mining" covers this use. These models are huge, big enough that they can just contain direct copies of copyrighted works. They've been shown to reproduce them relatively easily. The argument is that they've actually generalized enough or learned enough that they're now no longer the sum of the dataset. I can definitely see that being possible but the way the technology works it's really hard…

My reading of the relevant laws would actually lead me to believe that this is not a problem, as long as those reproductions are not returned and the eights holder did not opt out. But courts might decide differently. Regarding the copyright of returned material here is a good discussion: https://copyrightblog.kluweriplaw.com/2023/05/09/generative-...

> as long as those reproductions are not returned

That’s the author’s entire gripe. Brave reproduced a Wikipedia entry without attribution and then slapped a copyright on it to boot.

Re: The shady world of Brave selling copyrighted data for AI training

#76

Earlier quoted context omitted.

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

opera offers a free vpn and builtin adblocker i would use it daily if the UI/UX was better, or more similar to firefox

If am not wrong, Opera is owned by some Chinese company and they are known for doing some really shady stuff [0][1] in African countries.

[0] https://blogs.opera.com/africa/2022/05/free-data-with-opera-...

[1] https://www.androidpolice.com/2020/01/21/opera-predatory-loa...

Re: The shady world of Brave selling copyrighted data for AI training

#77
post #61

Earlier quoted context omitted.

Seems you're wrong. Here's an archive link to Braves page in '16 were their planned add replacement is explained: https://archive.md/W0k4j

One minor note: they made a slight change in plans from that initial design. How it works now is that when Brave replaces an ad, they put the new ad in a popup, not in-page

Correct. I've been using Brave since their very first versions on the desktop, and there never was any in-page ad insertion.

The one type of in-page modification they used to do is that they would add a "tip" button to the content creator of some social networks like Twitter or reddit. That had nothing to do with "replacing ads" though.

> replaces an ad, they put the new ad in a popup

Incorrect. There is no 1:1 replacement. You as the user can define how often you want to receive notifications, and even then the notifications only come when you are switching context between any action. It won't interrupt you while you are watching a video, working on google doc spreadsheet or reading though HN.

Re: The shady world of Brave selling copyrighted data for AI training

#78
post #61

Earlier quoted context omitted.

Seems you're wrong. Here's an archive link to Braves page in '16 were their planned add replacement is explained: https://archive.md/W0k4j

One minor note: they made a slight change in plans from that initial design. How it works now is that when Brave replaces an ad, they put the new ad in a popup, not in-page

"replaces an ad, they put the new ad"

That's not how it works. If you turn on Brave ads, they show up every once in a while, completely independently of webpage ads. And they work whether your ad blocker is on or off.

Re: The shady world of Brave selling copyrighted data for AI training

#79

Earlier quoted context omitted.

One minor note: they made a slight change in plans from that initial design. How it works now is that when Brave replaces an ad, they put the new ad in a popup, not in-page

Correct. I've been using Brave since their very first versions on the desktop, and there never was any in-page ad insertion. The one type of in-page modification they used to do is that they would add a "tip" button to the content creator of some social networks like Twitter or reddit. That had nothing to do with "replacing ads" though. > replaces an ad, they put the new ad in a popup Incorrect. There is no 1:1 repla…

You're responding to a comment that gave you a link to their inital plan, which was literally replacing the ads.

click on it, your horizon might be broadened by the added knowledge.

Re: The shady world of Brave selling copyrighted data for AI training

#80

Earlier quoted context omitted.

opera offers a free vpn and builtin adblocker i would use it daily if the UI/UX was better, or more similar to firefox

If am not wrong, Opera is owned by some Chinese company and they are known for doing some really shady stuff [0][1] in African countries. [0] https://blogs.opera.com/africa/2022/05/free-data-with-opera-... [1] https://www.androidpolice.com/2020/01/21/opera-predatory-loa...

True, right now lots of folks from the original Opera team (including CEO) work on Vivaldi. If one day Mozilla forces me to ditch Firefox, I will probably switch to this browser.
Post reply on HN