Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

51–60 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#51
This discussion on fair use are always quite anglocentric.

Atricle 3 and 4 of the EU 'Copyright in the Digital Single Market' give data miners quite extensive rights.

Move operation to the EU, train a foundational model, than train a constitutional model based on that.

As much as I hate the upcoming AI regulation, the CDSM is solid.

https://academic.oup.com/grurint/article/71/8/685/6650009 https://eur-lex.europa.eu/eli/dir/2019/790/oj

Update: Fixed wrong link

Re: The shady world of Brave selling copyrighted data for AI training

#52

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

This is (importantly) opt-in.

"Brave doesn’t follow the sneaky practices of other big tech search engines. The Web Discovery Project is opt-in, and the data collected under the Web Discovery Project has specific protections to ensure anonymity." per https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

Re: The shady world of Brave selling copyrighted data for AI training

#53
post #25

Earlier quoted context omitted.

If you don’t trust Brave then, yeah, they could be doing anything in the browser or on their servers - but that snippet you quoted is a slightly out of context statement from a big document about how they collect data like this, but _don’t_ collect or store it in a way that they could associate it with a user. If you don’t trust that they’re doing what they say they are, then the document doesn’t mean anything. Altho…

The rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.

This is opt-in only. https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

Re: The shady world of Brave selling copyrighted data for AI training

#54
post #39

Earlier quoted context omitted.

I disagree. People learning how to draw does degrade the future value of copyrighted work. Imagine the future where nobody was allowed to learn to draw, existing copyright value would skyrocket!

Arguments like this are great for getting your side to go "rah rah got 'em" and really, really bad for convincing anyone else. Legal judgments generally focus on actual impacts rather than quirks that might exist in hypothetical universes.

While that may be what your parent intended I'm not entirely sure and there does exist the philosophical level discussion here. Or market economics level I guess.

If your pool of people that can learn about topic X is restricted the outputs or their labor are more expensive. Now lift a continent of billions of people out of poverty, get them access to schooling, safety etc and see the market forces do the rest.

Now equate ChatGPT et al with said billion people. Just that it runs on electricity. If quality is good enough of course. Which is hard to decide right now because of hype.

Re: The shady world of Brave selling copyrighted data for AI training

#55

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

I think some initial news articles claimed this and everyone went with it. Which is basically how it worked, replacing ads, but in a different way, and actually more annoying.. block everyone else's ads...and have their own little popup ads, and if you enabled that, you'd get paid in BAT tokens per view too.

Re: The shady world of Brave selling copyrighted data for AI training

#56
post #40

> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…

The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true. If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false. For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Pl…

It actually doesn’t even matter if LLMs reproduce copyrighted data from their training. The issue is that a human copied the data from its source into memory for use in training, and this copy was likely not fair use under cases like MAI Systems.

The Supreme Court hasn’t ruled on a software case like this, as far as I know. But given the recent 7-2 decision against Andy Warhol’s estate for his copying of photographs of Prince, this doesn’t seem like a Court that’s ready to say copying terabytes of unlicensed material for a commercial purpose is OK.

I’m going to guess this ends with Congress setting up some kind of clearinghouse for copyrighted training material: You opt in to be included, you get fees from OpenAI when they use what you added. This isn’t unprecedented: Congress set up special rules and processes for things like music recordings repeatedly over the years.

https://scholarship.law.edu/cgi/viewcontent.cgi?referer=&htt...

Re: The shady world of Brave selling copyrighted data for AI training

#57

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

As a detractor and therefore a part of the set "all detractors", I do not believe this. I just don't buy their shady marketing and try not to support engine monoculture.

Re: The shady world of Brave selling copyrighted data for AI training

#58

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

You are a victim of the Mandella effect. There never was anything related to replacing ads in-page, yet if you ask all detractors what they don't like about it, that's the first point they bring up.

They don't replace ads in-page, but they do something very very similar.

They block the in-page ads and instead provide their own ads through popup notifications.

So they are replacing advertisements on websites.

Re: The shady world of Brave selling copyrighted data for AI training

#59
post #24

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

People still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.

It's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.

Re: The shady world of Brave selling copyrighted data for AI training

#60

Earlier quoted context omitted.

Have you worked in any project that required a forked browser? I did. When we folded less than two years later, one of the CTOs biggest stated regrets was that he went with Firefox instead of Chromium. The extension story in Firefox was easily 10x harder. Interfacing with the OS as well. Getting dbus services to work was a fool's errand.

Cool so your company folded but as I said, palemoon and waterfox seem to be running just fine. Thunderbird also works. I happen to own a brwoser extension and have both chromium and Firefox extensions. I kinda know myself.

> palemoon and waterfox seem to be running just fine.

GNU/Hurd is also a very interesting alternative OS, the design is a lot more elegant than GNU/Linux, it's still under active development and it has a surprising number of active users.

It's still a very bad idea to build the foundation of your tech stack on it.

Post reply on HN