Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

11–20 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#11
post #5

I firmly believe that corps like these don't deserve the benefit of the doubt. Google, Brave and really anyone big enough to allow themselves to do bad things and get away with it must adhere to a standard where they proactively show their stuff doesn't have malicious intents.

As always, if the product is free, you are the product...

Did you read the article? This is about a paid web crawling api that the author thinks is too good or something. Nothing about a free product

Re: The shady world of Brave selling copyrighted data for AI training

#12
post #9

I think this title is overstated. It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt. (Also, crawling as a service has been a thing for a while.)

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

Is there something wrong with accessing information that someone has posted for public access?

Re: The shady world of Brave selling copyrighted data for AI training

#13
post #9

I think this title is overstated. It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt. (Also, crawling as a service has been a thing for a while.)

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

I think the first one seems to be a case where Brave just has incomplete information about licensing so for the Wikipedia data and other CCthey need to provide a link.

The second doesn't seem like a problem to me as long as they respect robots.txt

Re: The shady world of Brave selling copyrighted data for AI training

#14
post #12
post #9

Earlier quoted context omitted.

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

Is there something wrong with accessing information that someone has posted for public access?

> Is there something wrong with accessing information that someone has posted for public access?

The Wikipedia example is glaring. They’re scraping content, stripping attribution and reselling it with a right to lock it down in a way that is not allowed by the original license.

Brave is laundering copyleft content while lying to their customers by selling a license they can’t give. If you’d like, you can sidestep the morality of copyright entirely and focus on the plagiarism and fraud.

Re: The shady world of Brave selling copyrighted data for AI training

#15
From article:

> without any worry for copyright infringement because Brave acts as a middleman.

This isn’t how law works. Unless Brave is explicitly indemnifying all their customers (which their lawyers would have to be insane to let them do), any trouble you could get in, is going to be 100% your problem. Pointing the finger at Brave could theoretically get them in trouble too, but would in no way let you off the hook.

Re: The shady world of Brave selling copyrighted data for AI training

#16
post #2

How long until IP works its way onto ai training data or ais themselves? Ie that for some specific instance, the training is intentionally wrong, so as to check and prove that there has been a breach of IP.

Depends, how do you distinguish humans acquiring knowledge by ingesting copyrighted content vs. a human using an AI that ingested copyrighted content?

Re: The shady world of Brave selling copyrighted data for AI training

#17
post #9

Earlier quoted context omitted.

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

I think the first one seems to be a case where Brave just has incomplete information about licensing so for the Wikipedia data and other CCthey need to provide a link. The second doesn't seem like a problem to me as long as they respect robots.txt

You didn't actually answer the question, at best you've sidestepped it by claiming that the dodgy shit is either by accident, or really not that bad. Maybe so.

But your original claim wasn't just "Brave are technically not doing anything illegal" or "they're no worse than the others". It was praising them for being better than the others, that they're the only ones trying to do the right thing. And for these example it's just not true, they're outright worse than the industry standard.

So, to repeat, what makes you think that "Brave is trying to do the right thing while other companies aren't even attempting"?

Re: The shady world of Brave selling copyrighted data for AI training

#18
Why use brave if my info is already being leaked by third parties? E.g. experian. Is it worth the inconvenience and their repeated tricky attempts at monetizing their security conscious niche? Not being facetious, just a real question from a non security conscious person.

Re: The shady world of Brave selling copyrighted data for AI training

#19
post #9

Earlier quoted context omitted.

> It seems like Brave is trying to do the right thing here vs other companies that don't even make the attempt I feel like I'm missing something. What the article claims they're doing is: 1. Misrepresenting what rights they have, and selling access to those rights. 2. Stealth-crawling the web, hiding from the webmasters just how much Brave is crawling their site, and making it impossible to block just their crawler.…

I think the first one seems to be a case where Brave just has incomplete information about licensing so for the Wikipedia data and other CCthey need to provide a link. The second doesn't seem like a problem to me as long as they respect robots.txt

I think you're missing the point. This is one example that uses a specific license, there are countless other licenses.

And you don't seem to have read the article either, because clearly it was explained that they don't respect robots.txt because they have no user-agent.

Re: The shady world of Brave selling copyrighted data for AI training

#20
post #16
post #2

How long until IP works its way onto ai training data or ais themselves? Ie that for some specific instance, the training is intentionally wrong, so as to check and prove that there has been a breach of IP.

Depends, how do you distinguish humans acquiring knowledge by ingesting copyrighted content vs. a human using an AI that ingested copyrighted content?

> how do you distinguish humans acquiring knowledge by ingesting copyrighted content vs. a human using an AI that ingested copyrighted content

Doesn’t matter when the content is reproduced verbatim, as Brave is doing. If I memorise your content and then repeat it as my own, I’m not somehow off the hook for copyright violation and plagiarism.

Post reply on HN