Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

121–127 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#121
post #52

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

This is (importantly) opt-in. "Brave doesn’t follow the sneaky practices of other big tech search engines. The Web Discovery Project is opt-in, and the data collected under the Web Discovery Project has specific protections to ensure anonymity." per https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

Opt-in or not, “sneaky” is a marketing term and not a UX principle. E.g. showing the user an example of real ROI from providing their data.

That said, stuff like Jedi Blue and Project Bernanke suggest Brave could just disclose they support competitive markets.

Re: The shady world of Brave selling copyrighted data for AI training

#122
post #90
post #71

Earlier quoted context omitted.

I think there are ways around it. The simplest would be to generate replacement data, for example by paraphrasing the original, or summarising, or turning it into question-answer pairs. In this new format it can serve as training data for a clean LLM. Of course the public domain data would be used directly, no need to go synthetic there. An important direction would be to train copyright attribution models, and diff-…

Would automated paraphrasing not be a derivative work of the original?

So you think any paraphrase of a copyrighted phrase is in copyright violation? That's like owning the idea itself. Is any utterance similar to this one now forbidden?

Re: The shady world of Brave selling copyrighted data for AI training

#123

Earlier quoted context omitted.

I don't know a lot about this particular approach but your comment that it's just using Google results is blatantly false. It all depends on the search engine that the brave user is leveraging, or no search engine if they type in the URL directly into the header.

Nonsense. This naive idea that Brave innocuously looks at user's traffic patterns. Google owns 95% of the market in most Western markets. There's no "blatantly false" about that. They scrape search engine results and present them as their own. Do 10,000 searches on Google and Brave and you'll see how similar they are. It's as simple as that, scraping by sleight of hand. Why can't they be a normal search engine - beca…

I would expect most people today find interesting content via social media, and not search engines.

Re: The shady world of Brave selling copyrighted data for AI training

#124
post #90

Earlier quoted context omitted.

Would automated paraphrasing not be a derivative work of the original?

So you think any paraphrase of a copyrighted phrase is in copyright violation? That's like owning the idea itself. Is any utterance similar to this one now forbidden?

I think if you automate paraphrasing from an original work to use that original work on bulk somehow, yes.

How do you even automate paraphrasing without training it on lots of original work? It's infringement all the way down.

Re: The shady world of Brave selling copyrighted data for AI training

#125

Earlier quoted context omitted.

I don't know a lot about this particular approach but your comment that it's just using Google results is blatantly false. It all depends on the search engine that the brave user is leveraging, or no search engine if they type in the URL directly into the header.

Nonsense. This naive idea that Brave innocuously looks at user's traffic patterns. Google owns 95% of the market in most Western markets. There's no "blatantly false" about that. They scrape search engine results and present them as their own. Do 10,000 searches on Google and Brave and you'll see how similar they are. It's as simple as that, scraping by sleight of hand. Why can't they be a normal search engine - beca…

I had a discussion about this on Twitter with Brendan Eich. He became hostile very quickly. He is not a very nice person.

Re: The shady world of Brave selling copyrighted data for AI training

#126
post #87

Earlier quoted context omitted.

> The answer is no, because you reading the article didn’t dramatically degrade its market value. How about if you read a news article to write a competing one rewording and possibly citing it (one of the most common practices in news)?

How about it? Do you not think it incurs a lot of negative effects?

What's the alternative, each news event is first to publish exclusivity regardless of quality? No synthesizing multiple stories into a linked narrative?

Re: The shady world of Brave selling copyrighted data for AI training

#127

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

We have these cropping up like ants. Mullvad Brave Opera Vivaldi Microsoft Heck zoho is in on a browser now What net gain does each of these companies provide over skinning chromium that isn't in Firefox? Last time I asked brave fanboys why they don't redskin Firefox and the response was "Firefox is pita to build" all the while we have projects like palemoon and waterfox that are hobby projects. If they can work with…

Uh. "cropping up like ants" is definitely a take. Not a good one, as most the browsers here had their first release date of 199*. I will list them out.

Mullvad, is the Tor Browser with the Mullvad VPN included, and released 2023. However, the Tor Browser, which it effectively is, is from 2002.

Brave, the one in this article, is from 2019.

Opera is from 1994.

Vivaldi is from 2015, and is developed by Opera's previous dev-team after a bad sale to a Chinese company.

Microsoft's first browser, Internet Explorer, is from 1995.

I can not comment about Zoho's browser, as i know little about it.

Post reply on HN