Live data from Hacker News

The shady world of Brave selling copyrighted data for AI training

stackdiary.com

111–120 of 127 posts

Re: The shady world of Brave selling copyrighted data for AI training

#111

Earlier quoted context omitted.

I'm Sampson, from the Brave team. The Web Discovery Project is a clever approach. For Brave to compete with Google, and offer a truly novel index of the Web, a novel approach must be taken. The WDP is an opt-in, privacy-preserving approach which gives Brave a fighting chance against the Search incumbants. Due to our preference of "Can't be evil" over "Don't be evil," the WDP is not only designed with privacy and anon…

It's not a clever approach, it's basically scraping Google results because that's where your users are searching. You follow the bread crumbs from Google searches. Cliqz entire history was based on this kind of thing, milking off other search engines by just deducting their ranking methods, it's parasitic. There's no cleverness about it.

I don't know a lot about this particular approach but your comment that it's just using Google results is blatantly false. It all depends on the search engine that the brave user is leveraging, or no search engine if they type in the URL directly into the header.

Re: The shady world of Brave selling copyrighted data for AI training

#112
post #102

Earlier quoted context omitted.

The Supreme Court declined to hear the case on appeal, which is a shade different from endorsing the decision after a hearing. That being said, it doesn’t take a lot of effort to differentiate these cases. Google was indexing copyrighted works and providing access to limited extracts. They weren’t transforming them into new works and then selling access to those new works over APIs.

OpenAI is also providing access to limited extracts. Google wasn't selling this over an API, they were providing "free" access to it while displaying ads to the user. Would the courts see this manner of monetization to be different enough that settled case law wouldn't apply?

OpenAI isn’t doing anything like what Google was doing with Books. It’s not hard for laymen to see that, and it’s going to be obvious to any judge who hears a case.

Imagine OpenAI had invented a software program that turned any written text into an animated cartoon enacting the text. That would obviously be creating a derivative work and outside fair use bounds. That they mix a bunch of works (copyrighted and otherwise) into a piece of software doesn’t allow them to escape that basic analysis.

Google showed a “clip” of the original work, no different in scope than Siskel & Ebert showing a clip of a film as they reviewed it. The uses are not comparable.

Re: The shady world of Brave selling copyrighted data for AI training

#113
post #33

Earlier quoted context omitted.

Microsoft is gambling on the hope that model training will be ruled fair use. This makes it seem that outcome is unlikely.

Do you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?

Yes it is considered fair use but it's also completely irrelevant because we're talking about a computer program not a person.

Re: The shady world of Brave selling copyrighted data for AI training

#114

It's always surprising to me when I hear people using the brave browser... It's by a company that initially tried to replace their blocked ads with their own "safe and non-intrusive" ads as far as I remember, until they backpaddled because of the outrage. It's also a for-profit company and you're not the customer, as you're not paying them money. I'd be way more worried how they're using the data they're collecting o…

There are multiple ways you can pay Brave.

https://brave.com/firewall-vpn/ https://account.brave.com/?intent=checkout&product=search https://brave.com/search/api/

Re: The shady world of Brave selling copyrighted data for AI training

#115

Earlier quoted context omitted.

Arguments like this are great for getting your side to go "rah rah got 'em" and really, really bad for convincing anyone else. Legal judgments generally focus on actual impacts rather than quirks that might exist in hypothetical universes.

While that may be what your parent intended I'm not entirely sure and there does exist the philosophical level discussion here. Or market economics level I guess. If your pool of people that can learn about topic X is restricted the outputs or their labor are more expensive. Now lift a continent of billions of people out of poverty, get them access to schooling, safety etc and see the market forces do the rest. Now e…

Your sentiment is exactly what I intended, albeit I was terse and a little facetious. ChatGPT is like introducing a bunch of new skilled labor, it’s just for the first time this skilled labor isn’t human. The fact that this skilled labor learned from copyrighted material is like saying human labor learned from copyrighted material.

Re: The shady world of Brave selling copyrighted data for AI training

#116

Earlier quoted context omitted.

It's not a clever approach, it's basically scraping Google results because that's where your users are searching. You follow the bread crumbs from Google searches. Cliqz entire history was based on this kind of thing, milking off other search engines by just deducting their ranking methods, it's parasitic. There's no cleverness about it.

I don't know a lot about this particular approach but your comment that it's just using Google results is blatantly false. It all depends on the search engine that the brave user is leveraging, or no search engine if they type in the URL directly into the header.

Nonsense. This naive idea that Brave innocuously looks at user's traffic patterns.

Google owns 95% of the market in most Western markets. There's no "blatantly false" about that.

They scrape search engine results and present them as their own.

Do 10,000 searches on Google and Brave and you'll see how similar they are. It's as simple as that, scraping by sleight of hand.

Why can't they be a normal search engine - because they need to scrape others. Simples.

Re: The shady world of Brave selling copyrighted data for AI training

#117
post #97

Earlier quoted context omitted.

My reading of the relevant laws would actually lead me to believe that this is not a problem, as long as those reproductions are not returned and the eights holder did not opt out. But courts might decide differently. Regarding the copyright of returned material here is a good discussion: https://copyrightblog.kluweriplaw.com/2023/05/09/generative-...

That's clearly not enough. There's a continuum between producing exact input copy and having genuine creativity because the model actually learned something. A model that just reformats code and changes all the variable names would pass your test and yet be clearly a copyright violation. This whole argument requires that the neural network weights do something creative because they learned from the code instead of ju…

The window of possible actual creativity may be limited and variable.

There are a lot of pretty complex prompts, where if you asked a group of reasonably skilled programmers to implement, they'd produce code that was "reformatted and changed variable names" but otherwise identical. Many of us learned from the same foundational materials, and there are only a handful of non-pathological ways to implement a linked list of integers, for example.

With code it may be more obvious, in that you can't as easily obfuscate things with synonyms and sentence structure changes. Even with prose, there is going to be a tendency to use "conventional" language choices, driving you back towards a familiar-looking mean.

Re: The shady world of Brave selling copyrighted data for AI training

#118

Earlier quoted context omitted.

> A judge willing to commit the Butlerian Jihad[0] might even say that regurgitation does not matter and that all AI outputs are derivative works of the entire training set[1]. A judge can’t “commit” the butlierian jihad. A jihad is a mass event caused by some fraction of the population believing in some cause. Which kinda gets to a point that seems to be missed. Copyright law is not “intrinsic” - nobody thinks that…

Copyright is a unique case in which the law represents a bargain struck in the 1970s that hasn't been updated since. Everyone ignores it because it's nearly impossible to actually enforce copyright on individual infringers. But that doesn't mean copyright is meaningless: any activity which is large enough to be legible [0] to the state will be forced to bend itself to fit within the copyright bargain. And AI training…

Butlerian jihad is a good reference point. Something so bad happened that a large enough portion of the population was convinced to destroy thinking machines, and this no-computer norm was held in human society for a crazy long time (been too long for me to remember how long elapsed before Chapterhouse, which I think is the book where thinking machines start returning). It was a core belief of humanity that computers were bad, not a law imposed by a judge or legislature.

So say a US judge did impose severe restrictions on LLMs through US copyright law. The giant companies that are using LLMs will just move to another country. And just like tax law, others will be happy to have them. Would the US start blocking inbound internet traffic from countries that don’t have the same interpretation of copyright? That seems very unlikely.

The point is that the only way LLMs get the butlerian jihad treatment is if the people rise up against them. Right now, that is nowhere close to happening.

Re: The shady world of Brave selling copyrighted data for AI training

#119
post #53

Earlier quoted context omitted.

The rest of the document is worst. They say they are using your computer to crawl pages you visit and report back to their server. Even Google doesn't do that.

This is opt-in only. https://support.brave.com/hc/en-us/articles/4409406835469-Wh...

If you're shilling for Brave, you should reveal you affiliation with them.

What is it?

Re: The shady world of Brave selling copyrighted data for AI training

#120

> Simply observe the event in which a user does a query q in Brave and then, within one hour, does the same query on a different search engine. What we do is to move the script that detects bad-queries to the browser, run it against the queries that the user does in real-time and then, when all conditions are met, send the following data back to our servers. Wait. Brave browser sends back to Brave Search engine about…

And Google gets the same data joining your cookies ever since Google Plus unified auth across their properties a decade ago. Wait you mean you thought G+ was supposed to compete against Facebook-the-product and not just Facebook-the-ad-network? Oops

Brave is perfectly OK with having oopsies too

Post reply on HN