Earlier quoted context omitted.
People still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.
It's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.
The shady world of Brave selling copyrighted data for AI training
81–90 of 127 posts
Re: The shady world of Brave selling copyrighted data for AI training
#82> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…
Unpopular opinion time: A ML model is clearly a derivative work of its input. Here's what I think would be fair: Anyone who holds copyright in something used as part of a training corpus is owed a proportional share of the cash flow resulting from use of the resulting models. (Cash flow, not profits, because it's too easy to use accounting tricks to make profits disappear). In the case of intermediaries (e.g., social…
Do you mean this in a copying sense or a mathematical sense?
What if it's only storing 1 byte per input document?
Re: The shady world of Brave selling copyrighted data for AI training
#83> Fair use is a doctrine in the law of the United States that allows limited use of copyrighted material without requiring permission from the rights holders. It provides for the legal, non-licensed citation or incorporation of copyrighted material in another author's work under a four-factor balancing test: > 1) The purpose and character of the use, including whether such use is of a commercial nature or is for nonp…
The entire fair use claim is derived not from any legal basis, but rather, that "it has to be fair use" because it would be legally catastrophic for OpenAI et al if it weren't true. If you look at the core argument in favour of fair use, it's that "LLMs do not copy the training data", yet this is obviously false. For Github copilot and ChatGPT examples of it reciting large sections of training data are well known. Pl…
The problem is that filtering the training set is naively O(n^2) and n is already extremely large for DALL-E. For LLMs, it's comically huge, plus now you have to do substring search. I've yet to hear OpenAI talk about training set deduplication in the context of LLMs.
As for the legal basis... nobody's ruled on AI training sets in the US. Even the Google Books case that I've heard cited in the past (even by myself) really only talks about searching a large corpus of text. If OpenAI's GPT models were really just a powerful search engine and not intelligent at all, they'd actually be more legally protected.
My money's still on "training is fair use", but that actually doesn't help OpenAI all that much either, because fair use is not transitive. Right now, such a ruling would mean that using AI art is Russian roulette: if your model regurgitates, the outputs are still infringing, even if the model is fair use. Novel outputs aren't entirely safe, though. A judge willing to commit the Butlerian Jihad[0] might even say that regurgitation does not matter and that all AI outputs are derivative works of the entire training set[1].
This logic would also apply in the EU. Last I checked the TDM exception only said training is legal, not that you could sell the outputs. They don't really respect jurisprudence the way the Anglosphere obsesses over "precedent", so copyright exceptions are almost always decided by legislatures and not judges over there, and the likelihood of a judge saying that all outputs are derivative works of the training set regardless of regurgitation is higher.
[0] In the sci-fi novel Dune, the Butlerian Jihad is a galaxy-wide purge of all computer technology for reasons that are surprisingly pertinent to the AI art debate.
Yes, this is also why /r/Dune banned AI art. No, I have not read Dune.
[1] If the opinion was worded poorly this would mean that even human artists taking inspiration to produce legally distinct works would be violating copyright. The idea-expression divide would be entirely overthrown in favor of a dictatorship of the creative proletariat.
[2] "Music and Film Industry Association of America" - an abbreviation coined for an April Fools joke article about the MPAA and RIAA merging together.
Re: The shady world of Brave selling copyrighted data for AI training
#84Earlier quoted context omitted.
Correct. I've been using Brave since their very first versions on the desktop, and there never was any in-page ad insertion. The one type of in-page modification they used to do is that they would add a "tip" button to the content creator of some social networks like Twitter or reddit. That had nothing to do with "replacing ads" though. > replaces an ad, they put the new ad in a popup Incorrect. There is no 1:1 repla…
You're responding to a comment that gave you a link to their inital plan, which was literally replacing the ads. click on it, your horizon might be broadened by the added knowledge.
lalaland1125 is making claims about what they actually did, and those claims are not correct.
Re: The shady world of Brave selling copyrighted data for AI training
#85Earlier quoted context omitted.
People still like to defend Brave when it gets caught on shady things over and over again. I guess there are no too many other options. For some people it is already too difficult to install uBlock or know its existence.
It's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.
I used to recommend Firefox, but Mozilla has totally jumped the shark (privacy violations [multiple], wastes too much money, blocks APIs that are useful with no real security risks while approving APIs with little use that do have security risks, etc, very user hostile).
Chromium is obviously not trustworthy at this point, let alone Chrome. So that leaves like, Safari and Opera?
Brendan Eich is the CEO of Brave, and I trust him. Mozilla was good until he was ousted for political reasons.
Re: The shady world of Brave selling copyrighted data for AI training
#86Earlier quoted context omitted.
It actually doesn’t even matter if LLMs reproduce copyrighted data from their training. The issue is that a human copied the data from its source into memory for use in training, and this copy was likely not fair use under cases like MAI Systems . The Supreme Court hasn’t ruled on a software case like this, as far as I know. But given the recent 7-2 decision against Andy Warhol’s estate for his copying of photographs…
So how is that supposed to work with people sending it legally obtained copyrighted materials for an analyze?
“Write a review of this short story: …” – probably fine.
“Rewrite this short story to have a happier ending: …” – probably not.
Re: The shady world of Brave selling copyrighted data for AI training
#87Earlier quoted context omitted.
Do you think a human learning something from reading is fair use? Or are we all copyright violators because reading that article altered our connectomes, and we may recall parts of it later?
The point being raised is quite specific. Not sure if you’re willingly ignoring it or what? The answer is no, because you reading the article didn’t dramatically degrade its market value. An AI ingesting all content on the internet and then being ultra-effective at frontrunning that content for a large number of future readers does degrade its market value (and subsumes it into the model’s value).
How about if you read a news article to write a competing one rewording and possibly citing it (one of the most common practices in news)?
Re: The shady world of Brave selling copyrighted data for AI training
#88The websites a Brave user browses are anonymously relayed to their servers for indexing/training. So, they crawl the web without a crawler and the website operators can't do anything about it. That's genius!
Re: The shady world of Brave selling copyrighted data for AI training
#89Earlier quoted context omitted.
It's because a lot of people are bought into BAT (Brave's cryptocurrency) and have a strong financial incentive to shill Brave.
I do not use BAT or any crypto. Brave just works, and it blocks ads automatically when I tell friends to install it on their computers. I used to recommend Firefox, but Mozilla has totally jumped the shark (privacy violations [multiple], wastes too much money, blocks APIs that are useful with no real security risks while approving APIs with little use that do have security risks, etc, very user hostile). Chromium is…
Brave is like 99% of Chromium + uBlock…
Re: The shady world of Brave selling copyrighted data for AI training
#90Earlier quoted context omitted.
I don't know what a fair settlement would be but I'm looking forward to a copyright-holder suing OpenAI to obtain one. These companies have no value if copyright can be enforced on their training data.
I think there are ways around it. The simplest would be to generate replacement data, for example by paraphrasing the original, or summarising, or turning it into question-answer pairs. In this new format it can serve as training data for a clean LLM. Of course the public domain data would be used directly, no need to go synthetic there. An important direction would be to train copyright attribution models, and diff-…