Earlier quoted context omitted.
> automatically rephrased an original text - something like the Encyclopaedia Britannica - to preserve the meaning but not have identical phrasing Note that it's very hard to do this starting from a single source, because in order to be safe from any copyright concern you'd have to only preserve the bare "idea" and everything else in your text must be independent. But LLM's seem to be able to get around this by looki…
The concept clustering across multiple sources allows you to rephrase more accurately while retaining meaning - however the point I'm making is if you then point that program at Encyclopaedia Britannica and simply rephrase it then charge for access to the rephrased version - should you be allowed to do that?
Thomson Reuters wins first major AI copyright case in the US
181–188 of 188 posts
Re: Thomson Reuters wins first major AI copyright case in the US
#182Earlier quoted context omitted.
> Yes, they do To clarify: veggieroll said training models wouldn't be viable, you said it'd just require licensing like everyone else already manages, I said most other cases don't use millions/billions of works, you're saying that yes they do? I feel like there must be a misunderstanding here, because that doesn't make much sense to me. Even for making a movie, which I think would be the most onerous of traditional…
> I said most other cases don't use millions/billions of works, you're saying that yes they do? I assumed we were talking about logistics, not tech. I'm sure it will be technically possibly to use less training data overtime (Deepseek is more or less demonstrating that in real time. Maybe there's copyright data but I'd be surprised if it used anything close to 80 TB like competittorz). I know hindsight is 20/20, but…
Still uncertain what you mean - the logistics of creating something? Logistics as in transporting goods? Either way I think veggieroll's point on viability still stands.
> Deepseek is more or less demonstrating that in real time. Maybe there's copyright data but I'd be surprised if it used anything close to 80 TB like competittorz
* GPT-4 is reported to have been trained on 13 trillion tokens total - which is counting two passes over a dataset of 6 trillion tokens[0]
* DeepSeek-V3, the previous model that DeepSeek-R1 was fine-tuned from, is reported to have been pre-trained on a dataset of 14.8 trillion tokens[1]
Can't find any licensing deals DeepSeek have made, so vast majority of that will almost certainly be unlicensed data - possibly from CommonCrawl and shadow libraries.
[0]: https://patmcguinness.substack.com/p/gpt-4-details-revealed
[1]: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/DeepSee...
> > > Let's not pretend these companies can't do this through normal channels.
> > I'm not sure that there really has been a normal channel [...]
> There isn't.
Then, surely it's not just pretending?
A while back, as a side project, I'd had a go at making a tool to describe photos for visually impaired users. I contacted Getty to see if I could license images for model training, and was told directly that they don't license images for machine learning. Particuarly given that I'm not massive company, I just don't think there really are any viable paths at the moment except for using web-scraped datasets.
> So they'd need to do it the old fashioned way with agreements .
I'm sceptical of whether even the largest companies would be able to get sufficient data for pre-training models like LLMs from only explicit licensing agreements.
> I don't exactly pity their herculean effort. Those same companies spend decades suing individuals for much pettier uses and building those precedent up (some covered under free use).
I feel you're conflating two groups: model developers that have previously been (on average) supportive of fair-use, and media companies (such as the ones currently launching lawsuits against model training) that lobbied for stronger copyright law. Both are acting in self-interest, but I'd disagree with the idea that there was any significant switching of sides on the topic of copyright.
> Content creators now need to take extra precautions so they aren't stolen from because they don't even bother trying to respect robots.txt.
The major US players claim to respect robots.txt[2][3][4], as does CommonCrawl[5] which is what the smaller players are likely to use.
You can verify that CommonCrawl respects robots.txt by downloading it yourself and checking.
If OpenAI/etc. are lying, it should be possible for essentially anyone hosting a website to prove it by showing access from one of the IPs they use for scraping[6]. (I say IPs rather than useragent string because anyone can set their useragent string to anything they want, and it's common for malicious/poorly-behaved actors to pretend to be a browser or more common bot).
[2]: https://platform.openai.com/docs/bots
[3]: https://support.anthropic.com/en/articles/8896518-does-anthr...
[4]: https://blog.google/technology/ai/an-update-on-web-publisher...
[5]: https://commoncrawl.org/faq
[6]: https://openai.com/gptbot.json
> Was all that velocity worth it? Who benefitted from this outside of a few billipnaires? We can't even say we beat China on this.
There's been a large range of beneficial uses for machine learning: language translation, video transcription, material/product defect detection, weather forecasting/early warning systems, OCR, spam filtering, protein folding, tumor segmentation, drug discovery and interaction prediction, etc.
I think this mainly comes back to my point that large-scale pretraining is not just for LLM chatbots. If you want to see the full impact, you can't just have tunnel-vision on the most currently-hyped product of the largest companies.
> Humans inherit their data and slowly structure around that. Maybe if AI models collaborated together as humanity did, I would sympathize more with this argument.
Machine learning in general (not "OpenAI") is a fairly open and collaborative field. Source code for training/testing is commonly available to use and improve; papers documenting algorithms, benchmarks, and experiments are freely available; arXiv (Cornell University's open-access preprint repository) is the place for AI papers, opposed to paywalled journals; and it's very common to fine-tune someone's existing pretrained model to perform a new task (transfer learning) opposed to training from scratch.
I'd attribute a lot of the field's success to building off each others' work in this way. In other industries, new concepts like transformers or low-rank-adaptation might still be languishing under a patent instead of having been integrated and improved on by countless other groups.
> AI can evolve organically but it instead devolved into a thieve's den.
Unclear what you mean by organically - evolution still needs data.
Re: Thomson Reuters wins first major AI copyright case in the US
#183Earlier quoted context omitted.
The intention of copyright is to protect useful work. The detail of how to do that in fair way that doesn't block other people is complex[1] - you can never cover all possibilities in a written law - that's why you have people interpreting them and making judgements . All I'm saying is the guiding light in that interpretation is copyright is there to protect the justifiable work of people in a fair way. Somebody taki…
> The intention of copyright is ..."to promote the progress of science and useful arts". I don't see anything in there about rewarding 'work' irrespective of whether that work involves any kind of creativity. > If those notes could have been created mechanically directly from the original source - why didn't the copier do that That's actually a very good question. In practice, I do absolutely agree that the notes inv…
Not sure where you got that quote from, but I'd say the work aspect is implicit in the "promote the progress" - ie progress requires that people are able to get paid in their work to progress science or the useful arts.
If the progress was trivial and required no work then it wouldn't need protection or promotion.
And sure it's phrased that way to get the balance between fair use and protection - but if there was no need of protection then copyright wouldn't need to exist - as free reuse is the default.
Re: Thomson Reuters wins first major AI copyright case in the US
#184Earlier quoted context omitted.
My biggest concern is, what happens when countries like China, who aren't restricted by this, far outpace western countries in this technology? Do we just shrug and accept our far inferior models? LLMs are a productivity multiplier (similar to a search engine), so it'll have a large impact on the economy if licensing costs prohibit large scale training.
"Our" models are not inferior. There is plenty of data, and the next frontier is prediction-time compute and data synthesis. Shouldn't the Chinese worry that they are depressing the publication of commercial IP?
Re: Thomson Reuters wins first major AI copyright case in the US
#185Re: Thomson Reuters wins first major AI copyright case in the US
#186Re: Thomson Reuters wins first major AI copyright case in the US
#187Earlier quoted context omitted.
The concept clustering across multiple sources allows you to rephrase more accurately while retaining meaning - however the point I'm making is if you then point that program at Encyclopaedia Britannica and simply rephrase it then charge for access to the rephrased version - should you be allowed to do that?
The underlying problem is that "meaning" in the ordinary sense still includes plenty of copyrightable elements. If you point a typical LLM program at some arbitrary text and tell it to "rephrase" that, you'll generally end up with a very close paraphrase that still leaves intact to a huge extent the "structure, sequence and organization" (in a loose sense) of the original. So it turns out that you're still in breach…
ie the way to avoid copyright is to double down on the copying?
I can see how, for a human, you could argue that there is creativity in splicing those bits together into a good whole - however if that process is automated - is it still creative - or just automated theft?
Re: Thomson Reuters wins first major AI copyright case in the US
#188Earlier quoted context omitted.
Add to this, the brain is constantly processing raw sensory data from the moment it became viable, even when the body is "sleeping". It's using orders of magnitude more data than any model in existence every moment, but isn't generally deemed "intelligent" enough until it's around 18 years old.
It’s unlikely that sensory data contributes to cognitive ability in humans. People with sensory impairments, such as blind people, are not less cognitively capable than people without sensory impairments. Think of Helen Keller, who, despite taking in far less sensory information than the average person, was still more intelligent than average.
* Preprocessed since the data is actually of 1D streams of characters, and not 2D colour points (as with vision models).