Live data from Hacker News

Japan’s government will not enforce copyrights on data used in AI training

technomancers.ai

361–370 of 426 posts

Re: Japan’s government will not enforce copyrights on data used in AI training

#361
post #357

Earlier quoted context omitted.

The Japanese article does explicitly state it if you run it through a translator, and also this is from May 11 > Additionally, the group raised other issues that Article 30-4 of the Copyright Law, which permits the use of a copyrighted work for machine learning, does not include procedures for gaining permission in advance from copyright holders. The article permits the use of copyrighted material such as text and im…

From what I think is the original source: > まずAIによる情報解析についての我が国の法制度(著作権法)について確認したところ、我が国において、非営利目的であろうと、営利目的であろうと、複製以外の行為であろうと、違法サイトなどから取得したコンテンツであろうと、方法を問わず情報解析のための作品利用はできると永岡大臣が明言しました。 > Confirming the legal system (copyright law) wrt. data analysis by AI in our country, Minister Nagaoka clearly stated that in our country, whether for non-profit purposes or for profit purposes, whether an act other than reproductio…

Well, its time to use the many software source leaks out here to create an even more powerful copilot.

Re: Japan’s government will not enforce copyrights on data used in AI training

#362
post #138

I think this should generally be true. The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. It may produce stuff that violates copyright, but the way you use or distribute the product of the model that can violate copyright. Making it write code that’s a clone of copyright code or making it make pictures with copy right imagery in it or…

> The aggregation performed by model training is highly lossy and the model itself is a derived work at worst and is certainly fair use. Lossy or not, the training data provides value . If all the various someones had not spent time making all the stuff that ends up as training data, then the model it trains would not exist. If you are going to use someone else's work in order to make something that you are going to…

> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated

This is insane. I learned mathematics from Houghton-Mifflin and Addison-Wesley copyrighted textbooks. Are you telling me that if I use this knowledge to, say, calculate matrices, I am violating copyright?

If I read Brandon Sanderson, absorb his writing style, setting, characters, and plot devices into my subconscious, and then a year later pen a short story that happens to be Sanderson-ish, do I have to write a check to Brandon?

No. LLMs are doing roughly the same thing our own brains do, at least at a high level: absorb information, and utilize that information in different ways. That's it.

Re: Japan’s government will not enforce copyrights on data used in AI training

#363
post #152

Earlier quoted context omitted.

Which I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.

I've heard this FAX meme since like early 2000, but I have yet to encounter one in my 4 years living here, and most Japanese people I joke this too just gets as perplexed by the joke. I wonder, where do people find those FAX machine? Have I lived in a tech/startup bubble in Tokyo and missed it? I didn't even see it at the local Ward office in the suburb. Some stuff are still old-school (hanko etc) but it to me seem l…

Have you seen the big printers at many konbinis? Those are also fax machines.

I have only sent two faxes in my life, and both were after I moved to Japan. The first was right after I arrived, I ordered something from Amazon and I realized I had written my address wrong, when I tried to fix it my account to blocked, and to unblock it I had to send a hand-written fax with my name and address. The second time was when I got a letter from the Sapporo police telling me that I had lost my driving license there. Again, to recover it I had to send a hand-written fax explaining the situation, so they could prove my identity. How that is considered a secure procedure is beyond me, but such is Japan.

Re: Japan’s government will not enforce copyrights on data used in AI training

#364
post #360
post #355

Earlier quoted context omitted.

LLM training is a compression algorithm, LLM weights are a compressed dataset of the training materials, and executing LLMs is accessing the compressed source material. That the compression technology relies on parameterizing the copyrighted material, and as a result can produce hallucinations remixing the copyrighted material, is super cool but doesn't change that this is compression at rest. All the hullabaloo abou…

Can I talk to gzip in natural language and have it produce novel output not contained within any of its source files? If not, I think your comparison is deeply flawed. LLM's are not simply compression algorithms.

Maybe if someone were to build it. You can't talk to LLMs in natural language either. They have a very precise query language. The natural language component is an additional feature bolted onto the front.

Also the output of LLMs is not novel, it's a derivative work of the training dataset. LLMs can't produce anything not present in the training set. They are operating on parameterized classifications of artwork instead of the artwork itself, which is the new technology; but they can't do anything except combine those existing Legos in new ways.

We also talk to databases in what was considered 'natural language' for that era - SQL. The LLMs are no different, which is why we have prompt engineering same as we have database engineering. Just because the database is processing the dataset and storing it in a proprietary way, approachable via query language, does not change the fact that the original data is still in there, and everything that comes out is derivative.

Re: Japan’s government will not enforce copyrights on data used in AI training

#365
post #229

Earlier quoted context omitted.

> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way. > Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry. Is it really new? Hum…

>> Is it really new? Humans have always learnt by studying what's out there already. "Humans" being the important word here. I don't understand why people keep trying to compare training a model to humans learning through reading etc. They are very different things. Learning done by machines at enormous scale and done to benefit private companies financially is not the same as humans learning.

> They are very different things

Yes, but also very similar. We learn very well by spaced repetition, and by practicing things. Our whole nervous system stores information in a similar way. Is the brain and the individual neurons more complex? Yes, sure, but that doesn't negate the core similarities.

> Learning done by machines at enormous scale and done to benefit private companies financially is not the same as humans learning.

Yes, that's the important difference. That in the end if you train a robot you get a program that's easy to copy/scale, the marginal cost of using it is orders of magnitude lower than what you'd get if you would do it with humans in the loop.

We already have fair use in copyright, because there are important differences between the various forms and modalities of human imitation.

And of course maybe it's time to rename copyright to usageright. After current copyright doesn't even apply in most cases. (The results are not derivatives, there's sufficient substantial transformative difference, etc. That said, the phrasing in the US constitution still makes sense: "... the exclusive Right to their respective Writings and ..." ... if we interpret Right to mean all rights, including even the right of who can read/see it.)

Re: Japan’s government will not enforce copyrights on data used in AI training

#366

This article is an example of emerging AI-bro tactics that completely mirrors crypto-bro tactics: they pick any piece of news and reinterpret it to fit an agenda. While the article is in English, the link to source is in Japanese. The only external source I found suggests the discussion is about promoting open data and open science from research institutions [1] [1] https://asianews.network/japan-to-promote-use-of-ge…

There are no such thing as "AI bros", AI training costs millions of dollars. It's a very different field than cryptocurrency where mining could be done by individuals in the early days.

You see corporate tactics being deployed in a wide variety of ways to secure a social, legal, and economic moats, since apparently they have limited technology only based moat. It's similar to how google files various amicus briefs against copyrights.

Google v. Oracle (2020): This was a landmark case in which Google was a party. The case revolved around whether copyright protection extended to a software interface. Google argued that APIs, which allow different software programs to communicate with each other, should not be subject to copyright. Google filed an amicus brief in its own case, arguing that a decision in favor of Oracle would stifle innovation in the tech industry.

Authors Guild v. Google (2015): This case involved Google's project to digitize millions of books from public and university libraries. The Authors Guild argued that Google was infringing on authors' copyrights. Google argued that its use of the books was fair use because it only showed snippets of the books in search results. Google filed an amicus brief in this case as well.

Aereo Case (2014): Google filed an amicus brief in support of Aereo, a company that provided a service for streaming broadcast television over the internet. Broadcasters sued Aereo, arguing that the service infringed on their copyrights. Google argued that a ruling against Aereo could have broad implications for cloud storage services.

Viacom v. YouTube (2012): In this case, Viacom sued YouTube, which is owned by Google, for copyright infringement. Google filed an amicus brief arguing that YouTube was protected by the safe harbor provisions of the Digital Millennium Copyright Act (DMCA), which protect service providers from liability for user-generated content.

Big Tech is often against Copyrights because they believe they have the economic moat to beat other players.

Re: Japan’s government will not enforce copyrights on data used in AI training

#367
post #268

Earlier quoted context omitted.

I don't think the comparison with human learning holds. NNs and humans don't learn the same way - humans can fairly quickly generalise what they have learned and, most importantly, go beyond what they've learned. I haven't see that happen with neural networks or GPTs; at best, you're getting the average of what it has 'learned'. There's human learning and there's neural network 'learning' and they're a different thin…

NN's absolute can go beyond what they have learned and aren't just producing the "average". Some good examples outside the typical LLM/images work: * Deep Mind's work on AlphaFold, which generates predictions on proteins that haven't been seen before * AlphaGo which plays games better than any human (so clearly can't be "the average") If we look at LLMs, something like writing code in the style of Shakespear isn't re…

> Deep Mind's work on AlphaFold, which generates predictions on proteins that haven't been seen before

I have used AlphaFold a bit in my own work, and if I showed it 'unusual' proteins like rare mutants it usually generated garbage. Some evidence for this exists in the literature; see for example https://www.biorxiv.org/content/10.1101/2021.09.19.460937v1 or https://academic.oup.com/bioinformatics/article/38/7/1881/65... or

>AlphaFold recognizes a 3D structure of the examined amino acid sequence by a similarity of this sequence (or its parts) to related sequences with already known 3D structures

https://www.biorxiv.org/content/10.1101/2022.11.21.517308v1

Re: Japan’s government will not enforce copyrights on data used in AI training

#368
post #152

Earlier quoted context omitted.

Which I find bizarre given how backwards Japan is in the adoption of other technologies. Eg their continued reliance on paper records and fax machines.

I would not say we are super reliant on paper/fax anymore. But, it is still quite common. I did receive a Fax at the office last week on Thursday. Oh, and some of our stuff uses Dialup, usually in relation to that older infrastructure. Got thrown for a loop this wednesday when ye good old Dialup (acoustic handshake) audio started screeching across the office. Did not realize we still used it at all!

I've sent faxes in America within the last 2 years; they're still a thing here too.

Re: Japan’s government will not enforce copyrights on data used in AI training

#369
post #358
post #290

Earlier quoted context omitted.

> Lossy or not, the training data provides value. If we ignore the issue of machine learning for now; It's not the job of copyright to prevent people extracting value from a copyrighted work. If it was, then it would be possible for copyright holders to launch lawsuits that block entities from using the knowledge that was published in copyrighted reference material. Or the rights holder of a cookbook would be able to…

The purpose of copyright is to create moral and economical incentives to authors to create new copyrightable works, but giving those authors a time limited state granted monopoly. Reproduction is only one aspect to this. As an example, applying a song to a video require addition permissions even if the party has permissions for reproduction. The owners to a record can also disallow the use of a song in a political ev…

Your examples are not inherent rights that copyright law explictly grants to rights holders. They are clever side effect of how the holder licenses out their monopoly on reproduction.

Holders rarely grant unrestricted reproduction rights to anyone. Reproduction rights licenses always come with a bunch of explicit restrictions, for example: "You may reproduce this novel, in print, unmodified, only for retail sale, in North America, on this quality of paper, for the next 5 years" and so on.

The party has the license to reproduce the song as a standalone audio recording, but attaching it to a video and reproducing the combined work isn't covered and the party must enter into negotiations with the rights holder for a new license. Such licenses often only grant the rights to reproduce it with that exact video and not a different one later, which allows the rights holder to gain control over which videos their song is attached to.

Same thing with holding morality over political events. The rights holder was careful to add a bunch of restrictions to that public performance license they sell. Sure, the politician might have bought a licence, but they forgot to check the small print that blocks their type of event from actually using it.

-----

Machine learning is kind of like compression, yes... It can be a useful analogy at times.

But it is absolutely nothing like lossy video compression. It's not compressing a single file or object. The only way you could train it on a 4k video and get a 420p video out is if that model was extremely over-fitted. The resulting model would likely be bigger than a 420p h264 video file and useless for anything else.

The way that machine learning is like a compression is that it find common patterns across it's entire training set and merges them in very lossy ways.

And it's actually very much like how a human brain works. Your brain doesn't start from scratch for every single human face you recognise. Instead, your brain has built up a generic understanding of the average human face, grouping by clusters of features. Then to remember a given human face it just remembers which cluster of features it's close to and then how it differs... Which is a form of lossy compression.

> No person can produce a 420p video just by consuming a 4k video

But many people do remember entire songs, complete with lyrics and music. And people with musical skills can (and often do) reproduce that song from memory as a cover... Which is copyright infringement if preformed publicly or otherwise distributed.

> nor can any machine learning model gain the emotional constructs and social contexts that human brains get from learning.

There are many things which large LLMs like chatgpt are absolutely incapable of doing. People do over hype their capabilities.

But in my experiments, chatgpt is actually quite good at tasks that require interpreting emotions and social contexts. Does it actually understand these emotions and social contexts? shrug. But if it doesn't actually understand that just proves that true understanding isn't actually needed to preform useful tasks in those areas.

Re: Japan’s government will not enforce copyrights on data used in AI training

#370
post #229

Earlier quoted context omitted.

> If you are going to use someone else's work in order to make something that you are going to profit off of, I believe that original author should be compensated. And should also be able to decide they don't want their work used in that way. > Note that I'm not talking about what existing copyright law says; I'm talking about how I believe we should be regulating this new facet of the industry. Is it really new? Hum…

>> Is it really new? Humans have always learnt by studying what's out there already. "Humans" being the important word here. I don't understand why people keep trying to compare training a model to humans learning through reading etc. They are very different things. Learning done by machines at enormous scale and done to benefit private companies financially is not the same as humans learning.

How is it meaningfully different with respect to this question?

If I go to a museum and look at a bunch of modern paintings, then go home and paint something new but “in the style of”, this is well-established as within my rights, regardless of how any of the painters whose work I studied and was inspired by might feel.

If I take a notebook and write down some notes about the themes and stylistic attributes of what I see, then go home and paint something in the same style, that too is fine - right? Or would you argue the notes I took are a copyright violation? Or the works I made using those notes?

Now let’s say I automate the process of recording those notes. Does that change the fundamentals of what is happening, with respect to copyright?

Personally, I don’t think so.

Post reply on HN