Live data from Hacker News

Who owns the code Claude Code wrote?

legallayer.substack.com

531–540 of 570 posts

Re: Who owns the code Claude Code wrote?

#532

Earlier quoted context omitted.

> training the LLM in violation of a license Bartz v. Anthropic found that this is fair use, so the license doesn't play into it.

If the trained LLM spits out large, recognizable portions of licensed code and you use it in your product don’t count on that case to keep you from defending yourself in court. The court found in Bartz v. Anthropic that training was fair use. They also found that pirating content to train against was not fair use, and Anthropic paid $1,500,000,000 in a settlement. There are licenses on most software source code. If y…

I find it pretty horrible that a company can pay a mere fine that is a small percentage of its total funding in exchange from materially benefiting from a conspiracy to commit a series of criminal acts.

If Anthropic hadn’t pirated training materials would they even exist? Would they still have been as competitive ?

Would they still have gotten every bit of VC funding in anticipation of future successes derived in part from past crimes?

What’s next ? Armed bank robbery when VC funding dries up?

Re: Who owns the code Claude Code wrote?

#533

I find it distasteful and disturbing that copyright infringement by the people training the LLM in violation of a license is considered contamination by the licensed code. It’s not contamination. The code didn’t seep into your codebase. If the LLM was trained in such a way that portions of code long enough to be protectable then the license was violated by humans. The liability for the problem doesn’t lie on the shou…

Other than putting something into the public domain I don't really know any open source licence that doesn't require at least attribution. One can assume that 99.9% of training data had some sort of license requirements, so just blindly using it is a copyright violation. People just don't seem to care.

Re: Who owns the code Claude Code wrote?

#534

Earlier quoted context omitted.

> training the LLM in violation of a license Bartz v. Anthropic found that this is fair use, so the license doesn't play into it.

If the trained LLM spits out large, recognizable portions of licensed code and you use it in your product don’t count on that case to keep you from defending yourself in court. The court found in Bartz v. Anthropic that training was fair use. They also found that pirating content to train against was not fair use, and Anthropic paid $1,500,000,000 in a settlement. There are licenses on most software source code. If y…

Also fair use is much more limited in the EU. Don't know how it applies here or if there where any rulings. Are you going to stop doing business with the EU (and Japan etc.)?

Re: Who owns the code Claude Code wrote?

#535

Earlier quoted context omitted.

> training the LLM in violation of a license Bartz v. Anthropic found that this is fair use, so the license doesn't play into it.

If the trained LLM spits out large, recognizable portions of licensed code and you use it in your product don’t count on that case to keep you from defending yourself in court. The court found in Bartz v. Anthropic that training was fair use. They also found that pirating content to train against was not fair use, and Anthropic paid $1,500,000,000 in a settlement. There are licenses on most software source code. If y…

The seller of the code has no visibility on the training set of the LLM. If the situation you're describing ends up being illegal, responsibility should fall on the LLM provider to provide tools to detect such overlap with their training sets, and on the clients to run the tools.

The provider of the LLM should want to enable this and to take on that responsibility (I mean take it from the clients), otherwise no one will want to use the tool. Maybe there could be AI tool-use lawsuit insurance, but I feel like that's worse than the copyright infringement detection tool for everyone involved.

I can see the tool happening in the EU, but nowhere else basically, especially in the US, the government sees "AI dominance" as a national priority and a national security priority.

Re: Who owns the code Claude Code wrote?

#536

Earlier quoted context omitted.

Interesting, though, that ownership of the code can still be transferred to the employer. So it's in the public domain (because not human authored) but owned by the employer (because the human and/or LLM was employed by the employer)? I don't really understand how this works.

When you write code by hand, you are the author. As part of your contract with your employer you grant copyright and authorship to your employer by default (as stated in the contract). The LLM is not employed by you or your employer, because you can't enter contracts with non human or non human organizations. When you license a non-LLM code generation service (like a page that creates a website for you), that company…

yeah this is what I understand - that it's the copyright ownership that is being transferred. But if there's no copyright to transfer, what does the company end up owning?

Re: Who owns the code Claude Code wrote?

#537
post #353

Earlier quoted context omitted.

Interesting, though, that ownership of the code can still be transferred to the employer. So it's in the public domain (because not human authored) but owned by the employer (because the human and/or LLM was employed by the employer)? I don't really understand how this works.

Note: IANAL I think what this means is that the employee may not be the copyright owner for multiple reasons, which are possibly applicable simultaneously. It does not imply that the employer owns copyright over the work that is in public domain, which would be a contradiction.

yeah, that makes sense

Re: Who owns the code Claude Code wrote?

#538
post #8

This is all well and good as an intellectual exercise, but in real life none of this matters. Almost no one thinks their code is copyrightable or seriously thinks their code is a moat. I've written the same chunks of code for a number of employers as has every engineer. We've all taken chunks from stack overflow and other places without carefully considering attribution. This comes up in a few places as a kind of vin…

> Almost no one thinks their code is copyrightable or seriously thinks their code is a moat. You'd be surprised! Among non-software management types, they often think of the code as extremely valuable IP and a trade secret. I'm a CTO and I've made comments before to non/less technical peers about how the code (generally speaking) isn't that big of a secret, and I routinely get shocked expressions. In one case the com…

I’ve worked at too many places where I mused that if someone gave the source code to the competitors, it’d likely drive the competitors out of business as they tried to use it.

Keeping it proprietary probably has the greatest value in preserving the company’s reputation…

Re: Who owns the code Claude Code wrote?

#539

Earlier quoted context omitted.

I would have assumed the opposite is true. Do you have any data to back that up?

You would assume that there is more proprietary code available to read on the internet than GPL code? Do you have any rationale for that assumption? Basically all GPL code is available on the web and there is a vast amount of it. I barely see any current non-FOSS code on the internet, although I think it would be fair to count the big projects who have been using pseudo-OSS licenses lately as proprietary. Wouldn't a…

Are you aware of non-GPL FOSS licenses?

Re: Who owns the code Claude Code wrote?

#540
post #329

Earlier quoted context omitted.

Filing isn't the gate, registration is. Copyright Office requires you to disclose AI involvement and disclaim the AI-generated parts. Zarya of the Dawn is the example — applicant filed for the whole graphic novel, got partial registration on the human-written text, refused on the Midjourney images. The reproducibility of the prompt isn't really the test. The test is whether a human made the expressive choices.

Your comments are getting classified by our software as LLM-generated or (more likely) LLM-edited. It's impossible to be certain, of course, but if this is the case—can you please not do this? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 . LLMs are amazing of course and we use them heavily ourselves - but not for modifying tex…

Wow, yes sir! I was using Claude to write faster. But I understand. Thanks for the note.
Post reply on HN