Live data from Hacker News

AI tooling must be disclosed for contributions

github.com

421–430 of 482 posts

Re: AI tooling must be disclosed for contributions

#421
post #239

Earlier quoted context omitted.

It's not just about how you got there. At least in the United States according to the Copyright Office... materials produced by artificial intelligence are not eligible for copyright. So, yeah, some people want to know for licensing purposes. I don't think that's the case here, but it is yet another reason to require that kind of disclosure... since if you fail to mention that something was made by AI as part of a co…

> if you fail to mention that something was made by AI as part of a compound work you could end up losing copyright over the whole thing The source you linked says the opposite of that: "the inclusion of elements of AI-generated content in a larger human-authored work does not affect the copyrightability of the larger human-authored work as a whole"

This is what you get for skimming. :D

Just to be sure that I wasn't misremembering, I went through part 2 of the report and back to the original memorandum[1] that was sent out before the full report issued. I've included a few choice quotes to illustrate my point:

"These are no longer hypothetical questions, as the Office is already receiving and examining applications for registration that claim copyright in AI-generated material. For example, in 2018 the Office received an application for a visual work that the applicant described as “autonomously created by a computer algorithm running on a machine.” 7 The application was denied because, based on the applicant’s representations in the application, the examiner found that the work contained no human authorship. After a series of administrative appeals, the Office’s Review Board issued a final determination affirming that the work could not be registered because it was made “without any creative contribution from a human actor.”"

"More recently, the Office reviewed a registration for a work containing human-authored elements combined with AI-generated images. In February 2023, the Office concluded that a graphic novel comprised of human-authored text combined with images generated by the AI service Midjourney constituted a copyrightable work, but that the individual images themselves could not be protected by copyright. "

"In the Office’s view, it is well-established that copyright can protect only material that is the product of human creativity. Most fundamentally, the term “author,” which is used in both the Constitution and the Copyright Act, excludes non-humans."

"In the case of works containing AI-generated material, the Office will consider whether the AI contributions are the result of “mechanical reproduction” or instead of an author’s “own original mental conception, to which [the author] gave visible form.” The answer will depend on the circumstances, particularly how the AI tool operates and how it was used to create the final work. This is necessarily a case-by-case inquiry."

"If a work’s traditional elements of authorship were produced by a machine, the work lacks human authorship and the Office will not register it."[1], pgs 2-4

---

On the odd chance that somehow the Copyright Office had reversed itself I then went back to part 2 of the report:

"As the Office affirmed in the Guidance, copyright protection in the United States requires human authorship. This foundational principle is based on the Copyright Clause in the Constitution and the language of the Copyright Act as interpreted by the courts. The Copyright Clause grants Congress the authority to “secur[e] for limited times to authors . . . the exclusive right to their . . . writings.” As the Supreme Court has explained, “the author [of a copyrighted work] is . . . the person who translates an idea into a fixed, tangible expression entitled to copyright protection.”

"No court has recognized copyright in material created by non-humans, and those that have spoken on this issue have rejected the possibility. "

"In most cases, however, humans will be involved in the creation process, and the work will be copyrightable to the extent that their contributions qualify as authorship." -- [2], pgs 15-16

---

TL;DR If you make something with the assistance of AI, you still have to be personally involved and contribute more than just a prompt in order to receive copyright, and then you will receive protection only over such elements of originality and authorship that you are responsible for, not those elements which the AI is responsible for.

--- [1] https://copyright.gov/ai/ai_policy_guidance.pdf [2] https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...

Re: AI tooling must be disclosed for contributions

#422
post #364

Earlier quoted context omitted.

Why would these people disclose their use of AI? These are not responsible and thoughful users of AI. The slop producers won't disclose, and the responsible users who produce high quality PRs with AI will get the "AI slop" label. At this point, why even disclose if the AI-assisted high-quality PR is indistinguishable from having been manually written (which it should be)? No point.

> Why would these people disclose their use of AI? Because lying about your usage of AI is a good way to get completely kicked out of the open source community once caught. That's like asking 'why should you bother with anti-cheating measures for speedruns'. Why should we have any guidelines or regulations if people are going to bypass them? The answer I hope should be very obvious. > high quality PRs with AI will ge…

Getting kicked out from open source community for lying about using AI? Haha, good one.

Re: AI tooling must be disclosed for contributions

#423

Earlier quoted context omitted.

> Have you considered that it is simply singers-performers who like to sing and would like to earn a bit of money from it, but don’t have many original songs if their own? Or, maybe you start to pay attention? They are selling their songs cheaper for TV, radio or ads. > Even pretending they were, if you compare between artists specialising in covers and big tech trying to expropriate IP They're literally working for…

> They are selling their songs cheaper for TV, radio or ads. I guess that somehow refutes the points I made, I just can’t see how. Radio stations, like the aforementioned venue owners, pay the rights organizations a flat annual fee. TV programs do need to license these songs (as unlike simple cover here the use is substantially transformative), but again: 1) it does not rip off songwriters (holder of songwriter right…

I know of restaurants and bars that choose to play cover versions of well-known songs because the costs are so much less.

Re: AI tooling must be disclosed for contributions

#424
post #348

Earlier quoted context omitted.

I’m curious … So “transformative” is not necessarily “derivative”? Seems to me the training of AI is not radically different than compression algorithms building up a dictionary and compressing data. Yet nobody calls JPEG compression “transformative”. Could one do lossy compression over billions of copyrighted images to “train” a dictionary?

A compression algorithm doesn't transform the data it stores it in a different format. Storing a story in a txt file vs word file doesn't transform the data. An llm is looking at the shape of words and ideas over scale and using that to provide answers.

No a compression algorithm does transform the data, particularly lossy ones. The pixels stored in the output are not in the input, they're new pixels. That's why you can't uncompress a jpeg. Its a new image that just happens to look like the original. But it even might not - some jpegs are so deep fried they become their own form of art. This is very popular in meme culture.

The only difference, really, is we know how a JPEG algorithm works. If I wanted to, I could painstakingly make a jpeg by hand. We don't know how LLMs work.

Re: AI tooling must be disclosed for contributions

#425
post #423

Earlier quoted context omitted.

> They are selling their songs cheaper for TV, radio or ads. I guess that somehow refutes the points I made, I just can’t see how. Radio stations, like the aforementioned venue owners, pay the rights organizations a flat annual fee. TV programs do need to license these songs (as unlike simple cover here the use is substantially transformative), but again: 1) it does not rip off songwriters (holder of songwriter right…

I know of restaurants and bars that choose to play cover versions of well-known songs because the costs are so much less.

I really doubt you would ever license any specific songs as a cafe business. You should be able to pay a fixed fee to a PRO and have a blanket license to play almost anything. Is it so expensive in the US, or perhaps they do not know that this is an option? If the former, and those cover artists help those bars keep their expenses low and offer you better experience while charging less—working with the system, without ripping off the original artists who still get paid their royalty—does it seem particularly parasitic?

Re: AI tooling must be disclosed for contributions

#426
post #171

Earlier quoted context omitted.

> Courts (at least in the US) have already ruled that use of ingested data for training is transformative. If you have code that happens to be identical to some else's code or implements someone's proprietary algorithm, you're going to lose in court even if you claim an "AI" gave it to you. AI is training on private Github repos and coughing them up. I've had it regurgitate a very well written piece of code to do a p…

That seems a real stretch. GPT 5 just invented new math for reference. What you are saying would be equivalent to saying that this math was obviously in some paper that mathematician did not know about. Maybe true, but it's a far reach.

New math? As in it just fucking Isaac Newton'd invented calculus? Or do you just mean it solved a math problem?

Re: AI tooling must be disclosed for contributions

#427

Earlier quoted context omitted.

How is that obviously proprietary? Aren't you implicitly assuming that the AI couldn't have written it on its own?

The idea that something that can't handle simple algorithms (e.g. counting the number of times a letter occurs in a word) could magically churn out far more advanced algorithms complete with tests is… well it's a bit of a stretch.

It's terrible at executing algorithms. This, it turns out, is completely disjoint from writing algorithms.

Re: AI tooling must be disclosed for contributions

#428
post #324

Earlier quoted context omitted.

Sure, screen recorders exist. And if contributions are that unwelcome, then it's better not to contribute. There has to be some baseline level of trust that the contributor is trying to do the right thing; I get enough spying from the corporations in my life.

I'm not talking screen recorder but a log file where I could be given it and use it for input and repeat the work exactly as it was done. Sort of like how one could do the same with their bash history file to repeat an initially exploratory analysis effort using the same exact commands. I'm surprised that isn't already a capability given the business interest with AI. One would think they would like to cache these pr…

Does that exist for VSCode?

If not, why would it exist for VSCode + a variety of CLI tools + AI? Anyhow, saving the exact prompt isn't super useful; the response is stochastic.

Re: AI tooling must be disclosed for contributions

#429

Earlier quoted context omitted.

This would be the first time ever that an LLM has discovered new knowledge, but the far reach is that the information does appear in the training data?

They've been doing it for a while. Gemini has also discovered new math and new algorithms. There is an entire research field of scientific discovery using LLMs together with sub-disciplines for the various specialization. LLMs routinely discover new things.

[deleted]

Re: AI tooling must be disclosed for contributions

#430
post #171

Earlier quoted context omitted.

> Courts (at least in the US) have already ruled that use of ingested data for training is transformative. If you have code that happens to be identical to some else's code or implements someone's proprietary algorithm, you're going to lose in court even if you claim an "AI" gave it to you. AI is training on private Github repos and coughing them up. I've had it regurgitate a very well written piece of code to do a p…

That seems a real stretch. GPT 5 just invented new math for reference. What you are saying would be equivalent to saying that this math was obviously in some paper that mathematician did not know about. Maybe true, but it's a far reach.

An example: https://medium.com/@deshmukhpratik931/the-matrix-multiplicat...

Obviously not ChatGPT. But ChatGPT isn't the sharpest stick on the block by a significant margin. It is a mistake to judge what AIs can do based on what ChatGPT does.

Post reply on HN