Live data from Hacker News

GCC steering committee announces AI policy

lwn.net

411–420 of 453 posts

Re: GCC steering committee announces AI policy

#411

Earlier quoted context omitted.

Of course the trained models, being derived from GPL code, become derivative work, and by extension - also GPL? Right?

No, you can't copyright an idea, only an expression of an idea, and LLMs operate at the level of ideas. They don't literally stitch together code from training.

Well, not exactly. A LLM is still a computer, doesn't have an intelligence (beside being called AI). That means that their output is a mere computation of their input data, and their input data it's the stuff that was used for the training.

If you imagine it as a "box" you feed into it material and a prompt and it spits out the same material rearranged to do what you did ask for. It does nothing more than a permutation of their input data, as does any computer program, of course in extremely complex and obscure way, but if you reason it abstractly it's the same things Turing theorized almost a century years ago, input -> BOX -> output.

So *of course* the output *is* a derived work of the input, and thus a GPL code should not really used as a training set.

Re: GCC steering committee announces AI policy

#412

Earlier quoted context omitted.

Code is communication. Open source is community. Engineering is understanding systems in depth. LLMs fundamentally hinder all three. I'm not sure you even need to look much further than that.

I disagree, and quote here Linus' recent comment: > In the kernel community we do open source because it results in better technology, not because of religious reasons. > And so we make decisions primarily based on technical merit. Not fear of new tools. You are not making an argument based on technical merit here.

I am making an argument based on the fundamental nature of the tool, which is to make stochastic, unvetted decisions our behalf — categorically different from everything else in our toolbox. The technical and social repercussions to this are potentially vast: no other tool comes close. So it is actually Linus who is making an unconvincing appeal to emotion (and authority) by pretending that LLMs resemble hammers.

Re: GCC steering committee announces AI policy

#413
post #382

Earlier quoted context omitted.

That specifically states “works created solely by machines” so probably doesn't cover AI-aided work? Though where you draw the line there is likely to be something that'll tax legal budgets for years to come…

From the 2025 Report on Copyright and Artificial Intelligence [1]: > The Office concludes that, given current generally available technology, prompts alone do not provide sufficient human control to make users of an AI system the authors of the output. Prompts essentially function as instructions that convey unprotectible ideas. While highly detailed prompts could contain the user’s desired expressive elements, at pr…

Again, that seems to be talking about fully AI generated work, like “works created solely by machines” from the previous document, not AI-aided work. It is only stating that the prompt is not considered sufficient human work because it comes before the generation process, but it says nothing of subsequent editing of combining with other stuff.

Re: GCC steering committee announces AI policy

#414

Earlier quoted context omitted.

I don't think the concerns are far fetched at all. Look at how image models spit out copyrighted stuff all the time. Midjourney has a bizarre EULA clause that if you use it to generate images that violate copyright and _they_ get sued, they can hold you liable downstream. Which is wild to me -- just don't train on things you don't own and this is not a problem! In images it's much more _obvious_, but I think code is…

Anthropic and co basically have the opposite policy (for paid users): if you get sued for copyright infringement, they will indemnify you. That means they're confident it's not an issue. By the way, do you have a source on Midjourney spitting out copyrighted stuff all the time? Does it happen at random or when users intentionally steer the prompt in that direction? I suspect it's the latter but I admit I'm not really…

These companies have already shown they're extremely reckless with copyright (Anthropic was penalized 1.5B for books, for example). I don't think they've earned that trust of "they looked at it so it must be ok"

Re: GCC steering committee announces AI policy

#415

Earlier quoted context omitted.

The thing they want is clear legal standing. If your submission is majority LLM generated, you can't really confirm that "your" code isn't a copy of something in the training set that has an incompatible license (or a clear copyright violation). That might not be a big deal for some projects, but being this is GNU and their entire identity is centered around free software and licensing, it's kind of a big deal to the…

That's just not an issue in practice. All major AI labs match generated code against their training data to prevent outputting verbatim copies. They're so confident they won't output copyrighted code that they offer copyright indemnification to paid users.

And yet Anthropic was penalized 1.5B. "Trust me bro" does not work with companies that pretty much show they can't be trusted on the daily.

Re: GCC steering committee announces AI policy

#416

Earlier quoted context omitted.

This is such a ridiculous take I hear all the time. The shitty ethics of some people do not represent the ethics of an entire profession. Most programmers are not automating people out of jobs. I've been professionally employed as a software developer for 20 years and I don't think a single thing I've ever written has ever replaced anyone's job. Hopefully sometimes it improves their job.

> I don't think a single thing I've ever written has ever replaced anyone's job. Doubt. Whatever it is your app does, I bet corporations would have been forced to hire more people to do it for them, were it not for you. And there's absolutely nothing "unethical" about it either. Toil is meant to be automated away. What I can't take is programmers thinking they're somehow above this. The only crime here is stopping be…

> The only crime here is stopping before AI replaces the CEOs and politicians. It should keep happening relentlessly until capitalism itself collapses and a post scarcity society is achieved.

Dude, this is why AI boosters scare the shit out of me. We're not moving towards a utopia, we're moving towards a dystopia. Leave me out of your shitty cult.

Re: GCC steering committee announces AI policy

#417

Earlier quoted context omitted.

How is this encouraging? They're sticking their heads in the sand and dooming themselves to irrelevance. All but the most strongly and wrongly ideologically motivated will contribute to other projects like LLVM when GCC asks them to code with rocks and sticks instead of taking advantage of arguably the most important invention in human history.

> arguably the most important invention in human history. Don't be ridiculous.

It's an invention that can invent things on its own, improve itself, do research, etc. The importance of such a thing should be obvious. The hockey stick that's coming for human progress overall is going to make the industrial revolution look like a flatline.

Re: GCC steering committee announces AI policy

#418
post #382

Earlier quoted context omitted.

From the 2025 Report on Copyright and Artificial Intelligence [1]: > The Office concludes that, given current generally available technology, prompts alone do not provide sufficient human control to make users of an AI system the authors of the output. Prompts essentially function as instructions that convey unprotectible ideas. While highly detailed prompts could contain the user’s desired expressive elements, at pr…

Again, that seems to be talking about fully AI generated work, like “works created solely by machines” from the previous document, not AI-aided work. It is only stating that the prompt is not considered sufficient human work because it comes before the generation process, but it says nothing of subsequent editing of combining with other stuff.

In the absence of more specific legislation or court decisions, the same rules that cover usage of any other public domain code would apply.

Re: GCC steering committee announces AI policy

#419

Earlier quoted context omitted.

That went over my head, but maybe it was meant to.

Think b) was meant to be don't eat at Chick-fil-A. IOW if you try to avoid giving money to bad people then you shouldn't cause someone to waste AI tokens. Otherwise you can.

Yes this. I was trying to highlight that wasting someone's tokens still makes AI companies money so if your goal is to spite some individual using AI to open a PR and doing things out of spite is ethical to you then go ahead. If your goal is to resist the proliferation of AI then you're probably shooting yourself in the foot.

Re: GCC steering committee announces AI policy

#420
post #377

Earlier quoted context omitted.

For that very reason they would also add very little LLM code over the next few years (until we get more court rulings) even if they allowed it. A little LLM code would not be usable without the rest of their code which remains covered by copyright. Its much the same as someone creating a fork of GPL code in which they make additions that they put in the public domain. All the original code and the fork as a whole wo…

You want to prevent the transition from a GPL codebase with some public domain code to a public domain codebase with some GPL code. One way to do so is to outright ban contributions leveraging tools that are able to generate public domain code at superhuman speeds. As you pointed out yourself, there's always the option to create a fork that does allow AI contributions, which may eventually force a re-assessment of th…

> You want to prevent the transition from a GPL codebase with some public domain code to a public domain codebase with some GPL code.

That would take a very long time if contributions are reviewed etc. By then any legal ambiguities would be clear.

> a public domain codebase with some GPL code

which would still be a GPL codebase

> One way to do so is to outright ban contributions leveraging tools that are able to generate public domain code at superhuman speeds.

Can they generate code that would pass the quality standards, and pass the processes, of a project like this at superhuman speed? There is a separate requirement that contributors must be able to understand code and answer questions about it so a human would have to review code before even trying to contribute it.

> As you pointed out yourself, there's always the option to create a fork that does allow AI contributions, which may eventually force a re-assessment of the policy if the gap in utility grows too large.

1. if you are right that LLMs will do well enough to create a huge gap, then that is inevitable. 2. if you are wrong about that then it is unnecessary to try to stop it.

Post reply on HN