Perhaps the most telling portion of their decision is: Quality concerns. Popular LLMs are really great at generating plausibly looking, but meaningless content. They are capable of providing good assistance if you are careful enough, but we can't really rely on that. At this point, they pose both the risk of lowering the quality of Gentoo projects, and of requiring an unfair human effort from developers and users to…
I've been using AI to contribute to LLVM, which has a liberal policy. The code is of terrible quality and I am at 100+ comments on my latest PR. That being said, my latest PR is my second-ever to LLVM and is an entire linter check. I am learning far more about compilers at a much faster pace than if I took the "normal route" of tiny bugfixes. I also try to do review passes on my own code before asking for code review…
Gentoo AI Policy
191–200 of 206 posts
Re: Gentoo AI Policy
#192Earlier quoted context omitted.
I get why water use is the sort of nonsense that spreads around mainstream social media, but it baffles me how a whole council of nerds would pass a vote on a policy that includes that line.
To be completely fair, AI really does use more water than other typical compute tasks, because AI takes A LOT of compute. No, it's not like email, or a web server. I can run an email server or apache on my rinky dink computer and get hundreds of requests per second. I can't run chatgpt, that requires a super computer. And of the stuff I can run, like deepseek, I'm getting very few tokens/s. Not requests! Tokens! Yes,…
As to actual numbers, they're not that hard to crunch, but we have a few good sources that have done so for us.
Simple first-principles estimate: https://epoch.ai/gradient-updates/how-much-energy-does-chatg...
Google report: https://arxiv.org/abs/2508.15734
Altman claim inside a blog post: https://blog.samaltman.com/the-gentle-singularity
Re: Gentoo AI Policy
#193Earlier quoted context omitted.
That’s what happened; the initial piracy was an issue, but those models were never released, and the models that were released were trained on copyrighted works they purchased.
That's not true, or they wouldn't have settled for 1.5bln specifically for training on pirated material. https://apnews.com/article/anthropic-copyright-authors-settl...
> A federal judge dealt the case a mixed ruling in June, finding that training AI chatbots on copyrighted books wasn’t illegal but that Anthropic wrongfully acquired millions of books through pirate websites.
With more details about how they later did it legally, and that was fine, but it did not excuse the earlier piracy:
> But documents disclosed in court showed Anthropic employees’ internal concerns about the legality of their use of pirate sites. The company later shifted its approach and hired Tom Turvey, the former Google executive in charge of Google Books, a searchable library of digitized books that successfully weathered years of copyright battles.
> With his help, Anthropic began buying books in bulk, tearing off the bindings and scanning each page before feeding the digitized versions into its AI model, according to court documents. That was legal but didn’t undo the earlier piracy, according to the judge.
Re: Gentoo AI Policy
#194Earlier quoted context omitted.
> Have you ever contributed to a very large project like LLVM? Oh, I did. Here's one: https://github.com/mariadb-corporation/mariadb-columnstore-e... > I would say clearly not from the comment. Of course, you are wrong. > It’s not so small that you can get everything in your head with only a reading. PSP/TSP recommends writing typical mistakes into a list and use it to self-review and to fix code before sending it in…
The personal dig was unwarranted. I apologise. > So, after reading code, one should write down what made him amazed and find out why it is so - whether it is a custom of a project or a peculiarity of code just read. Sorry but that’s delusional. The amount of people actually able to meaningfully read code, somehow identify what was so incredible it should be analysed despite being unfamiliar with the code base, mainta…
>The amount of people actually able to meaningfully read code, somehow identify what was so incredible it should be analysed despite being unfamiliar with the code base, maintain a list of their own likely error and self review is so vanishingly low it might as well not exist.
List of frequent mistakes gets collected after contributions (attempts). This is standard practice for high quality software development and can be learned and/or trained, including on one's own.LLVM, I just checked, does not have a formal list of code conventions and/or typical errors and mistakes. Could they have that list, we would not have the pleasure to discuss that. That PR we are discussing would be much more polished and there would be much less than several dozens of comments.
> If that’s the bare a potential new contributor has to cross, you will get exactly none.
You are making very strong statement, again.Re: Gentoo AI Policy
#195Earlier quoted context omitted.
> I can't help but wonder how these policies would be enforced One of the parties that decided on Gentoo's policy effectively said the same thing. If I get what you're really asking... the reality is, there's no way for them to know if a LLM tool was used internally, it's honor system. But I mean enforcement is just ban the contributor if they become a problem. They've banned or otherwise restricted other ones for be…
I see. So if I'm understanding correctly, then this policy serves as a kind of "legal ground" from which the maintainers can take action against perpetrators, right? To add a bit more context, when I was writing the original comment, I was mainly thinking of first-time contributors that don't have any track records, and how the policy would work against them.
They aren't government and it's not that bureaucratic. As with any group, if you break the guidelines/rules they just won't want to work with you.
> To add a bit more context, when I was writing the original comment, I was mainly thinking of first-time contributors that don't have any track records, and how the policy would work against them.
No matter what, somebody has to review the contribution. First time contributors get feedback, most of them correct their mistakes and some go on to be regular contributors (like me). Others never respond, and still more others make the same mistakes over and over again.
On the topic, Gentoo has projects like GURU where users can contribute new packages that maybe aren't ready for main tree or where a full developer wouldn't be interested, it's a good place to learn if interested in working towards becoming a developer: https://wiki.gentoo.org/wiki/Project:GURU
Re: Gentoo AI Policy
#196Earlier quoted context omitted.
LLMs perform better than doctors in a randomized trial: https://jamanetwork.com/journals/jamanetworkopen/fullarticle... And here: https://arxiv.org/html/2503.10486v1
> the use of an LLM did not significantly enhance diagnostic reasoning performance compared with the availability of only conventional resources. The other one isn't peer reviewed. Your précis doesn't appear to be warranted.
> The LLM alone scored 16 percentage points (95% CI, 2-30 percentage points; P = .03) higher than the conventional resources group.
Basically they setup the experiment as a control group and a LLM-assisted group. There was no difference between the two groups and that is what was reported in the top level finding that you quote.
Then they went back and said “wait, what if we just blindly trusted the LLM? What if we had a third group that had no doctor involved — just let the LLM do the diagnosis?” This retroactively synthesized group did significantly better than either of the actual experimental groups:
> The LLM alone scored 16 percentage points (95% CI, 2-30 percentage points; P = .03) higher than the conventional resources group … The LLM alone demonstrated higher performance than both physician groups, indicating the need for technology and workforce development to realize the potential of physician-artificial intelligence collaboration in clinical practice.
Re: Gentoo AI Policy
#197Earlier quoted context omitted.
Only bothering to mention it in response to one of many review comments is nearly the same as not disclosing it.
We might know the word "disclose" very different then. I'm amenable to taking issue with them not disclosing it up front , but then their guidelines - if the person above is to be believed - don't require it, and they did disclose it a few days after opening it. It was also not them responding to an allegation or anything, they disclosed it completely on their own terms. And that was two months ago. I find that latte…
It should be clear that my objection is to the mix of coc + ai in the context of llvm, not to this specific instance where someone is acting within the rules llvm has written down.
Re: Gentoo AI Policy
#198Earlier quoted context omitted.
Let's stop bullshitting, nobody here is going to contribute to Gentoo and is now put off because of this policy change. What we're looking at is mostly JavaScript monkeys who feel personally offended because they're unable to differentiate criticism of their tools from criticism of their own personal character. The outrage is purely theoretical.
As a JavaScript monkey I believe you have a point, and this was the core of my original question. How many contributors to gentoo are upset by this? Probably none. How many potential contributors to gentoo are upset by this? Maybe dozens? I'll be amazed if this has any notable negative outcomes for Gentoo and their contributions.
Re: Gentoo AI Policy
#199I don't understand this anti-AI stance. Either the code works and is useful, and it should be accepted, or it doesn't work and it should be rejected. Does it really matter who wrote it?
The linked API policy lists specific concerns in 3 categories: copyright, quality, ethical. Which one do you not understand?
Re: Gentoo AI Policy
#200Every time I encounter these kinds of policy, I can't help but wonder how these policies would be enforced: The people who are considerate enough to abide by these policies, are the ones who would have "cared" about the code qualities and stuff like that, so the policy is a moot point for these kinds of people. OTOH, the people who recklessly spam "contributions" generated from LLMs, by their very nature, would not r…