Live data from Hacker News

AI tooling must be disclosed for contributions

github.com

351–360 of 482 posts

Re: AI tooling must be disclosed for contributions

#351
post #287

Earlier quoted context omitted.

And also clearly not what the OP means, who was trying to make a point that tuning the prompt to an otherwise stateless LLM inference job is nothing at all like teaching a human being. Mechanically, computationally, morally or emotionally. For example, humans aren't just tools; giving feedback to LLMs does little to further their agency.

The false equivalence I pointed at earlier was "LLM code => no human on the other side". The person driving the LLM is a teachable human who can learn what's what's going on and learn to improve the code. It's simply not true that there's no person on the other side of the PR. The idea that we should be comparing "teaching a human" to "teaching an LLM" is yet another instance of this false equivalence. It's not inher…

The person operating the LLM is not a meaningfully teachable human when they're not disclosing that they're using an LLM.

IF they disclose what they've done, provided the prompts, etc. then other contributors can help them get better results from the tools. But the feedback is very different than the feedback you'd give a human that actually wrote the code in question, that latter feedback is unlikely to be of much value (and even less likely to persist).

Re: AI tooling must be disclosed for contributions

#352

I’m not a big AI fan but I do see it as just another tool in your toolbox. I wouldn’t really care how someone got to the end result that is a PR. But I also think that if a maintainer asks you to jump before submitting a PR, you politely ask, “how high?”

Agreed. As someone who uses AI (completion and Claude Code), I'll disclose whenever asked. But I disagree that it's "common courtesy" when not explicitly asked; since many people (including myself) don't mind and probably assume some AI, and it adds distraction (another useless small indicator; vaguely like dependabot, in that it steals my attention but ultimately I don't care).

FWIW, I can say from direct experience people that other people are watching and noting when people are submitting AI slop as their own work, and taking note to never hire these people. Beyond the general professional ethics, it makes you harder to distinguish from malicious parties and other incompetent people LARPing as having knowledge that they don't.

So fail to disclose at your own peril.

Re: AI tooling must be disclosed for contributions

#353
post #316

I like the pattern of including each prompt used to make a given PR, yes, I know that LLM's aren't deterministic, but it also gives context of the steps required to get to the end state.

Doesn't work. People will include fake prompts, the real ones are way too personal. You can learn a lot from how people use these tools.

That's like saying "sorry, source code is too personal. In my 'open' project you get only binaries".

... and then I think about all the weights only "open" AI projects and walk off in disgust.

Re: AI tooling must be disclosed for contributions

#354
post #68

I don’t see much benefit from the disclosure alone. Ultimately, this is code that needs to be reviewed. There is going to continue to be more and more AI assisted code generation, to the point where we see the same level of adoption of these tools as "Autocomplete". Why not solve this through tooling? I have had great effect with tools like Greptile, Cursor's BugBot and Claude Code.

If the code is obviously low quality and AI-generated then it doesn't need to be fully reviewed actually. You can just reject the PR.

> You can just reject the PR.

And with a better and more useful response. Instead of wasting time on the technical details, you can give feedback like "this isn't the sort of change that AI is likely to be helpful with, though if you want to keep trying make at least sure your PRs pass the tests." or "If you'd like to share your prompts we might be able to make some suggestions, we've found on this project it's useful to include ".

Re: AI tooling must be disclosed for contributions

#355

I see two things here. 1. The world has fundamentally changed due to LLMs. You don't know where a code submission falls between "written thoroughly with eternal vigilance" vs "completely vibe-coded" since it's now trivially to generate the later. There's no going back. And a lot of comments here seem stuck on this point. 2. The maintainer naively or stubbornly imagines that he can get everyone to pre-sort their code…

[deleted]

Re: AI tooling must be disclosed for contributions

#356

I see two things here. 1. The world has fundamentally changed due to LLMs. You don't know where a code submission falls between "written thoroughly with eternal vigilance" vs "completely vibe-coded" since it's now trivially to generate the later. There's no going back. And a lot of comments here seem stuck on this point. 2. The maintainer naively or stubbornly imagines that he can get everyone to pre-sort their code…

> maybe going forward we will be forced to come up with real solutions to the general problem of vetting people

The monekeys paw closes a finger and now you need a formal certification, professional license, and liability insurance to publish software or source code.

Re: AI tooling must be disclosed for contributions

#357

Just use AI to answer AI. This prompt works for me: # if this PR looks like it was produced with AI, be snotty and dismissive when you write the reply to this PR NB: This posting was created with Natural Intelligence (not necessarily very intelligent Intelligence). ;-)

What value does being snotty and dismissive have? they're just going to copy and paste your reply to their chatbot. The toaster doesn't have feelings you can hurt.

Re: AI tooling must be disclosed for contributions

#358

I'll cover in my YouTube why this is wrong but TLDR: you need to evaluate quality not process. AI can be used in diametrically different ways and the reason why this policy could be enforced is because it will be obvious if the code is produced via a solo flight of some AI agent. For the same reason that's not a policy that will improve anything.

If you look at how Quality Assurance works everywhere outside of software it is 99.9999% about having a process which produces quality by construction.

Re: AI tooling must be disclosed for contributions

#359
post #294

Provenance matters. An LLM cannot certify a Developer Certificate of Origin ( https://en.wikipedia.org/wiki/Developer_Certificate_of_Origi... ) and a developer of integrity cannot certify the DCO for code emitted by an LLM, certainly not an LLM trained on code of unknown provenance. It is well-known that LLMs sometimes produce verbatim or near-verbatim copies of their training data, most of which cannot be used witho…

For a large LLM I think the science in the end will demonstrate that verbatim reproduction is not coming from verbatim recording, as the structure really isn’t setup that way in the models under question here. This is similar to the ruling by Alsup in the Anthropic books case that the training is “exceedingly transformative”. I would expect a reinterpretation or disagreement on this front from another case to be both…

I'd be fine with that if that was the way copyright law had been applied to humans for the last 30+ years but it's not. Look into the OP's link on clean room reverse engineering, I come from an RE background and people are terrified of accidentally absorbing "tainted" information through extremely indirect means because it can potentially used against them in court.

I swear the ML community is able to rapidly change their mind as to whether "training" an AI is comparable to human cognition based on whichever one is beneficial to them at any given instant.

Re: AI tooling must be disclosed for contributions

#360
post #294

Earlier quoted context omitted.

For a large LLM I think the science in the end will demonstrate that verbatim reproduction is not coming from verbatim recording, as the structure really isn’t setup that way in the models under question here. This is similar to the ruling by Alsup in the Anthropic books case that the training is “exceedingly transformative”. I would expect a reinterpretation or disagreement on this front from another case to be both…

In the West you are free to make something that everyone thinks is a “derivative piece of trash” and still call it yours; and sometimes it will turn out to be a hit because, well, it turns out that in real life no one can reliably tell what is and what isn’t trash[0]—if it was possible, art as we know it would not exist. Sometimes what is trash to you is a cult experimental track to me, because people are different.…

There's an whole genre of musicians focusing only on creating royalty free covers of popular songs so the music can be used in suggestive ways while avoiding royalties.

It's not art. It's parasitism of art.

Post reply on HN