Live data from Hacker News

StackOverflow petition to allow removing AI generated content

openletter.mousetail.nl

111–120 of 120 posts

Re: StackOverflow petition to allow removing AI generated content

#111

Earlier quoted context omitted.

This doesn't work at scale. Stack overflow as a platform has been handling user generated input via moderators, voting, and testing. This is fine when there are only 26.8 million coders on the planet, most of which aren't posting on stack overflow regularly. With LLM's all of a sudden there is a huge influx of mediocre content on the platform that people can't handle. Inevitably this will erode trust in the platform.…

How are they going to check for LLM usage? I think it's way more likely that poor answers won't mention the usage of LLM's to generate the answer, while good answers aided by LLM's will more often mention it. Punishing honesty just seems incredibly counterproductive. Automatic detection is downright dystopian... being censored by an algorithm because it mistook my effort and work for a LLM.

Agree with the middle part - At the moment, the policy implemented by corporate is "Don't ask; don't tell". If someone says they used GPT or other AI for their answer, it's disallowed. If they try to hide the fact, there's not much the community can do to get it removed.

And while I'm not a moderator, as just a user I've flagged over 1,200 answers on Stack Overflow (and several of the smaller communities like Ask Ubuntu) that were subsequently removed. Automatic detection was never the sole criteria that was used to determine if it was AI - It's entirely possible to spot GPT content using multiple methods. I don't publicly talk about most of these, since we do have a group of users (sometimes spammers) who attempt to hide their use and make it more difficult to detect. See some of my additional notes on the topic on https://meta.stackexchange.com/a/389674/902710

Re: StackOverflow petition to allow removing AI generated content

#112

Either answers are good or not. It doesn't matter if they're generated by a 13-year-old in their bedroom, someone studying CS at university, a well-respected IC at a top tech company... or an AI. If answers are good, keep them. If they're bad, downvote them. If they're redundant or off-topic or gibberish, delete them. And to those asking why you would ever want AI-generated content on StackOverflow when you could jus…

If you as an answerer uses an LLM to generate content, then verify and vet it yourself based on your own knowledge before posting, I'd think it's fine. But spamming thousands of answers an hour automatically and wanting the community to do all the work is just not sustainable I feel. It'll also kill the sense of community if half the actors are bots.

Agreed - That's the basis of my "responsible use of AI on SO" post at https://meta.stackexchange.com/a/389675/902710

Re: StackOverflow petition to allow removing AI generated content

#113
post #48

Earlier quoted context omitted.

Agreed. I have found myself sometimes asking Chat-GPT for guidance related to obscure error messages and occasionally its more useful than google or stack overflow.

When it doesn’t hallucinate, which in my case it’s most of the time.

Sure, but if you try it out, you pretty quickly realize it's a hallucination. Unfortunately the type of GPT content we're now getting on Stack Overflow and its sibling sites is mostly unvalidated GPT hallucinations.

Re: StackOverflow petition to allow removing AI generated content

#114
post #51

I don't understand why you would, for the forseeable future, want to allow AI generated content on SO. Not with the current state of the art in generative AI. If I want an AI generated answer, with all the pros and cons specific to LLMs, I'll just open chatgpt or turn on copilot... SO answers (used to be) in an entirely different league in terms of trustworthiness. I'm totally on board with this moderator strike. Edi…

If the code works just fine, why remove it? For small algorithms it should be good enough.

"Code working" isn't necessarily black-and-white. For a new user (the one asking the question), the code may appear to solve the problem, but may have corner-cases or even security risks. That's entirely possible with user-generated code as well, of course, but GPT/AI allows it to be produced at a much higher rate, with the person who posted the answer often not being capable of (or not caring to) validate or correct it.

Re: StackOverflow petition to allow removing AI generated content

#115
post #66

Either answers are good or not. It doesn't matter if they're generated by a 13-year-old in their bedroom, someone studying CS at university, a well-respected IC at a top tech company... or an AI. If answers are good, keep them. If they're bad, downvote them. If they're redundant or off-topic or gibberish, delete them. And to those asking why you would ever want AI-generated content on StackOverflow when you could jus…

haha. no. answers are not either good or not. it's not binary. there is a lot of nuance and sometimes the questions and followups contain details that can take the answers in a completely new direction. just telling if an answer is good or not is a lot of work. sometimes it requires and expert to figure it out. when the answers are given by a human and it's a good faith effort both people in the loop benefit from it.…

> when the BS generation is automated there is zero incentive for a human being to even look and correct the answers. what is the incentive to do so?

You hit a good point here. If users can't be bothered to spend time, effort, energy and cognition into answering a question, why should the readers and correctors do so?

Re: StackOverflow petition to allow removing AI generated content

#117
post #54

I'd propose a different solution. For every single new question, auto-generate an answer with GPT and mark it as AI-generated.

At least the generated responses won’t start by asking me “why in the world would you want to do that?”

You nailed it.

Re: StackOverflow petition to allow removing AI generated content

#118

Anyone around who moderators on SO? Is there a general sense of alienation from corporate SO? My understanding is that in the early days, a lot of the devs at SO were actually recruited from the SO and Meta moderation userbase but probably that doesn't scale. Example: Ben: And then, April 29th, I was, again, sitting in the agency, doing my thing, and I also had my personal email open, and I suddenly got an email from…

I used to be heavily active in curating SO, and yes, there is an incredibly strong sense of alienation from the corporation (which is what drove me away).

In the old days, most of the staff, from devs to management to the CEO, were active users of the site and hung out on Meta and in chat. They were easy to reach and happy to answer questions; and whenever they made major changes to the site they would go to Meta to ask for feedback and adjust their plans accordingly.

There was a noticeable shift in this dynamic starting around ~2016. Around this time, the company stopped focusing development effort on the core site functionality, and instead prioritized side products & attempts to monetize the site (most of which failed). Feature requests on Meta were almost completely ignored, and site features that had been on previously-announced timelines/roadmaps were never delivered. But the community was still as strong as ever; and so people started implementing these missing features themselves in the form of bots and userscripts. This was the "golden age" of moderation bots, and a really fun time to be a part of the community -- power users ran heavily modified frontends that could display all kinds of additional information and automate repetitive actions, and integrate with bots to do things that Stack Exchange's systems were bad at -- like flagging spam and low-quality posts; identifying plagiarism; and detecting flamewars in comments.

As the company grew, they hired a ton of middle management who were not active participants in the site, and largely did not care about the day-to-day. This was alright when they left us alone, but around 2018-2019 they began to take an openly hostile stance towards the "power users". Here's an excellent post from that time summarizing the general sentiment: https://meta.stackexchange.com/questions/331513

The short version is: the company began blaming power users for things like the site's "unwelcoming" reputation (which is really a symptom of the site's outdated and opaque moderation tooling, and power users had been clamoring for better tools for years). They began a pattern of rolling out features and UI changes that took major steps backwards in usability and accessibility -- and due to all the negative feedback these changes received on Meta, they announced staff would no longer participate on Meta because it was "too negative". A high-level manager famously quipped that the opinion of Meta was not relevant as it represented "0.015%" of Stack Exchange's userbase -- despite the fact that that 0.015% was responsible for the majority of content & moderation activity contributed to the site.

In late 2019 it got a whole lot worse when, in rapid succession: 1) the company updated the site terms-of-service to illegally change the license of user-submitted content, and 2) an employee abruptly revoked a volunteer's moderation privileges without due process, and then went to news outlets making false accusations about that user's behavior (https://meta.stackexchange.com/questions/333965). Shortly after that, Stack Overflow fired several well-loved and highly respected staff moderators, for undisclosed reasons (https://meta.stackexchange.com/questions/342039/). A lot of people, including myself, left the community in the wake of this.

Since then -- at least from my outsider perspective of checking in once in a while to see what's going on -- it seemed for a while like the company was learning from its mistakes.They apologized for ignoring Meta, began asking for and listening to community feedback once again, created new policies to protect volunteers from the kind of abuse that happened in 2019, and began implementing some of those long-ignored and long-overdue feature requests. But in 2021 the company got bought out by a VC firm that is even more aggressive about trying to monetize the site; they started cranking up advertising and pushing generally unwanted side products, but they mostly left the community alone.

That brings us to generative AI. As soon as ChatGPT came out, a deluge of users began copy/pasting Stack Overflow questions into ChatGPT and copy/pasting its answers into the answer box (usually with no editing or fact-checking effort). General consensus among the community seems to be that ChatGPT produces wrong or unhelpful information to an unacceptable degree, and that allowing machine-generated content on Stack Overflow defeats the purpose of the site (you go to ChatGPT if you want answers from a machine, but you go to SO if you want answers from a human). The staff supported this consensus and made it official policy -- but at the same time the CEO kept making rambling blog posts about how "AI is the future about Stack Overflow" and launching sweeping initiatives within the company to do...AI related things? Nobody really knows what he's talking about.

That all leads up to the events that triggered strike: out of nowhere, the company suddenly announced a few days ago that it was overruling previous community consensus and prohibiting users from deleting content on the basis of being AI-generated.

Re: StackOverflow petition to allow removing AI generated content

#119
post #51

Earlier quoted context omitted.

If the code works just fine, why remove it? For small algorithms it should be good enough.

"Code working" isn't necessarily black-and-white. For a new user (the one asking the question), the code may appear to solve the problem, but may have corner-cases or even security risks. That's entirely possible with user-generated code as well, of course, but GPT/AI allows it to be produced at a much higher rate, with the person who posted the answer often not being capable of (or not caring to) validate or correct…

Yes, and SO already have plentiful of sample code that appear to solve the problem, but have huge flaws.

Re: StackOverflow petition to allow removing AI generated content

#120

Anyone around who moderators on SO? Is there a general sense of alienation from corporate SO? My understanding is that in the early days, a lot of the devs at SO were actually recruited from the SO and Meta moderation userbase but probably that doesn't scale. Example: Ben: And then, April 29th, I was, again, sitting in the agency, doing my thing, and I also had my personal email open, and I suddenly got an email from…

I used to be heavily active in curating SO, and yes, there is an incredibly strong sense of alienation from the corporation (which is what drove me away). In the old days, most of the staff, from devs to management to the CEO, were active users of the site and hung out on Meta and in chat. They were easy to reach and happy to answer questions; and whenever they made major changes to the site they would go to Meta to…

Wow, great summary. I did not know about the relicensing bit. That seems problematic!
Post reply on HN