Live data from Hacker News

Copilot regurgitating Quake code, including sweary comments

twitter.com

521–530 of 672 posts

Re: Copilot regurgitating Quake code, including sweary comments

#521

Earlier quoted context omitted.

Javascript books under "Law" is hilarious

I'm in for JavaScript Penal Code. Make that unwarranted type-coercing operator use punishable by law.

Instead of prison you go to callback hell

Re: Copilot regurgitating Quake code, including sweary comments

#522
post #509

Earlier quoted context omitted.

I don't know if I've ever heard anyone use the term "concentration camp" without qualifiers to refer to anything else than the nazi concentration camps (or something equivalent). Maybe it's just me, but I think it would have been more clear if you said internment camp if your intent was to refer to the broader context and not invoke a comparison to nazis.

Wikipedia redirects concentration camp to: https://en.wikipedia.org/wiki/Internment Where it also makes the point that the nazi camps were primarily extermination camps. Maybe take it up with them and get back to me if you feel truly passionate about this issue. >Maybe it's just me, but I think it would have been more clear Gosh, it's awfully ironic that this sentence would happen in a thread about how language polic…

> So, is it more important to you how people use the term concentration camp or the fact that ICE lock up children in internment/concentration/[ insert favorite word here ] camps?

Well, that escalated quickly.

I don't think I ever said anything for or against what ICE is doing, in fact I tried not to because the only thing I wanted to say was that when using the words "literally concentration camps" people might read that as "camps designed to kill people" since that is the way I've been taught it (in history classes) and heard it (in general use).

I don't even live in the US so I have no say in this in a democratic sense. If I did I'd be against the way migrants are treated and want more humane treatment, but I don't think that should be relevant to what I said.

Re: Copilot regurgitating Quake code, including sweary comments

#523
post #507

Earlier quoted context omitted.

That kind of bullshit phrasing can only get you so far. It's like if some corporate PR department told you "we're aware of the halting problem, and are taking steps to mitigate it." You would rightly laugh them out of the room. It's not going to work, and the people making these statements either don't understand how much they don't understand, or are deluding themselves, or are actively lying to us. An honest answer…

You're using some pretty strong language here, but do you have any more substantive criticisms of the analysis they present at https://docs.github.com/en/github/copilot/research-recitatio... ? They seem to think the incidence of meaningful (i.e. substantively infringing) recitation is very low, and that their solution in those cases will be attribution rather than elimination. Again, I'm not an ML expert, but that so…

The biggest issue with that analysis is that their model is clearly very able to copy code and change the variable names, copying code and changing variable names is very clearly still "copying", and the analysis doesn't seem to include that in its definition of "recitation event".

Re: Copilot regurgitating Quake code, including sweary comments

#524
post #365

Earlier quoted context omitted.

> I don't really have a problem typing. Im absolutely with you and want to upvote that part of the comment x100. Unfortunately it's often considered a fairly spicy opinions. Entire frameworks (Rails) are built around the idea of typing as little as possible. Others can't even be mentioned without the topic of boilerplate/keystroke count causing a flame war (Redux). A lot of engineers equate their value with the amoun…

I never commit something that I can easily google (with a high quality solution) to memory

[deleted]

Re: Copilot regurgitating Quake code, including sweary comments

#525
post #408

Earlier quoted context omitted.

Changing master to main was something Github did when they were taking heat for their contract with ICE. It was a nice bit of misdirection that cost them nothing, achieved nothing and garnered praise in some quarters. ICE, of course, runs an actual concentration camp which has a slightly more troublesome history than the word master. Language policing is to racism what recycling is to global warming - an attempt to s…

y'know it really seems like both purpose and outcome need to be closely examined here, if we're going to be emphasizing actual next to concentration camps. what's the paradigm of a concentration camp? if we go straight for Auschwitz we'll get nowhere, how about the Boer concentration camps? Origin of the term after all. What was the purpose? To concentrate the Boer population during a total war against them, so they…

Is this an indirect way of saying that you support ICE?

Coz if so Id really rather hear it straight rather than indirectly via an attempt to police my language.

Re: Copilot regurgitating Quake code, including sweary comments

#526
post #519

Earlier quoted context omitted.

https://gitgud.io/AuroraPurgatio/aurorapurgatio https://www.reddit.com/user/non-taken-name

I don't get it, that seems like standard fare for an R-rated movie? And then it seems like some complained because they decided to start editing it down to a PG-13 movie?

Essentially, from my understanding, there was a data leak they never commented on, they instituted a poorly made content filter without saying anything. The filter frequently has false positives and negatives, someone discovered they trained the game using content the filter was designed to block, meaning the ai itself would frequently output filter triggering stuff, more people found out their private unpublished stories were being read by third parties after a job ad and the stories were posted on 4Chan, people recognized stories they wrote that had triggered the filter that were posted, and then they started instituting no warning bans.

I might have missed something, but that's the gist of it.

Re: Copilot regurgitating Quake code, including sweary comments

#527

Earlier quoted context omitted.

You're using some pretty strong language here, but do you have any more substantive criticisms of the analysis they present at https://docs.github.com/en/github/copilot/research-recitatio... ? They seem to think the incidence of meaningful (i.e. substantively infringing) recitation is very low, and that their solution in those cases will be attribution rather than elimination. Again, I'm not an ML expert, but that so…

The biggest issue with that analysis is that their model is clearly very able to copy code and change the variable names, copying code and changing variable names is very clearly still "copying", and the analysis doesn't seem to include that in its definition of "recitation event".

I'd fully expect it to copy code and change variable names in a lot of cases--if it wants to achieve the goal of filling in boilerplate, how could it do anything else? That's pretty much the definition of boilerplate: it's largely the same every time you write it.

What's less clear to me is that Copilot regularly does that sort of thing with code distinctive enough that it could reasonably be said to constitute copyright infringement. If somebody's actually shown that it does, I'd love to see that analysis.

Re: Copilot regurgitating Quake code, including sweary comments

#528
post #465
post #385

Earlier quoted context omitted.

Ahh, so it's the most pointless interpretation of the phrase "filters to block offensive words", where it is stopping the user from causing offense to the AI rather than the other way around.

I believe the concept is to stop users from prompting the AI to generate offensive stuff specifically, and then publishing the so-generated stream of offensive stuff as negative PR for GitHub, in the same way the generated stream of offensive stuff coming from Microsoft’s AI was a big PR disaster.

Maybe, but even if so, filtering the output would also prevent this.

Re: Copilot regurgitating Quake code, including sweary comments

#529

I hate to be the one that says this but I think it‘s true: "So you are an SWE and you take a break from work to go to Hackernews to complain that Github's Copilot, which is an AI-based solution meant to help SWEs, is utter shit and completely unusuable. And then you go back to writing AI-based solutions for some other profession. Which is totally not shit or anything.“ Can anybody put this more elegantly?

SWEs create AI based solutions to X 'cause people pay them. Entrepreneurs and investors are the one who actually think they're the answer to everything.

Also, Copilot might (or might not) be useless or even interfere with real work. But it's probably low on the scale of awful things SWEs have helped create. The AI parole app is a thing that should haunt the nightmare of whoever created it, for example. But lots of AI apps may be useless but are probably also harmless so doing that might not be worst thing.

Re: Copilot regurgitating Quake code, including sweary comments

#530
post #507

Earlier quoted context omitted.

That kind of bullshit phrasing can only get you so far. It's like if some corporate PR department told you "we're aware of the halting problem, and are taking steps to mitigate it." You would rightly laugh them out of the room. It's not going to work, and the people making these statements either don't understand how much they don't understand, or are deluding themselves, or are actively lying to us. An honest answer…

You're using some pretty strong language here, but do you have any more substantive criticisms of the analysis they present at https://docs.github.com/en/github/copilot/research-recitatio... ? They seem to think the incidence of meaningful (i.e. substantively infringing) recitation is very low, and that their solution in those cases will be attribution rather than elimination. Again, I'm not an ML expert, but that so…

They had some people use the thing for a while, and concluded "Hey look, it doesn't seem to quote verbatim very often. Yay!" There is nothing in there that describes any sort of mitigation. The three sentences about an attribution search at the very end are aspirational at best, and are presented as "obvious" even though it's not at all clear that such a fuzzy search can be implemented reliably.

I use the halting problem as an analogy because their naive attempts to address this problem feel a lot like naive attempts to get around the halting problem ("just do a quick search for anything that looks like a loop," "just have a big list of valid programs," etc.). I can perform a similar analysis of programs that I run in my terminal and come to a similar "Hey look, most of them halt! Yay!" conclusion. I can spin a story about how most of the ones that don't halt are doing so intentionally because they're daemons.

But this approach is inherently flawed. I can use a fuzz tester to come up with an infinite number of inputs that cause something as simple as 'ls' to run forever.

Similarly, I can come up with an infinite number of adversarial inputs that attempt to make Copilot spit out training data. Some of them will work. Some of them will produce something that's close enough to training data to be a concern, but that their "attribution search" will fail to catch. That's the "open research question" that they need to solve.

We don't have a general solution to this problem yet, and we may never have one. They're trying to pass off a hand-wavey "we can implement some rules and it won't be a problem most of the time" solution as adequate. I don't see any reason to believe that it will be adequate. Every attempt I've seen at using logic to try and coax a machine learning model into not behaving pathologically around edge cases has fallen flat on its face.

Post reply on HN