Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

331–340 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#331
post #2

Doesn't really explain how co-pilot is stealing your community. I've used co-pilot and it works great until you are past boilerplate than it falls apart.

It doesn't, it's just a very easy to relate to argument. Generating snippets of code has nothing to do with a fully functional software package/product/service and an organic community around it. One might argue that a community could be more easily formed thanks to co-pilot because it increases developer productivity and lower the effort to contribute so OS projects actually benefit from co-pilot. If this sounds far…

IF they attributed the code snippet to the project they ripped it off from!!

But they deliberately don’t tell you thst… it’s just a code snippet floating in space like they invented it

Re: GitHub Copi­lot inves­ti­ga­tion

#332

Most of these points can also be raised against DALL-E 2, but software has one extra thorn: patents. It's a common advice to not read software patents[1] because the infringement penalties are lower if you did so unwittingly, that is, by reinventing the patented technique yourself. I wonder if using Copilot doesn't push the penalties back again to wilful infringement. Or worse, patent trolls poisoning the training da…

Has Dall-E 2 yet reproduced 1:1 anything from its training set?

Idk about dalle but with stable diffusion if you type in "Mona Lisa" or "Van Gogh" you have to fight pretty hard with your prompt to NOT get identical reproductions of those respective works

Re: GitHub Copi­lot inves­ti­ga­tion

#333
post #169

Earlier quoted context omitted.

> What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? A better future to me. I don't want pictures of my face training ML models, nor do I want my art, or my code. I don't want my face to be more recognizable to AI, and I don't want my work to contribute to the consolidation of power to a few big firms. And for what, wh…

what if a service could tell you everywhere your photo was on the Internet?

You mean google image search?

Re: GitHub Copi­lot inves­ti­ga­tion

#334
post #200
post #143

Earlier quoted context omitted.

If they continue that path, the future will be that OpenAI, Microsoft, Google etc. will pay larger and larger fines at least in the EU, until they are blocked entirely.

Which may be entirely justifiable Earlier HN thread today on a large chunk of OS code pasted almost verbatim by the CoPilot engine into a project, but stripped of any licensing references. Within the last few days, another thread on artists who have spent decades developing a unique and valuable style are making parallel complaints about Dall-E/SD/etc., where inputting "Xyz in the style of [Artist]" produces exactly…

Huh, I wonder if they decided to remove the licenses and other comments from code before training on it. That would almost be necessary to avoid comments ending up inside of code.

Re: GitHub Copi­lot inves­ti­ga­tion

#335
post #246
post #190

Earlier quoted context omitted.

You're not supposed to be able to use dominance in one market (git hosting) to gain dominance in another (AI powered code suggestions).

What do you mean "You're not supposed to"? Is there some law that forbids this? From my (potentially naive) POV this seems to be roughly equivalent to asking physics professors to stay away from mathematics since they are likely to have some relevant cross-domain expertise.

Yes, this is the basis of antitrust law

Re: GitHub Copi­lot inves­ti­ga­tion

#336
post #2

Doesn't really explain how co-pilot is stealing your community. I've used co-pilot and it works great until you are past boilerplate than it falls apart.

They don't really explain in a satisfactory way how Copilot is "stealing communities" even though they themselves complain that Microsoft hasn't provided "solid legal references". Copilot is an AI stunt, an exploration, trying something new and exciting with very mixed and not-so-useful results. This lawsuit, however, is just lawyers doing what they do for fun. I guess the retained Microsoft lawyers love it too. Glad…

I’d guess it is basically: you’re making a mural people can look at for free, but a business selling canvas reproduction of it without attribution.

If I solve a problem and I say: Sure, use my code if you want, but be sure to contribute any improvement back to the community - I wouldn’t be happy seeing a tool spewing it out everywhere. And it’s not even free. In this case, microsoft is literally making money on the back of millions of programmers. And without approval.

Re: GitHub Copi­lot inves­ti­ga­tion

#337

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

You should read the article before commenting. The WAY it was done with copilot is the problem: no attribution, just shoving all legal liability off on the end “programmer” without providing the attribution required TO COMPLY WITH LICENSES as the diligent programmer tries to clear all the code copilot handed it without meta data. Go read the article before arguing further, please. Otherwise you are wasting all of our…

I did read the article first. Start-to-finish. As others have pointed out, it's a very visually appealing website.

I hope that when a case on generative models hits the courts that it's found that training on data from the Internet counts as fair use. I hope that for the reasons I laid out in my comment, because I think that if it isn't then we are all in trouble since these tools will STILL EXIST, but they will be in the hands of the few instead of the many. My main reaction is to how short-sighted it seems like the authors and many others are being about this technology in general. They seem to think they can just wish it away.

I also think that training on data from the Internet is fair-use, but I'm not a lawyer and I haven't studied the law extensively, so who cares what I think about that.

Re: GitHub Copi­lot inves­ti­ga­tion

#339

One issue I see with Copilot is that they get free access to all open-source data on GitHub, but using GitHub APIs to download the data yourself isn't possible (rate limiting). This is an unfair advantage. Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't. I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.

> using GitHub APIs to download the data yourself isn't possible

Is it? Data storage would be prohibitive, but I can see ways to download the entirety of Github in a few weeks/months (assuming my size estimate is accurate).

Re: GitHub Copi­lot inves­ti­ga­tion

#340

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

It's a bit of a bait and switch (obviously not literally, these things didn't exist so nobody was ever promised they wouldn't be used).

But in terms of user behavior, it's rather the same. I used to make more stuff publicly available online than I do now, and the mass-scale surveillance and data modeling that big companies do off of publicly available stuff is a big part of that.

Generally that's how you get walled gardens - by abusing the commons - but here you'd need not just a walled garden but a TINY TINY invitation only one if you don't want people doing mass surveillance and data modeling (CoPilot is really more of the latter than the former, but any of this "scrape the whole internet" stuff is just a tiny little sidestep away from being used for more blatantly evil surveillance purposes - here we're training a generative model, they're we're de-annonymizing everything you've written anywhere...).

Is there a good solution to "BigCos are gonna do whatever they want with the shit you make" other than invite-only, paid-content type models?

Post reply on HN