Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

271–280 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#271

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

Well, what does the future where those materials are free use look like? You argue that desiring ownership of works that you created is “extremely emotional and cherry-picked”, but you do not provide a compelling argument why artists, photographers, and indeed programmers should be excluded from the conversation when it is their art, photography and code that is being appropriated in the first place .

I think that future has a lot of really cool and cheap tools that will help all those artists, photographers, programmers, etc get even more out of what they love doing. I think there will wind up being some job markets that shrink (not with the tech we have now though) for small-medium things in all of these fields (think logo design, stock photos, client libraries, simple out-of-the-box applications) and my hope is that those job markets that shrink cause others to grow due to the increased levels of productivity that these tools will give us.

Ultimately my argument is that these tools will allow human beings to accomplish more things with less and that these tools should be distributed to as many people as possible for as little cost as possible. Part of that belief comes from the fact that I think these tools are coming no matter what and I'm slightly concerned about the potential (although unlikely-looking) future where a small number of large corporations are the only ones controlling these tools and they just rent-seek on them.

Re: GitHub Copi­lot inves­ti­ga­tion

#272
post #187

Earlier quoted context omitted.

Copy & paste isn't "standing on the shoulders of others". It is more like being an intestinal parasite.

Not trying to overly advocate for copy-pasting here, but isn't copy-pasting just the ugly child of calling a library function? If it's a blind copy-paste it's pretty much the same effect. Surely you wouldn't call using a library being an intestinal parasite?

Isn't copying a poem and releasing it under your own name (i.e, stealing it) the same as referring to it in a footnote, using proper attribution?

Re: GitHub Copi­lot inves­ti­ga­tion

#273

Earlier quoted context omitted.

I mean, he clearly knew about that code in advance and used his prior knowledge to coax Copilot into spitting it out, yeah? Three characters can get you pretty far, that's 1 combination out of 125,580 (considering all english letters, upper and lower, along with most of the numbers and symbols on my keyboard), plus the description of a fairly complex algorithm. Also, this code is really just executing a mathematical…

I'm sure you see you've been downvoted. But I want to say I agree with you on the millions / billions bit. While I understand the rub about licenses, the fact is the vast majority of code is not all that original or unique. Some fringe amount is, and those edge cases are worth discussing. But the rest? Likely not all in all all that special. Yes, we get paid good money to do it. But is that a function of what it take…

This example is one of those rare pieces of code that is special though. It's the product of years of deliberate work by professor-level academics. This is exactly the kind of person who would have the least to fear from copilot if it really was just automating the boring plumbing parts and not shamelessly copying high-value, creative, insightful code.

Re: GitHub Copi­lot inves­ti­ga­tion

#274

Copilot is great and this is a waste of time.

Being a useful tool doesn't make it legal.

Technical progress takes precedence over pitiful intelectual property discussions. If you don't believe that, I am not sure what you are doing in a community like this.

Re: GitHub Copi­lot inves­ti­ga­tion

#275

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

There does seem to be more heat and noise then substantive discussions here.

Though that doesn't justify such a dismissive attitude towards ordinary HN commenters. As the way it's written implies that most are too stupid and overly emotional, which is more likely to fuel complaints instead of dousing them.

Re: GitHub Copi­lot inves­ti­ga­tion

#276
post #246
post #190

Earlier quoted context omitted.

You're not supposed to be able to use dominance in one market (git hosting) to gain dominance in another (AI powered code suggestions).

What do you mean "You're not supposed to"? Is there some law that forbids this? From my (potentially naive) POV this seems to be roughly equivalent to asking physics professors to stay away from mathematics since they are likely to have some relevant cross-domain expertise.

> Is there some law that forbids this?

Yes. It's called the Sherman Act, and it's the basis of anti-trust enforcement in the US.

https://www.ftc.gov/advice-guidance/competition-guidance/gui...

>>

I know lots of people here don't like it, but it is the law and that was the question; "this" in parent clearly meant "use dominance in one market to gain dominance in another" in grandparent, regardless of whether that's actually the central issue here or not.

Re: GitHub Copi­lot inves­ti­ga­tion

#277
post #169

Earlier quoted context omitted.

> What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? A better future to me. I don't want pictures of my face training ML models, nor do I want my art, or my code. I don't want my face to be more recognizable to AI, and I don't want my work to contribute to the consolidation of power to a few big firms. And for what, wh…

what if a service could tell you everywhere your photo was on the Internet?

Such a service would make a great and fantastic service for stalkers to bypass the usual difficulties in locating someone who has done the best to excise them from their life.

Re: GitHub Copi­lot inves­ti­ga­tion

#278
post #247

There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out . The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principl…

This was a comment made to me in a previous, similar discussion, discussing case law around Google's use of copyrighted books in building a search engine: https://news.ycombinator.com/item?id=32654478

I'm not sure I completely agree w/ the comment (nor do I think it vindicates CoPilot), but I think it does provide insight into why CoPilot is violating copyright.

Re: GitHub Copi­lot inves­ti­ga­tion

#279
post #187

Earlier quoted context omitted.

Copy & paste isn't "standing on the shoulders of others". It is more like being an intestinal parasite.

If Copilot was just a "copy and paste" tool very few people would find it useful and you wouldn't be here whining. So don't worry, Copilot is not just copying and pasting.

Critics of intellectual property theft are "whining". This is useful information in the next Microsoft IP lawsuit.

Re: GitHub Copi­lot inves­ti­ga­tion

#280
post #169

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? A better future to me. I don't want pictures of my face training ML models, nor do I want my art, or my code. I don't want my face to be more recognizable to AI, and I don't want my work to contribute to the consolidation of power to a few big firms. And for what, wh…

If they hire photographers to take photos of people in public and use them for training, there’s no law stopping them really. Your only real hope would be to always walk around in a burqa.
Post reply on HN