Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

151–160 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#151
post #38

Abolish all copyright. We're all happily pirating movies and music but code is for some reason sacred.

FWIW, not all of us "happily pirate movies and music".

I want there to be more good music and movies. I want to support artists who create entertainment I enjoy. I go out of my way to buy physical copies of music from artists, wherever possible from the merch table at their shows or from their own websites. I pay to go see movies on the big screen (partly because I like the big screen cinema experience, but also because I understand "opening week revenue" is a key performance indicator for the success of a movie).

I thing copyright is old, outdated, and probably not really fit for purpose for forms of creative work invented in the last 50 years. But I also thing creative workers need to get paid for their effort (juist the same as software developers), and absent a FAANG-style set for companies employing teams of songwriters, musicians, authors, and the like - on FAANG-style salaries, copyright seems to be the option that is working (however badly).

I'll join your "abolish all copyright" crusade as soon as there's an alternative that at least likely to possibly work as well (or better) than the system copyright allows. Just abolishing copyright and erasing the publishing/music/movie/art industries without a transition plan isn't a thing I can support. (At least a transition the artists/editors/producers/writers/etc. I'll admit there's a large chunk of management and legal in the fairly abusive parts of the music industry I wouldn't shed a tear if they all became homeless and destitute overnight...)

Re: GitHub Copi­lot inves­ti­ga­tion

#152

Never forget this is how people who dare to reverse engineer Windows are treated: https://www.theregister.com/2019/07/03/reactos_windows_resea... https://marc.info/?l=ros-dev&m=118775346131654&w=2 I don't use Github, but fuckers upload my code there anyway. Copyright is evil, but only large corporations having copyright, even more than they already do, is even worse.

This I feel like is one of the better points in the thread.

The asymmetry that exists in copyright law where large corporations can enforce their copyright to the point of breaking the law themselves (YouTube's content ID is another non-legal, but still very impactful example) is absolute bullshit.

Unfortunately I think that if training ML models on Internet-data is found not to be fair use then things will get harder for individuals training models and corporations will be barely inconvenienced as they can afford to pay for sources, make deals with other large institutions for data, etc.

Re: GitHub Copi­lot inves­ti­ga­tion

#153
post #81

I've been trained on open source code, and there are likely many algorithms that I've internalized that are very similar to the "standard" way of performing an operation. Is there a reason why an AI being trained on the same open source code isn't a similar situation? I agree that wholesale pasting of code chunks is an issue, but that hasn't been my experience with Copilot. I'm not arguing for Copilot here...I'm genu…

>Is there a reason why an AI being trained on the same open source code isn't a similar situation? You are a human. You know what's right or wrong. You know you can't just copy code 1:1 from public repositories without respecting their license. The AI doesn't know and doesn't care. It's a common problem with creative AIs that they will occasionally regurgitate near 1:1 copies of their training data, and I don't think…

I don't deny that there are examples that appear to be wholesale copying, and that is definitely an issue to be addressed. No doubt.

What I don't understand is why the rest of the service (where it doesn't appear to be pasting existing code) is being maligned when it behaves like a more powerful version of autocomplete.

Re: GitHub Copi­lot inves­ti­ga­tion

#154

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

Any law where the penalty is a fine only exists for the poor. Any regulation where the penalty is in the millions only exists for small businesses.

Re: GitHub Copi­lot inves­ti­ga­tion

#156
Maybe I'm in the minority, but I think the prospect of someone autocompleting and getting a snippet that came from me, they found it useful, and are going to incorporate it is great. It means my thoughts and logic are shaping culture in a mimetic feedback loop.

Re: GitHub Copi­lot inves­ti­ga­tion

#157
post #45

The example given is "sparse matrix trans­pose in the style of Tim Davis", but someone who wanted something with such specificity would be able to just take it from Github anyway, perhaps with a little more searching.

And would therefore have to follow the license of the code they took it from. That's exactly the point. Copilot is reproducing the same code but without the license.

A simple search on github reveals that those functions were reposted verbatim thousands of times, most people just copy and paste snippets of code they find useful, ignoring licenses. This highlights how all the power a license promises to hold is completely fictional. Any "in the style of Tim Davis" modifier only shows some kind of unwarranted self-importance complex on the part of the guy, thinking his style is widely known and distinctive (it's not). It's not the job of Copilot, the team that builds it, or the programmers that use it, to determine where the functions that were reposted thousands of times under all kinds of licenses originated.

This is the same case as with copyrighted photos in newspapers, a paper prints a photo somebody allowed them to use, but then it turns out that person did not have the right to use it in the first place. Did not stop newspapers from printing photos at all.

Here are the search terms: https://github.com/search?q=cs_transpose&type=Code

Re: GitHub Copi­lot inves­ti­ga­tion

#158
post #127

Earlier quoted context omitted.

> It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. I wonder how many people on HN would be on the side of the creators if we were talking about content created by Walt Disney and whether pirating was ethical?

I don’t think anyone at all is here arguing that an AI trained on a massive corpus of movies that outputs snippets of film based on a prompt would be illegal. Such a thing would literally be the same as Midjourney with images. The fact that you can likely coax any AI to output snippets close to some of the source material is not likely to really matter and would be as if you recreated a copyrighted work using any oth…

What’s the difference here then?

Re: GitHub Copi­lot inves­ti­ga­tion

#160
post #78

Earlier quoted context omitted.

Copilot wouldn't be shut down or neutered because "a few vocal people" protested against it. It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus. You act like Microsoft is trying to do a public service and people are angry about it. The reality is that they're taking billions of hours of work and using it to build a product t…

Quoted post unavailable.

No post body was provided.
Post reply on HN