Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

901–910 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#901
post #748

Earlier quoted context omitted.

My code was AGPL. OpenAI can go to h..l (Footnote: I like your poem. It conveys the concept much better than anywhere I'd ever seen before)

Thanks, but it's not my poem! You can find it here: https://blog.ninapaley.com/2009/12/15/minute-meme-1-copying-...

I think it became yours when you copied it :)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#902
post #581
post #423

Earlier quoted context omitted.

the solutions haven't arrived. neither have changes in lieu of having solutions. "trying" isn't an actual, present, functional change. and it just gets passed around as an excuse for companies to keep doing whatever they're doing.

Please recall how much the world changed in just the last year. What would be your expected timescale for the solution of this particular problem and why is it more important than instilling models with the ability to logically plan and answer correctly?

the timeline for LLMs and image generation has been 6+ years. it is not a thing where it "arrived just this year, and only just changing". it's been in a development for a long time. and yet.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#903
post #568
post #518

Earlier quoted context omitted.

"Now displaying 3 citations out of ~150,000,000.." [1] http://web.archive.org/web/20120608192927/http://www.google.... [2] https://steemit.com/online/@jaroli/how-google-search-result-... [3] https://www.smashingmagazine.com/2009/09/search-results-desi... [4] Next page :)

This is not answering the GP question and does not count as a satisfactory ranked citation list. The first one is particularly dubious. Also you didn’t clarify which statement was based on which citation. I didn’t see “dog” in your text. To help understand the complexity of an LLM consider that these models typically hold about 10,000 less parameters than the total characters in the training data. If one wants to ins…

You mean 10,000x less parameters? In other words, only 1 character for every 10,000 characters of input?

Yeah, good luck embedding citations into that. Everyone here saying it's easy needs to go earn their 7 figure comp at an AI company instead of wasting their time educating us dummies.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#904

Earlier quoted context omitted.

The reproduction of that material in an educational setting is protected by Fair Use.

I don't think that is relevant to my comment. Whether the material is purchased, borrowed from a library, or legally reproduced under "fair use", I'm still asserting that I don't "owe" the creators any of my profit that I earn from taking advantage of what I learned.

The parent comment asks whether an “engine” trained on copyrighted data is entitled to decide profit. Your comment is about a human receiving knowledge that facilitates profit. Of course, these are legally independent scenarios.

Take a college student who scans all her textbooks, relying on fair use. If she is the only user, is she obligated to pay a premium for mining?

What about the scenario in which she sells that engine to other book owners? What if they only owned the book a short time in school?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#905
post #93

Earlier quoted context omitted.

What? This is about whether one country wants to cede a massive economic advantage to another country.

On the other hand, you could also argue that if AI takes all financial incentives from professionals to produce original works, then the AI will lose out on quality material to train on and become worse. Unless your argument is there’s no need for anything else created by humanity, everything worth reading has already been written, and humanity has peaked and everyone should stop? Like all things, it’s about finding…

The whole "AI training blackhole" thing is a myth. As long as humans are curating the content generated by ML, the content generated is still valid training data. Remember, for every ML generated image you see online, someone had to go through countless attempts to get it to create exactly what they wanted.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#906
> For example, in 2019, The Times published a Pulitzer-prize winning, five-part series on predatory lending in New York City’s taxi industry. The 18-month investigation included 600 interviews, more than 100 records requests, large-scale data analysis, and the review of thousands of pages of internal bank records and other documents, and ultimately led to criminal probes and the enactment of new laws to prevent future abuse.

> OpenAI had no role in the creation of this content, yet with minimal prompting, will recite large portions of it verbatim.

This is the smoking gun. GPT-4 is a large model and hence highly likely to reproduce content. They have many such examples in the court filing https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...

IANAL but that's a slam dunk of copyright violation.

NYT will likely win.

Also why OpenAI should not go YOLO scaling up to GPT-5 which will likely recite more copyrighted content. More parameters, more memorization.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#907
post #764

Earlier quoted context omitted.

And god willing if there is any justice in the courts NYTimes will lose this frivolous lawsuit. Copyright law is a prehistoric and corrupt system that has been about protecting the profit margins of Disney and Warner Bros rather than protecting real art and science for living memory. Unless copy/paste superhero movies are your definition of art I suppose. Unfortunately it seems like judges and the general public are…

> Copyright law is a prehistoric and corrupt system that has been about protecting the profit margins of Disney and Warner Bros rather than protecting real art These types of arguments miss the mark entirely imho. First and foremost, not every instance of copyrighted creation involves a giant corporation. Second, what you are arguing against is the unfair leverage corporations have when negotiating a deal with a risi…

[dead]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#908
post #707
post #600

Earlier quoted context omitted.

Really? Because the GPT-3 paper talks about "...two internet-based books corpora (Books1 and Books2)..." (see pages 8 and 9) - https://arxiv.org/pdf/2005.14165.pdf Unclear what that corpora might be, or if its the same books2 you are referring to.

My guess is that this poster meant books3, not books2. books1 and books2 are OpenAI corpuses that have never (to my knowledge) had their content revealed. books3 is public, developed outside of OpenAI and we know exactly what's in it.

sorry, books3 is indeed what I meant.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#909

Earlier quoted context omitted.

I don't think that is relevant to my comment. Whether the material is purchased, borrowed from a library, or legally reproduced under "fair use", I'm still asserting that I don't "owe" the creators any of my profit that I earn from taking advantage of what I learned.

The parent comment asks whether an “engine” trained on copyrighted data is entitled to decide profit. Your comment is about a human receiving knowledge that facilitates profit. Of course, these are legally independent scenarios. Take a college student who scans all her textbooks, relying on fair use. If she is the only user, is she obligated to pay a premium for mining? What about the scenario in which she sells that…

I agree that they are different scenarios that may lead to different legal frameworks. My point though was that asserting that the "profit" motive is sufficient to conclude something is owed to the creators is faulty logic. Individuals can generate profit from what they learn and we don't generally require them to share their profits with the creators of copyrighted material that they used to educate themselves.
Post reply on HN