Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

491–500 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#491
post #452

Earlier quoted context omitted.

I mean maybe not the single most important development, but definitely a very important technological development with the potential to revolutionize multiple industries

Can I ask what industries with what application? I've seen lots of task like summarizing articles or producing text. The image and video work seems too rudimentary to be taken seriously. Is there something out there that seems like a killer application? I was amazed at the idea of the block chain but we never found a use for it outside of cryptocurrency. I see a similariy with AI hype.

Well front page of HN right now is an article about how AI aided in the development of a new antibiotic

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#492
post #344

I hate that this will likely be decided by a 75 years old judge that hasn't been close to a computer since getting his 15 years old grandson to fix his patience game

Yeah the whole thing with Zucc's trial highlighted this perfectly for me, a bunch of clueless old dolts who have no fucking idea how anything in the modern age works.

I'm eagerly awaiting the time where the people making these decisions at least have some sort of baseline level of understanding, otherwise these psychopathic megacorps will keep getting away with things based on technicalities and the judge's lack of knowledge.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#493

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

What do you actually believe, with that statement? Do you believe Libraries are operating illegally? That they aren't paying rightsholders?

Also: GPT is not a legal entity in the united states. Humans have different rights than computer software. You are legally allowed to borrow books from the library. You are legally allowed to recite the content you read. You're not allowed to sell verbatim recitation of what you read. This is, obvious, I think? But its exactly what LLMs are doing right now.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#494
post #375

Earlier quoted context omitted.

Can you imagine spending decades of your life, studying skin cancer, only to have some $20/month ChatGPT index your latest findings and spit out generically to some subpar researcher: "Here's how I would cure melanoma!" followed by your detailed findings. Zero mention of you. F-that. Attribution, as best they can, is the least OpenAI can do as a service to humanity. It's a nod to all content creators that they have b…

Can you imagine spending decades of your life studying antibiotics, only to have an AI graph neural network beat you to the punch by conceiving an entire new class of antibiotics (first in 60 years) and then getting published in Nature. https://www.nature.com/articles/d41586-023-03668-1

It looks like the published paper managed to include plenty of citations.

https://dspace.mit.edu/handle/1721.1/153216

As it should be.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#495
post #331
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

ChatGPTs birth as a research preview may have been an attempt to avoid these issues. It would have been unlikely to trigger legal anger for a free product which few use. When usage exploded, the natural inclination would be to hope for the best.

Google may simply have been obliged to follow suit.

Personally, I’m looking forward to pirate LLMs trained on academic content.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#496

Earlier quoted context omitted.

Well yeah, they’re being sued. They move very quickly to stop any obvious copyright violation paths.

And in a lawsuit, there's very much the question of intent as well. If OpenAI never meant to allow copyrighted material to be reproduced, shut it down immediately when it was discovered, and the NYT can't show any measurable level of harm (e.g. nobody was unsubscribing from NYT because of ChatGPT)... then the NYT may have a very hard time winning this suit based specifically on the copyright argument.

Intent isn't some magic way to claim innocence. Here negligence is very much at play. Were OpenAI negligent when they made the NYT articles available like this?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#497
post #229
post #189

Earlier quoted context omitted.

The ChatGPT subscription is more valuable because it's built on the theft of the NYT content and many other authors' work.

No. It is a technology that can massively accelerate human progress—I don’t buy the theory that OpenAI used NYTimes content that was not freely available to them. If you read all of the public internet you probably have lots of snippets of NYTimes articles. Regarding the reading of the wirecutter by a browser tool, I don’t know how much of it is available online without subscription (because I subscribe to the NYTime…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#498
post #367

Earlier quoted context omitted.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures. I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a cha…

It’s disingenuous to frame using data to train a model as a “view,” of that data. The simple cases are the easy ones, if ChatGPT completely rips a NYT article then that’s obviously infringement; however, there’s an argument to be made that every part of the LLM training dataset is, in part, used in every output of that LLM. I don’t know the solution, but I don’t like the idea that anything I post online that is openl…

All I can ever think about with how ML models work is that they sound an awful lot like Data Laundering schemes.

You can get basically-but-not-quite-exactly the copyrighted material that it was trained on.

Saw this a lot with some earlier image models where you could type in an artists name and get their work back.

The fact that AI models are having to put up guardrails to prevent that sort of use is a good sign that they weren't trained ethically and they should be paying a ton of licensing fees to the people whose content they used without permission.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#499
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

It’s likely fair use.

I think we need a lot of clarity here. I think it's perfectly sensible to look at gigantic corpuses of high quality literature as being something society would want to be fair use for training an LLM to better understand and produce more correct writing... but the actual information contained in NYT articles should probably be controlled primarily by NYT. If the value a business delivers (in this case the information of the articles) can be freely poached without limitation by competitors then that business can't afford to actually invest in delivering a quality product.

As a counter argument it might be reasonable to instead say that the NYT delivers "current information" so perhaps it'd be fair to train your model on articles so long as they aren't too recent... but I think a lot of the information that the NYT now relies on for actual traffic is their non-temporal stuff - including things like life advice and recipes.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#500
post #486
post #440

Earlier quoted context omitted.

robots.txt is not meant to be a mechanism of communicating the licensing of content on the page being crawled nor is it meant to communicate how the crawled content is allowed to be used by the crawler. Edit: same applies to humans. Just because a healthcare company puts up a S3 bucket with patient health data with “robots: *” doesn’t give you a right to view or use the crawled patient data. In fact, redistributing i…

Furthering the S3 health data thought exercise: If OpenAI got their hands on an S3 bucket from Aetna (or any major insurer) with full and complete health records on every American, due to Aetna lacking security or leaking a S3 bucket, should OpenAI or any other LLM provider be allowed to use the data in its training even if they strip out patient names before feeding it into training? The difference between this ques…

HIPAA is about more than just names. Just information such as a patient's ZIP code and full medical history is often enough to de-anonymise someone. HIPAA breaches are considered much more severe than intellectual property infringements. I think the main reason that patients are considered to have ownership of even anonymised versions of their data (in terms of controlling how it is used) is that attempted anonymisation can fail, and there is always a risk of being deanonymised.

If somehow it could be proven without doubt that deanonymising that data wasn't possible (which cannot be done), then the harm probably wouldn't be very big aside from just general data ownership concerns which are already being discussed.

Post reply on HN