Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

451–460 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#451
post #411

Earlier quoted context omitted.

A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights. The issue of replicating a style is probably more difficult.

> Presumably, if a passage of any significant length is cited verbatim (or almost verbatim), there would have been a way to track that source through the weights.

Figure this out and you get to choose which AI lab you want to make seven figures at. It's a really difficult problem.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#452

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

I mean maybe not the single most important development, but definitely a very important technological development with the potential to revolutionize multiple industries

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#453
post #399
post #366

Earlier quoted context omitted.

At the root, it seems like there's also a gap in copyright with respect to AI around transformative. Is using something, in its entirety, as a tiny bit of a massive data set, in order to produce something novel... infringing? That's a pretty weird question that never existed when copyright was defined.

Replace the AI model by a human, and it should become pretty clear what is allowed and what isn’t, in terms of published output. The issue is that an AI model is like a human that you can force to produce copyright-infringing output, or at least where you have little control over whether the output is copyright-infringing or not.

Its less clear than you think, and comes down more on how OpenAI is commercially benefiting and competiting with NYT than what they actually did. (See four factors of fair use)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#454

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#455

Earlier quoted context omitted.

If LLMs actually create added value and don't just burn VC money then they should be able to pay a fair price for the work of people they're relying upon. If your business is profitable only when you get your raw materials for free it's not a very good business.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

That is the case. It's just that the fair price is fairly low and is often covered by the government in the name of the greater good.

When for-profit companies seek access to library material they pay a much much higher price.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#457
post #345

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

It matters what ends up being best for humanity, and I think there are cases to be made both ways on this

People often get buried in the weeds about the purpose of copyright. Let us not forget that the only reason copyright laws exist is

> To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries

If copyright is starting to impede rather than promote progress, then it needs to change to remain constitutional.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#458
post #408
post #314

Earlier quoted context omitted.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

I would note that in the examples the NYT cites, the prompts explicitly ask for the reproduction of content. I think it makes sense to hold model makers responsible when their tools make infringement too easy to do or possible to do accidentally. However that is a far cry from requiring a little longer license to do the trainint in the first place.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#459

Earlier quoted context omitted.

A neural net is not a database where the original source is sitting somewhere in an obvious place with a reference. A neural net is a black box of functions that have been automatically fit to the training data. There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net.

> There is no way to know what sources have been memorized vs which have made their mark by affecting other types of functions in the neural net. But if it's possible for the neural net to memorize passages of text then surely it could also memorize where it got those passages of text from. Perhaps not with today's exact models and technology, but if it was a requirement then someone would figure out a way to do it.

Except it doesn’t memorize text. It generates text that is statistically likely. Generating a citation that is statistically likely wouldn’t really help the problem.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#460
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

We all remember when Aaron Swartz got hit with a wire tapping and intent to distribute federal crime for downloading JSTR stuff right? It's really disgusting, IMO, that corporations that go above and beyond that sort of behavior are seeing NO federal investigations for this sort of behavior. Yet a private citizen does it and it's threats of life in prison. This isn't new, but it speaks to a major hole in our legal sy…

Circumventing computer security to copy items en masse to distribute wholesale without transformation is a far cry from reading data on public facing web pages.
Post reply on HN