Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

861–870 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#861
post #708

Earlier quoted context omitted.

The world you’re hoping for will put all AI tech only within the hands of the established top 10 media entities, who traditionally have never compensated fairly anyway. Sorry but if that’s the alternative to some writers feeling slighted, I’ll choose for the writers to be sad and the tech to be free.

“Feeling slighted” is a gross understatement of how a lack of compensation flowing to creators has shaped the internet and the wider world over the past 25 years. If we have a problem with the way top media companies compensate their creators, that is a separate issue - not a justification for layering another issue on top.

I’m a creator myself and see the two futures ahead of me and free benefits me in the long term more than closed.

The tech can either run freely in a box under my desk or I’ll have to pay upwards of 15-20k a year to run it on Adobes/Google/etcs servers. Once the tech is locked up it will skyrocket to AutoCAD type pricing because the acceleration it provides is too much.

Journos can weep, small price to pay for the tech being free for us all.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#862

Earlier quoted context omitted.

So at first ChatGPT will copy journalists. Then journalists will stop working because nobody pays them. Then there will be no news. Some people may look at that situation and decide to start a new news business but that business will fail because ChatGPT will immediately rip it off. The end game is just no news other than volunteers.

> So at first ChatGPT will copy journalists. You still literally have not explained how this works. ChatGPT could write a news article, but it's not going to actively discover new social phenomena or interview people on the street. Niche journalism will continue having demand for the sole reason that AI can't reliably surface new and interesting content. So... again, how does a pre-trained transformer model scoop a j…

AI labs are working on and largely already have generative ai that can be actively updated. The generative ai scoops real journalists stories by watching their feed. This isn’t very different from the current status quo, it’s just a continuation of an already shitty situation for news organizations. If their revenue decreases even more than it already has they will cease to exist. Niche journalism barely has any demand today, it won’t take much more reduction in demand for it to not be worth the cost to produce. Just a few more people using big tech products as their news feed instead of the news organizations themselves is all it would take.

You can say that the people getting their news from the tech products will switch to paying news organizations in some way if the news starts to disappear but I highly doubt it seeing how people treat news today. And if it that does happen they’ll switch back again to the ai products as the centralization it can provide is valuable.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#863

Earlier quoted context omitted.

> I very much doubt the company [...] was selling the non-copyrighted software Well you'd be mistaken. Lately, it was custom software, for a particular client, and of no interest to others. Earlier, it was before software copyright was a thing, and computer manufacturers gave software away to sell the hardware. At the very beginning, yes, it was "very specific" hardware; it was Burroughs hardware, which used Burrough…

> Well you'd be mistaken. Lately, it was custom software, for a particular client, and of no interest to others. Earlier, it was before software copyright was a thing, and computer manufacturers gave software away to sell the hardware. Then I am not mistaken: the company was initially selling hardware, with the software being just a value add as you say (no copyright: no interest in trying to sell, exactly my point).…

> offered to take on the software maintenance for a much lower price

We didn't charge maintenance for this software. We would write it to close the sale of a computer. It was treated as "cost of sale". I'm sure it was cheaper (to us) than the various discounts and kickbacks that happened in big mainframe deals.

As far as Microsoft and Adobe is concerned, I wouldn't regard it as a misfortune if they had never existed. I'm not convinced that RedHat's existence is contingent on copyright.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#864

Earlier quoted context omitted.

I'm not for or against anything at this point until someone gets their balls out and clearly defines what copyright infringement means in this context. If you give a bunch of books to a kid all by the same author and then pay that kid to write a book in a similar style and then I go on to sell that book...have I somehow infringed copyright? The kids book at best is likely to be a very convincing facsimile of the orig…

There are two problems with the “kid” analogy: a) In many closely comparable scenarios, yes, it’s copyright infringement. When Francis Ford Coppola made The Godfather film, he couldn’t just be “inspired” by Puzo’s book. If the story or characters or dialog are similar enough, he has to pay Puzo, even if the work he created was quite different and not a literal “copy”. b) Training an LLM isn’t like giving someone a bo…

How is it a copy at all? Surely the model weights would therefore be much larger than the corpus of training data, which is not the case at all.

If it disgorges parts of NYT articles, how do we know this is not a common phrase, or the article isn't referenced verbatim on another, unpaid site?

I agree that if it uses the whole content of their articles for training, then NYT should get paid, but I'm not sure that they specifically trained on "paid NYT articles" as a topic, though I'm happy to be corrected.

I also think that companies and authors extremely overvalue the tiny fragments of their work in the huge pool of training data, I think there's a bit of a "main character" vibe going on.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#865
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

Why isn't robots.txt enough to enforce copyright etc? If NYT didn't set robots.txt properly, is their content free-for-all? Yes I know the first answer you would jump to is "of course not, copyright is the default", but it's almost 2024 and we have had robots.txt as industry de jure to stop crawling.

NYT seemed to claim paid subscriptions as well, which I'm not sure that bots can actually crawl.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#866
He he, people will wanna treat LLM learning different to their own learning.

I think it's fine, as long as it was fed publicly accessible content, without any payment or subscription then it's accessible to an LLM as it is to you and I and that's fair.

And for the people that screech about LLMs being different because they can mass produce derivative works; first of all, ALL works are derivative and if machine produced works are compelling enough to compete with human produced ones then clearly humans need to get better at it.

The automatic loom took over from weavers cause it was better, if it wasn't then people would still work as weavers.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#867

Earlier quoted context omitted.

[flagged]

It should link to one of the articles about TCP it used as a reference to write that info blurb, not the TCP spec. The problem is that those links doesn't link to where it got that text, it links to whatever that text linked to. Saying it is giving links is like saying that when I copy paste an article with links I am providing links to the source. No I am not, I am plagiarizing including plagiarizing those links. So…

Presumably it can't link them, because it's been train on the data, not built on top of it. Gpt model doesn't include the sum of all training data, that's not how machine learning works at all (and overfitting on such a large and diverse dataset would be a monumental fuck up)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#868

I have deeply mixed feelings about the way LLMs slurp up copyrighted content and regurgitate it as something "new." As a software developer who has dabbled in machine learning, it is exciting to see the field progress. But I am also an author with a large catalog of writings, and my work has been captured by at least one LLM (according to a tool that can allegedly detect these things). Overall, current LLMs remind me…

I don’t view LLMs as a fad. It’s like drummers and drum machines. Machines and drummers co-exist really well. I think drum machines, among other things, made drummers better.

Neither, and NYT editors use all sorts of productivity tools, inspiration, references, etc too. Same as artists will usually find a couple references of whatever they want to draw, or the style, etc.

I agree with the key point that paid content should be licensed to be used for training, but the general argument being made has just spiralled into luddism at people who are fearful that these models could eventually take their jobs; and they will, as machines have replaced humans in so many other industries, we all reap the rewards, and industrialisation isn't to blame for the 1%, our shitty flag waving vote for your team politics are to blame.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#869
post #417

Earlier quoted context omitted.

"probably the single most important development in human history" is the kind of hyperbole you'd only find here. Better than medicine, agriculture, electrification, or music? That point of view simply does not jive with what I see so far from AI. It has had little impact beyond filling the internet with low-effort content. I feel like the crypto evangelists never got off the hype train. They just picked a new destina…

Also the assumption a publication that’s been around for 150 years is disposable, not the web application that was created a year ago. I’ve been saying for a while that people’s credulity and impulse to believe absolutely any storyline related to technology is off the charts.

Been around for 150 years but I imagine the generations who it leans on are dying off. Nobody reads print media format anymore, we get our news elsewhere, for free and with varying political undertones, rather than the fixed one of a bought and paid for outlet.

Keep in mind these guys play both sides of every field they cover in their "news".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#870
post #693

Earlier quoted context omitted.

A human can't credit the source of each element of everything they've learnt. AI's can't either, and for the same reason. The knowledge gets distorted, blended, and reinterpreted a million ways by the time it's given as output. And the metadata (metaknowledge?) would be larger than the knowledge itself. The AI learnt every single concept it knows by reading online; including the structure of grammar, rules of logic,…

At the same time, there are situations where humans are expected to provide sources for their claims. If you talk about an event in the news, it would be normal for me to ask where you heard about it. 100% accuracy in providing a source wouldn’t be expected, but if you told me you had no idea, or told me something obviously nonsense, I would probably take what you said less seriously.

The raw technology behind it literally cannot do that.

The model is fuzzy, it's the learning part, it'll never follow the rules to the letter the same as humans fuck up all the time.

But a model trained to be literate and parse meaning could be provided with the hard data via a vector DB or similar, it can cite sources from there or as it finds them via the internet and tbf this is how they should've trained the model.

But in order to become literate, it needs to read...and us humans reuse phrases etc we've picked up all the time "as easy as pie" oops, copyright.

Post reply on HN