Live data from Hacker News

AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

tvnewscheck.com

1–10 of 17 posts

Re: AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

#5
> Ultimately, two things are becoming clear: AI needs journalism to thrive, and journalism needs to find a sustainable way to coexist with AI

Why would AI need journalism to thrive? This seems just thrown in there, I bet journalists would want that to be true but I don't see anything here stating why.

Re: AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

#7
post #5

> Ultimately, two things are becoming clear: AI needs journalism to thrive, and journalism needs to find a sustainable way to coexist with AI Why would AI need journalism to thrive? This seems just thrown in there, I bet journalists would want that to be true but I don't see anything here stating why.

Because after the initial scraping where the AI for example can tell you about each American president, users want the AI updated to know if any present is a convicted criminal (for example).

When we were all naive the AI bots came through and just took our content, now as everyone wises up, folks are laying down a cost for the next round of web scraping.

Imagine your bot could only scrape once a year.. your AI would quickly get out of date on many topical issues. This ‘latest updates gap’ is where the journalists and publishers see leverage.

Re: AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

#8
post #3

What about small publishers? Can they change their license to fight back the AI data gobblers? A watermark is easy to implement and prove if your content gets picked up in the next training set. Or is this something that is not possible?

AI companies have historically not given two figs about licenses or copyright. The law's murky enough right now that they can claim it's outside of the protections offered by copyright, so a publisher's only real recourse is to either poison the articles for AI consumers or block their crawler's access entirely.

Re: AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

#9
post #5

> Ultimately, two things are becoming clear: AI needs journalism to thrive, and journalism needs to find a sustainable way to coexist with AI Why would AI need journalism to thrive? This seems just thrown in there, I bet journalists would want that to be true but I don't see anything here stating why.

AI is not a primary source. Journalists (should) be doing more than just collecting data from online sources.

Re: AI's Desperate Hunger for News Training Data Has Publishers Fighting Back

#10
post #3

What about small publishers? Can they change their license to fight back the AI data gobblers? A watermark is easy to implement and prove if your content gets picked up in the next training set. Or is this something that is not possible?

> A watermark is easy to implement and prove if your content gets picked up in the next training set.

Is it? How would you watermark raw text? Images maybe, but I'm skeptical even there.

My high-school cousins tell me kids use one AI to write these days, and another to rewrite it to avoid AI detectors. I view fighting against this as a Sisyphean task.

Post reply on HN