Live data from Hacker News

Fear of AI just killed a useful tool

techdirt.com

51–60 of 309 posts

Re: Fear of AI just killed a useful tool

#51
post #4

This is an incredibly biased article, hinging entirely on the assumption that AI training is fair use.

"Fair use" only applies to instances of copying / redistributing. The hint is in the name: copy-right.

There's a notion, which seems to have taken off among creators who are paranoid about AI eating their livelihoods (which it might eat a chunk of) that copyright prevents people from doing anything with works they [legally] acquired other than personally read, listen, or watch it.

That's not how copyright, as it has existed in the past, works. You can do all the algorithmic processing of your ebook collection that you want. You might be able to display small portions of a book to others, depending on the situation.

Quoting one or two paragraphs out of an entire book seems like reasonably safe fair use, but that won't stop a copyright-maximalist creator (or their publisher) from suing you, and won't stop some copyright-maximalist judge from ruling against you, so it's probably best to minimize the amount of content from a book that you redisplay directly. But you can do all the analysis and statistics generation you want, and display those results to others.

It remains to be seen what judges will do with AI generation of works based on ingesting gigantic amounts of copyrighted work. The entire framework of copyright is going to be broken, and until Congress steps in and changes it, judges are going to go every which way. There's no bright line for 4-factor analysis; it's always been a gut-level "is this a reasonable use that doesn't impact commercial sales too much". There's no possible rational way to draw a line. AI models can generate a painting of a new subject only loosely in the style of a contemporary painter, which would not be copyright infringement, or it can generate a near-clone of an existing work with the right prompting, and depending on how clever the prompter is, a lot of intermediate stages of likeness. Who decides how close to an existing work is too close?

Re: Fear of AI just killed a useful tool

#52

For this particular example, the tool doesn't seem like it's a big deal. It just analyzes works for data. I'm not sure how this would be any different from a literary critic doing the same thing manually. In general, though, I think artists would be less hostile to technological innovations if the people imploring them to "figure out how to embrace the technology rather than fear it" weren't actively trying to destro…

I agree with you. There’s a pattern that I see a lot, of having:

1. large powerful players doing something not entirely helpful;

2. victims of that protesting that change vehemently; all that in vain because the players are powerful and have sheltered themselves from criticism, usually via lobbying;

3. regulatory capture or protests go after a smaller player, which is widely advertised to accuse 2. of going too far — even when the problem in 1. is still entirely there, and now ignored.

It’s definitely the case with globalization (large conglomerate benefit, people protest, and a small artisan who started selling abroad is featured being victimized by tariffs), fossil fuel (large oil extractor, climate advocate, farmer seeing fertilizer prize go up), immigration, American cultural hegemony, car dominance over cities, etc.

That pattern allows larger players still doing harm to wash their morals. I feel like we need better antibodies to say: No, this does not absolve them.

Re: Fear of AI just killed a useful tool

#53
I just think it's really funny that a third or so of the article is the author struggling to figure out why this would be useful to anyone.

> scanned and analyzed a whole bunch of books and would let you call up really useful data on books. [...] Frankly, all of that sounds amazing. And amazingly useful. Even more amazing is that he built it, and it worked. It would produce useful analysis of books.

> This is all quite interesting. It’s also the kind of thing that data scientists do on all kinds of work for useful purposes. Smith built Prosecraft into Shaxpir, again, making it a more useful tool.

Author's general illiteracy aside, he's really giving the game away here. I can't even think about the ethical implications of the project, because why would I care to count the number of adverbs and passive voice in all books ever, and why would you need a state of the art LLM-powered AI to do it?

Re: Fear of AI just killed a useful tool

#54
> The Gizmodo article has a ridiculously wrong “fair use” analysis, saying “Fair Use does not, by any stretch of the imagination, allow you to use an author’s entire copyrighted work without permission as a part of a data training program that feeds into your own ‘AI algorithm.’” Except… it almost certainly does? Again, we’ve gone through this with the Google Book scanning case, and the courts said that you can absolutely do that because it’s transformative.

I could be wrong but I'm pretty sure Fair Use doesn't mean you can download a dump of Library Genesis and feed that into your system!

Re: Fear of AI just killed a useful tool

#55

I’m with the artists on this one. Our obsession with converting everything into input for an algorithm that spits out an ill-defined number (what the hell is “vividness”?) needs to stop. We already tried this with human communication and gave birth to the dystopian nightmare that is social media, why keep repeating our mistakes?

That's silly. Humans review books all the time, using very similar words. Where's the outrage over that? This is manufactured, stretched, overhyped objections. I believe it's all as the OP suggests, because the word AI is in there. Not because anything illegal or immoral is going on. In fact it's a terribly useful tool, and once the mob cools off it'll likely return.

> That's silly. Humans review books all the time, using very similar words. Where's the outrage over that?

Easy: humans are not machines. "X does it all the time, so I should be able to do it" is never a valid conclusion. It depends on the situation.

> In fact it's a terribly useful tool, and once the mob cools off it'll likely return.

Maybe this tool in particular does not "abuse" the books. Maybe this tool in particular is terribly useful. But you can't blame authors and artists for taking a stance against those new algorithms that provably have the potential to automatically "steal" from their work. You can believe that asking ChatGPT to "write a novel in the style of X" is not abusing the copyright, that's fine. And the authors can answer that they fear it has the potential to break their source of revenue to a point where they won't want to publish anything anymore. And they are entitled to it. And maybe someday we come up with licenses that prevent the use as training data (how in the world could one conclude today that "it is most definitely fair use", given that this is a very new way of using IP material?).

Re: Fear of AI just killed a useful tool

#56

Earlier quoted context omitted.

That's silly. Humans review books all the time, using very similar words. Where's the outrage over that? This is manufactured, stretched, overhyped objections. I believe it's all as the OP suggests, because the word AI is in there. Not because anything illegal or immoral is going on. In fact it's a terribly useful tool, and once the mob cools off it'll likely return.

You are exactly modeling the chauvinistic Silicon Valley attitude that is causing the outrage in the general population to begin with. “Our algorithms are pretty much the same as human art criticism, so put down the pitchforks you unenlightened scum” is up there with telling them to eat (a Stable Diffusion generated picture of) cake.

People starved while it was suggested they eat cake. Not sure how that relates - are the rights around art crit not the same as AI crit?

Re: Fear of AI just killed a useful tool

#57
post #34

Earlier quoted context omitted.

What will hurt artists is, when in 10 years, all publishers are demanding that the vividness score (TM) be at least a 95% “because that’s what drives sales”. Which is what will happen if the authors don’t proactively stop it from happening. Look at how the music industry has evolved over time.

Or it could help me find terser books I like, people will still have preferences and if the author tries to pander to only the largest market segment I'd argue that's on them.

I think it’s much more likely you would get the book equivalent of crap SEO sites spammed out to satisfy numerical measures of quality.

Re: Fear of AI just killed a useful tool

#58
I found this article frustratingly vague on how prosecraft.io actually worked. As far as I can tell, the author scraped the web for books, including in-copyright books. Then he analyzed it with techniques based on "classical" natural language processing techniques, rather than transformers or deep learning. He appears to have retained the books he scraped for future analysis. The site itself seems to use only snippets.

However, the apology [0] says that the creator did not "intend" to participate in AI that can "create zero-effort impersonations of artists." I'm not sure if the wording is unintentionally vague, or if there is some way his project could be used in that way.

For what it's worth, the Computational Story Lab's hendometer [1] seems to have largely out-of-copyright books from Project Gutenberg, plus the Harry Potter series.

[0]: https://blog.shaxpir.com/taking-down-prosecraft-io-37e189797...

[1]: https://hedonometer.org/books/v3/863/

Edit: Apparently he was working on an LLM project. https://twitter.com/stealcase/status/1688721685585809408. It's unclear whether he was planning to use the books he scraped (although as @stealcase points out, GPT-Neox itself was trained on books that were pirated).

Re: Fear of AI just killed a useful tool

#59

I’m with the artists on this one. Our obsession with converting everything into input for an algorithm that spits out an ill-defined number (what the hell is “vividness”?) needs to stop. We already tried this with human communication and gave birth to the dystopian nightmare that is social media, why keep repeating our mistakes?

We repeat the mistakes because in the short term, someone finds it profitable, hence a prisoner's dilemma type situation.

If an AI tool was killed, I consider it a victory. That's because even if there are some small useful applications of AI, AI on the whole will certainly put most creatives out of business.

Instead, I propose the following: anyone who is interested in preventing AI from taking over their craft should join me in a coalition of ban AI from their own business. By placing a notice that your work is "100% AI FREE", you are doing something akin to the fair-trade/sustainably sourced sticker on chocolate or other food products: you are letting consumers know that your work was made by a human, so that they can support you.

If enough people get in on this, and pledge to support only those creators who don't use AI, then we can make AI an unprofitable venture and hopefully kill it forever!

I already put a 100% AI FREE badge on my YouTube channel, which means that I will never use AI for writing scripts, editing videos, producing images, etc. Moreover, I also pledge to support other creators who pledge never to use AI, by buying their products over others!

Re: Fear of AI just killed a useful tool

#60
post #18

Earlier quoted context omitted.

What do you even mean by this comment? Have you considered the possibility that people are smart in ways that you are not considering, rather than just labeling it “dumb”?

Have you considered ... in ways that you are not considering...? I am pretty confident they haven't. Sounds like you've set yourself up for a reverse "true scotsman" here ;)

Nice catch, thanks for pointing it out.
Post reply on HN