Live data from Hacker News

Fear of AI just killed a useful tool

techdirt.com

281–290 of 309 posts

Re: Fear of AI just killed a useful tool

#281

I'm not a lawyer, but neither is this Mike guy. I'm quite suspicious about how confident he is in stating that all of this is legally fine. From the creator himself: "When I ran out of books on my own shelves, I looked to the internet for more text that I could analyze, and I used web crawlers to find more books." I'm annoyed by the phrasing. Because you're obscuring that you pirated commercial books. A commercial bo…

Let's split the discussion though.

The authors clearly didn't support the idea of the project, regardless of the source for the data.

How the developer acquired the data is a different discussion. Unless you have clear proof they pirated the content, why disparage them?

Re: Fear of AI just killed a useful tool

#282

I just think it's really funny that a third or so of the article is the author struggling to figure out why this would be useful to anyone. > scanned and analyzed a whole bunch of books and would let you call up really useful data on books. [...] Frankly, all of that sounds amazing. And amazingly useful. Even more amazing is that he built it, and it worked. It would produce useful analysis of books. > This is all qui…

Imagine seeing the trends in books throughout certain time period, wars, etc. Or larger trends over the history of all written works. Or all kinds of other neat and useful information that can impact decisions we make today. Do you think all analysis of things you aren't personally interested in is useless?

Re: Fear of AI just killed a useful tool

#283
post #278

Earlier quoted context omitted.

This project was not generative AI. Comments are saying this project, which is not at all similar to generative ai, seemed to be okay. But you keep replying to say essentially “but if it was generative ai then authors have a legitimate reason to be angry”. There is no need to shoehorn that debate into this particular situation, and I see no merit in defending authors that had a knee jerk reaction to this project on t…

I think it is not completely off topic. Here is how I see it: Engineers tend to globally think that LLMs are not really a problem for copyright holders. At least those who develop LLMs pretty clearly don't give a damn. And on top of that, it is in their interest to not be constrained by copyrights. If this is my feeling (that engineers globally don't care about copyright holders), then it seems reasonable to me that…

To rephrase in my own understanding of what you wrote:

1) Some engineers (or more broadly, software developers) do not respect copyright

2) Therefore you reasonably are skeptical of projects related to material under copyright.

3) It is not always obvious if a project is respectful of copyright.

Now, applying these #1,#2,#3 you believe they justify the outrage for this particular project.

I disagree, because outrage combined with a lack of understanding (#3) is pretty much my definition of a knee-jerk reaction and vastly counterproductive to the interests of copyright holders because it will make the dismissiveness you predict a self-fulfilling prophecy.

Re: Fear of AI just killed a useful tool

#284

Earlier quoted context omitted.

> However, the apology [0] says that the creator did not "intend" to participate in AI that can "create zero-effort impersonations of artists." I'm not sure if the wording is unintentionally vague, or if there is some way his project could be used in that way. This seems FUDdy. "Intend" isn't in the apology at all, and the wording that is there says clearly that generative AI came after prosecraft, so there's no way…

I apologize for the quotes around intend. I wrote it without, then I forgot it was a paraphrase and added them back again. Unfortunately, I cannot edit my comment to fix that. I do think “intend” is a reasonable paraphrase of “never wanted to.” (Edited to add) I don’t think prosecraft was a finished project and he was definitely still working on his other tool for writers that incorporates some of the same tools. > T…

So you can edit your OP and this comment but can’t edit “intend”?

Re: Fear of AI just killed a useful tool

#285

Earlier quoted context omitted.

I apologize for the quotes around intend. I wrote it without, then I forgot it was a paraphrase and added them back again. Unfortunately, I cannot edit my comment to fix that. I do think “intend” is a reasonable paraphrase of “never wanted to.” (Edited to add) I don’t think prosecraft was a finished project and he was definitely still working on his other tool for writers that incorporates some of the same tools. > T…

So you can edit your OP and this comment but can’t edit “intend”?

HN comments cannot be edited once two hours have past.

Re: Fear of AI just killed a useful tool

#286

Earlier quoted context omitted.

The thing with this project is that it had no conceivable way of threatening the original authors, financially or otherwise. The analyses it produced weren't replacements of the original works, and served as a writing tool, not as something that generated new content. In this case, the shutdown seemed to hinge entirely on the irrational fear of anything with "AI" written on it and an unshakeable conviction that this…

>> had no conceivable way of threatening the original authors, financially or otherwise how so? what's inconceivable about it? >> seemed to hinge entirely on the irrational fear how are the authors "fearing without reason" or "illogically fearing" ?

> how so? what's inconceivable about it?

Authors make money through sales of their work. This tool was a writing aid that analyzed text and included some copyrighted works in its dataset. There was no way of retrieving these books in full, and the excerpts that were allegedly shown to the users were used in an analytical context, unlike the original works. So, this website couldn't replace ownership of the actual book for readers, and had no capacity to hurt their sales. Basically, the way the service used this data should be transformative enough as to not have any impact on the authors.

> how are the authors "fearing without reason" or "illogically fearing" ?

I called it "fear" because there was no strong argument on the authors' side as to why this tool is bad. I called it "illogical" because I think that it's no coincidence that this controversy only came up now, in 2023. Back in 2017 and onward, the existence of this tool didn't appear to generate pushback. My pet theory is that in 2023, now that we have good generative AI, the advancements have spawned an entire subset of people that view anything "AI" as inherently tainted and immoral. The original complaints seem to lack understanding of this tool and multiple people have conflated it with generative AI, despite it having nothing to do with that.

Re: Fear of AI just killed a useful tool

#287
post #79

Earlier quoted context omitted.

"Fair use" only applies to instances of copying / redistributing . The hint is in the name: copy -right. There's a notion, which seems to have taken off among creators who are paranoid about AI eating their livelihoods (which it might eat a chunk of) that copyright prevents people from doing anything with works they [legally] acquired other than personally read, listen, or watch it. That's not how copyright, as it ha…

Say I make a tool where you can enter the title of a book, and get the full text of the book without paying for it. I assume we all agree that would be illegal, right? Now say that instead of distributing that tool as an executable , I distributed it as a library . It would contain all the same books as the illegal executable above, but some developers would need to write an actual executable that would use the libra…

> Say I make a tool where you can enter the title of a book, and get the full text of the book without paying for it.

Let me introduce you to the Library of Babel[1].

But you need to know the hex! you complain. But that's basically how all of the "AI outputs copyrighted works!!!" gimmicks work. They're impractical unless you know exactly what you want it to reproduce. You can't just casually pick up a copy of Harry Potter like you would in a real library.

So is the Library of Babel illegal? What's the difference?

[1]: https://libraryofbabel.info/browse.cgi

Re: Fear of AI just killed a useful tool

#288
post #116

Earlier quoted context omitted.

Even Facebook's Llama was trained on books3, a dump of pirated books.

It's so mind blowing to me that it made it past corporate legal. I don't get what defense there could be besides "lmao try and stop me, nerds"

Fair use is basically the whole defense.

Re: Fear of AI just killed a useful tool

#289
post #123

Earlier quoted context omitted.

> If you want to do this kind of thing, let authors opt-in (or publishers). If it's fair use, why should you have to do that? The same copyright law protecting author's ownership rights over their art also provide "fair use" to other people. Someone may disagree with current fair use laws (and I suspect many outraged here do not), but that's a broader issue not related to this particular tool. It just 100% seems like…

> also provide "fair use" to other people "How much of someone else's work can I use without getting permission? Under the fair use doctrine of the U.S. copyright statute, it is permissible to use limited portions of a work including quotes, for purposes such as commentary, criticism, news reporting, and scholarly reports." https://www.copyright.gov/help/faq/faq-fairuse.html Limited portions, not the entire work.

Limited potions can be reproduced In the derived work you are distributing. Summaries and statistics of the work are almost certainly fair use.

Re: Fear of AI just killed a useful tool

#290
post #191

Earlier quoted context omitted.

I do not buy the premisses. Your example states that it can provide the full text of any book. An LLM cannot do that. They can produce something in the same style and setting. When an actual human author mimics other writers styles, it is not illegal, why exactly should it be illegal for an author to use an LLM to do it?

> When an actual human author mimics other writers styles, it is not illegal, why exactly should it be illegal for an author to use an LLM to do it? There is a fundamental difference of scale. Say I write a blog post about some technical thing I know. You read it, learn from it (and other sources), and then you write your own blog post with your understanding. You may link to my post (if you believe it is heavily ins…

I do not disagree with any of what you wrote. That is also an entirely different line of reasoning than your first argument.

That said, LLMs today, cannot do this in a meaningful way. If an author cannot write a better book than ChatGPT, then that author would not be able to live of their writing anyway. And the authors that use ChatGPT to write a book, but still put the effort into fine tuning it, will not be able to this at scale. You also need someone to line out the plot and twists and turns if it is to be a full length book.

Let's assume that in 10 years, LLMs are at the point where you cannot distinguish between a well written book by an author and one generated by entirely by an AI. Suppose we have two authors, one that is long dead and whos works are public domain and a young one that is just starting. An LLM trained on the author whos work is in the public domain can generate books that is just like the original works. But what if the young author writes in a similar way, is that now legal or illegal to generate the same content? It's impossible to know if the young authors work have been used for training.

My take on it, is that LLMs are pretty stupid. They cannot come up with new and novel things. So if a writer writes something that is different (i.e. new and novel), how do we protect that? We cannot prevent it from being used for training, so the next logical step is to protect it the same way as we protect technology with patents. But that come with its own class of problems, say if two people write the same way independently, only one can have the right to do it. That is not the solution either.

I do not have the answer, but I am certain that trying to ban LLMs, or dictating what and how is not the answer. Perhaps the authors that can write in a new and novel way, and knows how to use AI will proliferate because they embrace it.

Post reply on HN