Live data from Hacker News

Fear of AI just killed a useful tool

techdirt.com

231–240 of 309 posts

Re: Fear of AI just killed a useful tool

#231

open source it on IPFS and walk away, maybe call it something else and just drop the hashes on 4chan anonymously this reaction deserves the inevitable, as his only problem was it being attributed to him

How can he do it anonymously when we already know he's the only one who has it?

Lie and claim to have been hacked.

Re: Fear of AI just killed a useful tool

#232
post #206

Earlier quoted context omitted.

Why didn't you just say that, instead of posing a hypothetical about software that may itself contain full book text which can be used to display (in this case fair-use) passages to end users? lol I think the disconnect between your point of view and mine is that I see "training an LLM on copyrighted text" the same as a person reading copyrighted text, which is perfectly legal. And I see violating copyright as a pers…

> Why didn't you just say that, instead of Because I answered to a post that was talking about drawing the line for fair use. I just shared my view of how I see it. To me, OpenAI should be responsible for not giving copyrighted material to users if they are not allowed to do it. This means that they should be sued every single time someone manages to extract what is considered as copyrighted material from their softw…

> This means that OpenAI should be sued every single time someone manages to extract what is considered as copyrighted material from their software.

I agree! If GPT4 is outputting copyrighted material beyond what is considered fair-use (i.e. substantively more than what is provided by say, google books), I agree that is copyright infringement.

Indeed it is about the output, and making stuff available that people would otherwise have to pay for (or more precisely, enough of the copyrighted work that a person would have reason not to pay for the original work, causing a material loss to the original author) - that is a fineable violation imo.

Something else to think about... I work in biotech and have published articles in scientific journals on cellular and molecular level disease sequelae (such articles are also protected by copyright). Models trained on scientific literature are now being used for novel drug discovery and disease treatment pathways. These models are already outputting suggestions that seem very promising. Shall we also not provide these models access to the full corpus of scientific literature? It would significantly handicap these models to not have access to copyrighted scientific works. On one hand, some proportion of researchers will retain their jobs that would have otherwise been outsourced to LLMs (perhaps even myself). On the other hand, some amount of future patients will suffer or die from a disease that would have otherwise been cured.

Re: Fear of AI just killed a useful tool

#233
post #67

If you want to do this kind of thing, let authors opt-in (or publishers). Yes, it will take effort and probably go slow, but if the tool is really useful and amazing, it should be doable. I suspect the authors are put-off by a couple things: - the text of the works scanned seems like it may be from pirated sources. That poisons the project, no matter what it does with the scans, for many authors. - the use of these s…

Because the authors were AI-fearful luddites. From "Book" to "Program that judges books" lies well beyond any argument that the use of the derivative work could supersede the original. It's such clear cut transformative use that the authors come across as grossly misinformed about copyright law as a whole.

Perhaps there is an argument for generative AI possibly superseding the original, in that people might start asking an AI to generate them stories "in the style of x" instead of buying the author's books, but this wasn't that. It was just some fun data analysis of books.

Re: Fear of AI just killed a useful tool

#234
post #194

Earlier quoted context omitted.

> It is not their job to learn how the black box works If you have not learned the basics of how something works, you have no right for your opinion on it to be considered valid. Period. Invalid opinions do harm to democracy and endanger our way of life.

> you have no right for your opinion on it to be considered valid. Period. That is so wrong it is actually dangerous. Do I need to understand how a nuclear bomb works for my opinion on it to be considered valid? Obviously not. I only need to understand the consequences of it. It does not matter at all how it works, if I am against the fact that it will kill a whole lot of people. > Invalid opinions do harm to democra…

You can of course have any opinion you want. But this is not just about the authors having an opinion. It's about them starting a harassment campaign based on just faulty facts and making no attempt at verifying them.

If we work from the nuclear bomb analogy, you certainly don't need to be a nuclear physicist to protest nuclear bombs. You just need to have some a reasonably correct high level understanding of the impact of a nuclear bomb. But that's not what is happening here. This is more like storming the Belgian embassy to stop Belgium from using their nuclear arsenal to trigger a chain reaction in the atmosphere: totally detached from reality in every aspect.

As far as I can tell from your messages on this, you think that the harassment was entirely justified. Is that correct?

Re: Fear of AI just killed a useful tool

#235

Earlier quoted context omitted.

While I would agree in theory that a project like this would be best with opt-in, in reality that would just not work. Publishers would never opt-in to it, if they even respond to your requests at all.

Then don't do it? Or, if you do it, do it privately and don't share it on the internet? I'm not sure why this is a difficult idea; if asking for something and getting permission to do it is so difficult that 'would just not work. Publishers would never opt-in to it' ...then, it seems really obvious that even if you want to do it, technically can do it and you could maybe make a legal argument to doing it doesn't viol…

Should I need the publisher's permission to write a review of a book? Personally, I find that idea abhorrent. This sounds like an interesting project, unambiguously protected under fair use doctrine, both as analysis and as transformative, and the authors got their knickers in a twist because they are scared of that which they do not understand.

Re: Fear of AI just killed a useful tool

#236
post #144
post #61

My summary of the case: Someone did statistical analysis of a bunch of texts and created a tool that evaluates your text according to the developed model. Writers accused him of plagiarizing/using the content of their works.

As an aside, this would be completely legal in Japan, as classification and statistical analysis are protected as fair-use. I wonder if similar language exists in other copyright systems, but I would imagine it is likely the opposite...

This was unambiguously fair use under American copyright law, too.

Re: Fear of AI just killed a useful tool

#237
post #175

Copyright issues aside. My personal opinion is that these tools are mostly useful to suck the "soul" out of a book. They give you templates and stuff and useful statistics to help you go to the lowest common denominator. The problem is more visible in the movie industry, where they have had script templates for a hundred years now (actual time interval pulled out of a*), but it's starting to show up in books too. For…

I wonder about a parallel for paintings. What if there was an analysis stating exactly the brushes the painter used, the number of strokes, the exact pigments, etc? Would that, in your opinion, "suck the 'soul' out of a painting"? I could see this as a brilliant learning tool. A tool to provide deep insight into something that would be very challenging to quantify personally. I think all this would make future author…

The cave paintings in France have been studied this way, starting with Leroi-Gourhan's work and then accelerating with the use of computers. It's defi itely shed some important light on the artists who made them tens of thousands of years ago, and I don't think it made the paintings any less wonderful.

Re: Fear of AI just killed a useful tool

#238
post #123

Earlier quoted context omitted.

> If you want to do this kind of thing, let authors opt-in (or publishers). If it's fair use, why should you have to do that? The same copyright law protecting author's ownership rights over their art also provide "fair use" to other people. Someone may disagree with current fair use laws (and I suspect many outraged here do not), but that's a broader issue not related to this particular tool. It just 100% seems like…

> also provide "fair use" to other people "How much of someone else's work can I use without getting permission? Under the fair use doctrine of the U.S. copyright statute, it is permissible to use limited portions of a work including quotes, for purposes such as commentary, criticism, news reporting, and scholarly reports." https://www.copyright.gov/help/faq/faq-fairuse.html Limited portions, not the entire work.

Copyright pertains to reproduction of the work. The statistics this tool provided are not reproductions at all. It did also provide quotes, which were not extensive and certainly not the entire work.

Re: Fear of AI just killed a useful tool

#239
post #140

Earlier quoted context omitted.

That's twitter generally. If your engagement only reaches the level of twitter, you aren't really engaging at all.

So as long as that's all the engagement there is, we're free to ignore it and carry on, correct?

I would think so. If someone is shouting & stomping their feet in the public town square about my project, but I never go anywhere near the town square anyway, I don’t think I’m going to shutdown my project. It’s just too bad the person who created this tool happened to walk through the town square.

Re: Fear of AI just killed a useful tool

#240
post #116

I found this article frustratingly vague on how prosecraft.io actually worked. As far as I can tell, the author scraped the web for books, including in-copyright books. Then he analyzed it with techniques based on "classical" natural language processing techniques, rather than transformers or deep learning. He appears to have retained the books he scraped for future analysis. The site itself seems to use only snippet…

Even Facebook's Llama was trained on books3, a dump of pirated books.

It's so mind blowing to me that it made it past corporate legal. I don't get what defense there could be besides "lmao try and stop me, nerds"
Post reply on HN