Live data from Hacker News

How can I prevent my site from being a free dataset for LLMs?

news.ycombinator.com

51–60 of 66 posts

Re: How can I prevent my site from being a free dataset for LLMs?

#51
post #45

Earlier quoted context omitted.

Attribution comes to mind.

Am I wrong when I don't attribute my understanding of words to the dictionary I read for any particular word?

No, but you're wrong when you use that argument in this situation.

Re: How can I prevent my site from being a free dataset for LLMs?

#52

Earlier quoted context omitted.

Am I wrong when I don't attribute my understanding of words to the dictionary I read for any particular word?

No, but you're wrong when you use that argument in this situation.

Can you convince me that's not equivalent to what LLMs do with their training sets? My understanding is that's a useful analogy?

Re: How can I prevent my site from being a free dataset for LLMs?

#53

Why don't you want it to be used as training data? You want visitors to be able to freely benefit from your work. What's wrong with AI also benefiting? Or more specifically the AI's eventual users?

I’ll give that a shot; - AI learns what’s on the site - Visitors stop coming to the site, because the information is now freely available in parsed/summarized form from the AI - Blogger stops posting because there are no visitors - AI stops learning because there is no new content

So, that's kinda what [search engine] does, isn't it? But there's no hand wringing about that anymore?

Re: How can I prevent my site from being a free dataset for LLMs?

#54

Earlier quoted context omitted.

No, but you're wrong when you use that argument in this situation.

Can you convince me that's not equivalent to what LLMs do with their training sets? My understanding is that's a useful analogy?

You can't plagiarize by copying a single word you learned. You can't plagiarize by learning ideas or common expressions and reusing them.

If you read copywritten material and then pass it off as your own you are plagiarizing. Words in a dictionary don't come under that, but I'd bet that if you released a new dictionary that was mostly copied from the old one, most people would consider that plagiarism as well.

Re: How can I prevent my site from being a free dataset for LLMs?

#55

Earlier quoted context omitted.

Can you convince me that's not equivalent to what LLMs do with their training sets? My understanding is that's a useful analogy?

You can't plagiarize by copying a single word you learned. You can't plagiarize by learning ideas or common expressions and reusing them. If you read copywritten material and then pass it off as your own you are plagiarizing. Words in a dictionary don't come under that, but I'd bet that if you released a new dictionary that was mostly copied from the old one, most people would consider that plagiarism as well.

I completely agree with this; but my understanding about how LLM work is that they don't copy meaningful segments of text from any specific source. Instead, they predict the next block of text, which they'd only do if they've seen that idea/sequence enough times with context to rank the prediction high enough.

I haven't seen any service copy out large block of text enough to make me think it's reasonable to call their output plagiarized.

Meaning, if the LLM I use will only repeat an idea that many someone's have written about, such that it's seen the idea, or parts of that idea many times. Why is that still plagiarism? Or rather, worthy of direct attribution? Or why was I wrong to use the argument about citing a dictionary here?

(I'm aware that a number of people are working on giving memory so AI can quote from pages like wikipedia. But I don't think it's fair to call that "training data")

Re: How can I prevent my site from being a free dataset for LLMs?

#56
post #49
post #35

Earlier quoted context omitted.

> I want my AI trained on everything "your" AI ? > If you don't want AI learning from your work, then don't publish it. If you don't want me stealing and reusing your licensed open source code don't make it public If you don't want me to steal your car don't park it on public roads See how dumb that is ?

> If you don't want me stealing and reusing your licensed open source code don't make it public A practical matter, larger point completely aside: a nonzero number of individuals and corps will indeed use licensed code internally if they come across it and they feel it helps their goals.

Oracle and Microsoft will send you a million dollar bill if you tried that with their products. You would be surprised by who rats out a company for a reward

Re: How can I prevent my site from being a free dataset for LLMs?

#57

Earlier quoted context omitted.

I’ll give that a shot; - AI learns what’s on the site - Visitors stop coming to the site, because the information is now freely available in parsed/summarized form from the AI - Blogger stops posting because there are no visitors - AI stops learning because there is no new content

So, that's kinda what [search engine] does, isn't it? But there's no hand wringing about that anymore?

People visit the site because they want more info if the summary isn't enough. No way to do that with chatGPT. This limitation probably means search engines are safe for now

Re: How can I prevent my site from being a free dataset for LLMs?

#58

Earlier quoted context omitted.

You can't plagiarize by copying a single word you learned. You can't plagiarize by learning ideas or common expressions and reusing them. If you read copywritten material and then pass it off as your own you are plagiarizing. Words in a dictionary don't come under that, but I'd bet that if you released a new dictionary that was mostly copied from the old one, most people would consider that plagiarism as well.

I completely agree with this; but my understanding about how LLM work is that they don't copy meaningful segments of text from any specific source. Instead, they predict the next block of text, which they'd only do if they've seen that idea/sequence enough times with context to rank the prediction high enough. I haven't seen any service copy out large block of text enough to make me think it's reasonable to call thei…

"use will only repeat an idea that many someone's have written about, such that it's seen the idea, or parts of that idea many times"

As you go narrower with a query only one source of truth is available at that point it does plagiarize.

Re: How can I prevent my site from being a free dataset for LLMs?

#59
post #58

Earlier quoted context omitted.

I completely agree with this; but my understanding about how LLM work is that they don't copy meaningful segments of text from any specific source. Instead, they predict the next block of text, which they'd only do if they've seen that idea/sequence enough times with context to rank the prediction high enough. I haven't seen any service copy out large block of text enough to make me think it's reasonable to call thei…

"use will only repeat an idea that many someone's have written about, such that it's seen the idea, or parts of that idea many times" As you go narrower with a query only one source of truth is available at that point it does plagiarize.

Very interesting, any chance you've got a citation for this? I'd like have some sort of proof next time I tell someone that they will happily plagiarize from single sources.

Re: How can I prevent my site from being a free dataset for LLMs?

#60
post #58

Earlier quoted context omitted.

"use will only repeat an idea that many someone's have written about, such that it's seen the idea, or parts of that idea many times" As you go narrower with a query only one source of truth is available at that point it does plagiarize.

Very interesting, any chance you've got a citation for this? I'd like have some sort of proof next time I tell someone that they will happily plagiarize from single sources.

Even if it's not copying word for word, if it's not citing its data, it's still plagiarism. Plagiarism includes copying ideas without crediting their source.
Post reply on HN