Now is the time to give LLMs access to the ACM digital library
101–110 of 185 posts
Re: Now is the time to give LLMs access to the ACM digital library
#102Earlier quoted context omitted.
The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph. Derivative work or transformative? It's not the same.
AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for…
Re: Now is the time to give LLMs access to the ACM digital library
#103Re: Now is the time to give LLMs access to the ACM digital library
#104Earlier quoted context omitted.
How do you square away the idea that you do science for the increase in knowledge of human kind, but then say that a particular use of that knowledge is verboten? I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science. Obviously there are the moral / ethical arguments about AI in general here to weigh against - tho…
Good question. Unfortunately, academic knowledge is widely ‘verboten’ already. Everything under paywall, researchers having to pay up to $10,000 to publish in open access in some venues, rare books unavailable even to top universities. Access to knowledge and information is increasingly difficult for everyone. That said, what matters here is the social contract, what do I bring to society and what do we get from tech…
I’m sorry, but I can’t buy this argument. Making information more available does not make it more discriminatory. Nobody is saying it will only be available in the best models and withheld from other models or services like the ChatGPT free plan. Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it. Nothing about this shrinks access or makes it more discriminatory.
I understand that you’re upset about the use of the content, but I think you need to admit that your stance is the one trying to restrain use of the content. Training LLMs on it can only bring knowledge to a wider audience, not restrict it.
Whether or not that’s a good or fair idea is a separate discussion, but arguing that this makes access to the knowledge more discriminatory and locked away is 180 degrees backwards.
Re: Now is the time to give LLMs access to the ACM digital library
#105Earlier quoted context omitted.
The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.
Everybody gives more power to the larger players. You won't work for me for $10/hr but you'll work for a guy who has more money for more. There's nothing wrong with that. Money is just a fungible unit representing value offered.
Re: Now is the time to give LLMs access to the ACM digital library
#106Earlier quoted context omitted.
Bartz v. Anthropic PBC, No. 24-cv-05417 (N.D. Cal. June 23, 2025) Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)
Did... you read any of these? Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for There's also these parts: > its use of pirated books to create such library does not constitute fair use. Which indicates that there ar…
I just pointed out none took place.
Anyway I'm not replying in this thread anymore.
Re: Now is the time to give LLMs access to the ACM digital library
#107Earlier quoted context omitted.
AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for…
Are you sure you didn’t have RAG enabled and it wasn’t putting the paper in its context for your prompt? I find it hard to believe that anything but the most commonly published papers/code would exist directly in model weights.
1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through
2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail
3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer
4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used
There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information
Re: Now is the time to give LLMs access to the ACM digital library
#108Earlier quoted context omitted.
If it was a non profit that trained the model - would that change your mind?
The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.
And when the open weights models distill all the content out of the majors anyway?