Live data from Hacker News

Now is the time to give LLMs access to the ACM digital library

cacm.acm.org

171–180 of 185 posts

Re: Now is the time to give LLMs access to the ACM digital library

#171
post #136

Earlier quoted context omitted.

Why do you think big AI is going to just share all it’s ‘knowledge’?

Because you can access it for free? As it was the case from the first day LLMs became a thing in public consciousness? Or did I miss the change, and ChatGPT, Claude and Gemini no longer have a free tier anymore? And then the next tier that costs peanuts for anyone in the west, that gives you more access to slightly fresher models?

That is interacting with it - which people are getting charged more and more for.

The first ones are free.

And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.

Re: Now is the time to give LLMs access to the ACM digital library

#172

Earlier quoted context omitted.

> most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't I’m sorry, but I can’t buy this…

> Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it. Is that not why they’re shredding the books when they’re done with them?

First, this has nothing to do with the topic. ACM provides digital access. They don't have a single copy of a physical book which they're going to send to another company for destruction. Like I said, they're not deleting the source material or removing it from circulation.

As the other commenter posted, the original report that AI companies were shredding books was based on a second-hand retelling of a rumor, embellished for "AI bad" headlines.

There are actual bookshops talking about this are saying that most of the books are things like "How to master Microsoft Word 96" and that's why they're rare. They're also saying that the destination shipment is going to FBA (fulfilled by Amazon). They think it's an flipping operation trying to find arbitrage opportunities.

Also the reason AI companies have to destroy books is because they've been legally forbidden from using digital copies available. They had to pay a large settlement for it. So it's not some conspiracy to deprive the world of knowledge. It's what the courts told them they had to do.

Re: Now is the time to give LLMs access to the ACM digital library

#173
post #135

Earlier quoted context omitted.

How are you sure that what big ai is doing is “the increase in knowledge of human kind”? It is good for their business model; but it may lead to monopoly unlike anything acm ever had and very dubious prospect for academia.

I understand the LLM production companies have a funding mechanism but do they really have a business model?

Yes they do. They're drawing so much attention for the funding precisely because the business model is so attractive and has such a large TAM.

Re: Now is the time to give LLMs access to the ACM digital library

#174
post #171

Earlier quoted context omitted.

Because you can access it for free? As it was the case from the first day LLMs became a thing in public consciousness? Or did I miss the change, and ChatGPT, Claude and Gemini no longer have a free tier anymore? And then the next tier that costs peanuts for anyone in the west, that gives you more access to slightly fresher models?

That is interacting with it - which people are getting charged more and more for. The first ones are free. And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.

> which people are getting charged more and more for.

Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock.

There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.

Re: Now is the time to give LLMs access to the ACM digital library

#175
post #107

Earlier quoted context omitted.

I've seen: 1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through 2. AI models plagiarising stack overflow answers word for word, quite recently about the r…

If it never regurgitated the exact same thing, would you accept AI then?

For me there's two separate problems:

1. The plagiarism aspect, and that most of the training data was used without permission

2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)

Re: Now is the time to give LLMs access to the ACM digital library

#176
post #171

Earlier quoted context omitted.

That is interacting with it - which people are getting charged more and more for. The first ones are free. And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.

> which people are getting charged more and more for. Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock. There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.

Says no one managing a corp budget!

Re: Now is the time to give LLMs access to the ACM digital library

#177
post #176

Earlier quoted context omitted.

> which people are getting charged more and more for. Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock. There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.

Says no one managing a corp budget!

I’m talking about price for a given output.

You’re talking about companies using more tokens.

Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.

Re: Now is the time to give LLMs access to the ACM digital library

#178
post #176

Earlier quoted context omitted.

Says no one managing a corp budget!

I’m talking about price for a given output. You’re talking about companies using more tokens. Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.

I’m talking about they can charge you whatever, and you don’t own anything so you don’t get a say.

Re: Now is the time to give LLMs access to the ACM digital library

#179
post #178

Earlier quoted context omitted.

I’m talking about price for a given output. You’re talking about companies using more tokens. Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.

I’m talking about they can charge you whatever, and you don’t own anything so you don’t get a say.

Whatever was released in the open as weights is forever out there and beyond reach of any corporate interest.

Re: Now is the time to give LLMs access to the ACM digital library

#180
post #175

Earlier quoted context omitted.

If it never regurgitated the exact same thing, would you accept AI then?

For me there's two separate problems: 1. The plagiarism aspect, and that most of the training data was used without permission 2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)

1. I said in my hypothetical it would not reproduce exact content. 2. Okay, others find it useful. So what?
Post reply on HN