Live data from Hacker News

EU's AI Act: ChatGPT must disclose use of copyrighted training data

artisana.ai

61–70 of 72 posts

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#61

Today the planet bears the load of over eight billion autonomous agents grabbing training data from all the other agents. This intellectual thievery must stop.

So much goes into open source and proper licensing and attribution, think about how much you directly or indirectly benefit from that ?

Just saying that we should go for a free for all and trash IP ownership won’t be good because those with money today will crush those without and take everything that was publicly available and owned without giving back.

This is IMO what Open AI have done.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#62
post #7

It's so absolutely obvious that the concept of intellectual property is not going to survive. What's the point of this agonizing life support?

It's not just about IP. Have you considered how much an LLM trained on scraping websites might know about you? Would it reveal that information if I prompted it with something like "Create a short story about HN user jMyles and his home life"? It might be best to know what these models have actually been trained on.

100s of millions of people used it for quite some time now, many on daily basis. Do we have any evidence that it “knows about you”? Why can’t we say that the “internet” knows too much about you?

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#63
post #62

Earlier quoted context omitted.

It's not just about IP. Have you considered how much an LLM trained on scraping websites might know about you? Would it reveal that information if I prompted it with something like "Create a short story about HN user jMyles and his home life"? It might be best to know what these models have actually been trained on.

100s of millions of people used it for quite some time now, many on daily basis. Do we have any evidence that it “knows about you”? Why can’t we say that the “internet” knows too much about you?

To your latter question, you very much could. But pretty much the whole point of these models is to create connections between various bits of information. It obviously knows about some people (otherwise how could you ask it about famous people?). If it's been trained on everything you can trawl from the net, then why wouldn't it know about you?

But I don't know if people have been asking these bots about themselves. I don't have access to ChatGPT4. Anyone checked this?

I suspect the ChatGPT models haven't been trained on everything available online, so it wouldn't know much, but perhaps the next generation will.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#65
post #62

Earlier quoted context omitted.

100s of millions of people used it for quite some time now, many on daily basis. Do we have any evidence that it “knows about you”? Why can’t we say that the “internet” knows too much about you?

To your latter question, you very much could. But pretty much the whole point of these models is to create connections between various bits of information. It obviously knows about some people (otherwise how could you ask it about famous people?). If it's been trained on everything you can trawl from the net, then why wouldn't it know about you? But I don't know if people have been asking these bots about themselves.…

Internet is basically a bunch of nodes with interconnected data. Any search index might be considered as an interconnected graph. Same way with LLMs, it’s nothing new, might be considered faster depending on the definition.

Comparing to celebrities, data about people are so sparse that it would look like noise. I would be surprised if it encoded anything useful.

Half of the internet and all media were obsessed with making it say the f-word and tricking it into saying that it would kill all the people. Attacking from the privacy is quite obvious but I didn’t see it mentioned. I asked about 10 random friends and myself and it didn’t recognize the names despite having plenty of search results.

In one of the interviews they mentioned that it was trained on 10% of the web. So should have enough data.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#66
post #59
post #43

Earlier quoted context omitted.

> Many parts of Github would not exist without intellectual property laws. I do not think this is true. I think most devs on GitHub operate as if there are no IP laws. I think if they went away, almost nothing would change (some noise around "license" fields would go away).

Licensing is very important to open source. It is literally what drives and protects large scale open source innovation and stops it going extinct. It’s actually worth learning about GPL 3.0 and copy left licensing. Anyway if that goes away. A lot of innovation will too. If privately owned LLMs go on consuming everything, not giving back to the communities that make them what they are. It may erode the system.

If there is no IP law there is no longer any need for Copyleft. Copyleft is a means to an end.

https://c4sif.org/2022/05/against-intellectual-property-afte...

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#67
post #60

Earlier quoted context omitted.

Why is an app needed? Maybe it's different in the US but we've had an app to get a taxi for quite a while. I've never seen any need to install that app. It's easier to call, tell how many people need a ride, tell when it's needed and get a confirmation that it's on its way and long it'll take for it to be there. Why do I need an app?

You’re on hacker news and you ask “why do I need an app?”, red alert. I agree with you by the way.

If I may paraphrase a book title "The best app is no app"? :-) Besides, I thought we're supposed to be building services you can talk with now so you don't have to fiddle with old fashioned buttons anymore? Sure the human component doesn't scale that well, but it's pretty localized service and speech recognition is amazing.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#68
post #66
post #59

Earlier quoted context omitted.

Licensing is very important to open source. It is literally what drives and protects large scale open source innovation and stops it going extinct. It’s actually worth learning about GPL 3.0 and copy left licensing. Anyway if that goes away. A lot of innovation will too. If privately owned LLMs go on consuming everything, not giving back to the communities that make them what they are. It may erode the system.

If there is no IP law there is no longer any need for Copyleft. Copyleft is a means to an end. https://c4sif.org/2022/05/against-intellectual-property-afte...

Lol but there will never be an end to IP law while there are people with money and influence. Utopia is an idea, not a reality.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#69
post #12

We, in Europe, are jumping the gun way too soon and this will have serious consequences to the industry, here. At this moment we barely understand how or why LLMs do what will be their role in society. Why is it that some bureaucrats want to regulate and based on what given the status of the industry? The only thing I see is the industry moving elsewhere just as it is starting to develop which is a shame.

I wish instead that more technologies were more regulated from the start. Take social media as an example, we are just realizing how bad it is for a lot of peaople and it is now too late to go back.

Re: EU's AI Act: ChatGPT must disclose use of copyrighted training data

#70
post #69
post #12

We, in Europe, are jumping the gun way too soon and this will have serious consequences to the industry, here. At this moment we barely understand how or why LLMs do what will be their role in society. Why is it that some bureaucrats want to regulate and based on what given the status of the industry? The only thing I see is the industry moving elsewhere just as it is starting to develop which is a shame.

I wish instead that more technologies were more regulated from the start. Take social media as an example, we are just realizing how bad it is for a lot of peaople and it is now too late to go back.

That works both ways. Imagine that they had regulated the internet as it was starting to appear in the 90s? It actually happened in a small way here in Portugal. They regulated the .pt domain in such a restrictive way that it became irrelevant and everyone went with .com or other options. Thankfully they didn't regulate the internet services themselves or Europe would be a digital backwater these days.

This time they want to regulate the services even before they are functional which is crazy. They even call it the Artificial Intelligence Act when it is not clear if there is intelligence involved. It is also strange the insistence that companies have to disclose if the models were trained with copyrighted material. Google and Wikipedia, to name a couple, use plenty of copyrighted material and that seems to be ok and any issues in that department are already regulated.

Post reply on HN