Live data from Hacker News

Many AI safety orgs have tried to criminalize currently-existing open-source AI

1a3orn.com

1–10 of 405 posts

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#2
There are many types of safety. For example, protecting your profits! I'd imagine that if we trace the money back these organisations will look a lot like lobbyists for existing companies in the AI space. I recall Microsoft's licence enforcement effort was done with that sort of scheme, I think they used the BSA [0]. It has been a while though so maybe it was a different group.

Anyway, point being, if they can lobby for something unpopular under a different brand, that is how to do it. Much less PR risk.

[0] https://en.wikipedia.org/wiki/Software_Alliance

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#5
There is a bigger reason why the end of open source AI might be close: as soon as training data becomes licensed, that’s it for open source AI. Poof.

I wish I could be more eloquent on this point, but I’ve mostly just been depressed about this seeming inevitability.

Hopefully it won’t be the case. But how could it be otherwise? Hundreds of thousands of people are mad at openai and midjourney for doing exactly what open source AI needs to do in order to survive: fine tune or train from scratch.

As soon as some politician makes it a platform issue, it seems like the law will simply be rewritten to prevent companies from using training data at will. It’s such a compelling story: "big companies are stealing data owned by small businesses and individuals." So even if the court cases are decided in OpenAI’s favor, it’s not at all clear that the issue will be settled.

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#6
1. Licensing of data will be huge bottleneck 2. Uncensored results will be used against opensource models questions hovering in dark or grey area 3. Limited compute compared to big corp and model size gap 7B for opensource and closed source would be magnitude bigger

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#7
AI safety people are hypocrites. If they practiced what they preached, they'd be calling for all AI to be banned, ala Dune. There are AI harms that don't care about whether or not the weights are available, and are playing out today.

I'm talking about the ability of any AI system to obfuscate plagiarism[0] and spam the Internet with technically distinct rewords of the same text. This is currently the most lucrative use of AI, and none of the AI safety people are talking about stopping it.

[0] No, I don't mean the training sets - though AI systems seem to be suspiciously really good at remembering them, too.

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#8

There is a bigger reason why the end of open source AI might be close: as soon as training data becomes licensed, that’s it for open source AI. Poof. I wish I could be more eloquent on this point, but I’ve mostly just been depressed about this seeming inevitability. Hopefully it won’t be the case. But how could it be otherwise? Hundreds of thousands of people are mad at openai and midjourney for doing exactly what op…

[deleted]

Re: Many AI safety orgs have tried to criminalize currently-existing open-source AI

#10

There is a bigger reason why the end of open source AI might be close: as soon as training data becomes licensed, that’s it for open source AI. Poof. I wish I could be more eloquent on this point, but I’ve mostly just been depressed about this seeming inevitability. Hopefully it won’t be the case. But how could it be otherwise? Hundreds of thousands of people are mad at openai and midjourney for doing exactly what op…

The advantage that open AI (not the company) has is that if using copyrighted content as training data without licensing it is found to be illegal, they can just keep doing it. There's plenty of FOSS software basically designed to violate copyright law (comic readers, home media center servers/clients, torrent clients) that big tech cannot compete with lest they face legal consequences. Basically what I'm saying is that the open source community will continue to use books3 and scraped images to train while facebook and the like get stuck in legal quagmire.

Of course, this ignores the fact that popular "open" models of today were actually trained with facebook or other companies' computational resources, so unless some cheaper way to train models were developed we would actually be stuck with proprietary models trained with lots of compute but unable to use unlicensed training data, and open models that can use whatever data they like but must operate in the shadows without access to much compute for training.

Post reply on HN