Live data from Hacker News

Llama and ChatGPT Are Not Open-Source

spectrum.ieee.org

111–120 of 130 posts

Re: Llama and ChatGPT Are Not Open-Source

#111

Earlier quoted context omitted.

That's not true though. Every word of this sentence, for instance, has a correct spelling and grammatical rules as to where the words and punctuation go. To English learners, that might be very difficult. I'm learning a new language now and I feel the pain. The textbook and instructor, in this case, is way more correct than the collective opinion of my fellow students. The vast majority of things actually follow this…

> Every word of this sentence, for instance, has a correct spelling and grammatical rules as to where the words and punctuation go. As a linguistic descriptivist, hahahahahahahaha.

An grammar so no because no rules yes Have. then its they're simple she Great understand your of problem in yes?

Or did I break some rules there?

We're not talking Oxford commas here. There are fundamental structures.

For instance, If I made that "are there fundamental structures" then it's a question.

Basic grammar is very real and not some subjective ephemeral thoughtstuff with wildly different opinions. It's so obvious that you don't even question it.

That's the way most things are - how to operate a faucet, close a window, put on a shirt, open a cabinet, use a fork... Most things in life aren't controversial.

Even the controversial ones like what's art and what's music, it's just a small periphery that's questioned. A symphony played by an orchestra is music, a painting in a museum is art - even these subjective categories substantially have extremely wide agreement.

True controversy is an outlier. (Ie, is 4'33" by John Cage music?)

You can claim whatever you want but with just things that's just exercising the freedom to be wrong.

Re: Llama and ChatGPT Are Not Open-Source

#112

I think we have forgotten the foundation of open source which was free software which was about empower the user. Instead of all these complicated criteria, there are only 2 criteria that get to the heart of the question: 1. Can you run this on your own hardware 2. Can you use this directly or tweak this as a community to make porn without a bunch of barriers This gets to the heart of the matter: Can you run the soft…

You're describing freeware. Open source means the source code is available.

Re: Llama and ChatGPT Are Not Open-Source

#113
post #74

Earlier quoted context omitted.

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

There is a lot of mental gymnastic, to submit an artwork publicly on the internet, allowing everyone to copy your arstyle, make derivative art of it, have other learn from it, but if ever a machine 'learn' from it, it's "stealing". It's not because you don't like a derivative work, that the derivative work is "stealing" your content. Saying it's stealing is wrong, it's lying to get your point accross. You blame "how…

> have other learn from it, but if ever a machine 'learn'

Why did you put one of those in quotes and not the other?

Re: Llama and ChatGPT Are Not Open-Source

#114

Earlier quoted context omitted.

So I can have a giant database of all the most recent copyrighted works, on my home computer, as long as I claim I'm using it for training a model? And if I happen to listen to some music or watch some movies too, who's going to know? I think an argument could be made that it's fine for an LLM to learn from copyrighted works, but maybe it should have to go to the library to do it. Having your own human-accessible cop…

Sure, I think you’re right, it’s reasonable to expect a company to at least buy a single copy of each book (or whatever it eats) it ingests if the work could not otherwise be legally obtained for free. Or something reasonable along those lines

Yes, let’s shut out the small players now go Microsoft can bilk us

Re: Llama and ChatGPT Are Not Open-Source

#115

Earlier quoted context omitted.

> Every word of this sentence, for instance, has a correct spelling and grammatical rules as to where the words and punctuation go. As a linguistic descriptivist, hahahahahahahaha.

An grammar so no because no rules yes Have. then its they're simple she Great understand your of problem in yes? Or did I break some rules there? We're not talking Oxford commas here. There are fundamental structures. For instance, If I made that "are there fundamental structures" then it's a question. Basic grammar is very real and not some subjective ephemeral thoughtstuff with wildly different opinions. It's so ob…

It was more the correct spelling claim I object to. Considering all variations of English.

Re: Llama and ChatGPT Are Not Open-Source

#116

"Mark Dingemanse, a coauthor of this report, had a particularly strong assessment of the Llama 2 model: "Meta using the term `open source' for this is positively misleading: There is no source to be seen, the training data is entirely undocumented, and beyond the glossy charts the technical documentation is really rather poor. We do not know why Meta is so intent on getting everyone into this model, but the history o…

It seems Dr Dingemanse understands the problems introduced by so-called "tech" companies, e.g., Google. "For basic visitor statistics I use Matomo, an excellent open source alternative for Google Analytics. IP addresses are anonymized and data never leaves the server." https://markdingemanse.net/credits

Edit: removed unreasonable tirade

Re: Llama and ChatGPT Are Not Open-Source

#117
post #76
post #59

Earlier quoted context omitted.

You're not addressing the massive abuse of the commons this still represents. If artists don't have the right to tell you to fuck off for using their work in training data, they're less likely to publicly show that work, which hurts them because they become less visible and hurts the AI because the training data gets worse.

Go on youtube and type "copy arstyle". Now tell me how artists were not stealing from each other.

> Now tell me how artists were not stealing from each other.

Key words being "from each other". AI only takes, it doesn't give anything back, it doesn't inspire. Their ultimate goal is to absorb the entire human history of art and then displace millions of people who received nothing from this transaction they were forced into, just so that some billionaire can afford another yacht.

I have no problems with AI models, but if you want to use art, writing, code, etc. for training, you should restrict your use to public domain works, ask for consent, or commission it. Obviously that would cost a lot of money so big tech companies are once again looking for a free ride.

Re: Llama and ChatGPT Are Not Open-Source

#118
post #76

Earlier quoted context omitted.

Go on youtube and type "copy arstyle". Now tell me how artists were not stealing from each other.

> Now tell me how artists were not stealing from each other. Key words being "from each other". AI only takes, it doesn't give anything back, it doesn't inspire. Their ultimate goal is to absorb the entire human history of art and then displace millions of people who received nothing from this transaction they were forced into, just so that some billionaire can afford another yacht. I have no problems with AI models,…

> AI only takes, it doesn't give anything back, it doesn't inspire.

Totally wrong. There is even a well known counter example: dream-like videos generated by ai wasn't something we had before. This statement tells me you never used it.

> just so that some billionaire can afford another yacht

You are heavily mistaken, that the scenario where ai learning isn't considered as fair use of material. In this case, only megacorps will be able to trains their own AI by spending billions in content.

> Obviously that would cost a lot of money so big tech companies are once again looking for a free ride.

You are constructing your own story while ignoring what is happening.

Adobe and Dall-e models were trained with datasets they mostly had rights over. OpenAI partnered with Shutterstock, and Adobe have their own photo stock. Big tech companies didn't had problem looking for image content. Stable Diffusion on the other hand is open source & open research, but didn't had any rights on most of the image they trained.

Re: Llama and ChatGPT Are Not Open-Source

#119
post #74

Earlier quoted context omitted.

There is a lot of mental gymnastic, to submit an artwork publicly on the internet, allowing everyone to copy your arstyle, make derivative art of it, have other learn from it, but if ever a machine 'learn' from it, it's "stealing". It's not because you don't like a derivative work, that the derivative work is "stealing" your content. Saying it's stealing is wrong, it's lying to get your point accross. You blame "how…

It's a lot of mental gymnastic to think machine learning = human learning. Especially on HN, where people should understand scale matters a lot in real world. It's generally consider ok to sell your fanart on comic festivals. But do you think it's okay that Disney starts selling fanart of One Piece without the publisher's permission? A car and a pair of legs both move you from A point to B point. So why do laws treat…

> It's a lot of mental gymnastic to think machine learning = human learning.

That's why I have put it in quotes.

> But do you think it's okay that Disney starts selling fanart of One Piece without the publisher's permission?

It's why it's called "fair use". Fanart isn't the only case of fair use. Parody for example allows it.

And AFAIK, no court yet have decided that AI "learning" is not fair use or what the limite are.

Re: Llama and ChatGPT Are Not Open-Source

#120
post #113
post #74

Earlier quoted context omitted.

There is a lot of mental gymnastic, to submit an artwork publicly on the internet, allowing everyone to copy your arstyle, make derivative art of it, have other learn from it, but if ever a machine 'learn' from it, it's "stealing". It's not because you don't like a derivative work, that the derivative work is "stealing" your content. Saying it's stealing is wrong, it's lying to get your point accross. You blame "how…

> have other learn from it, but if ever a machine 'learn' Why did you put one of those in quotes and not the other?

Because these two are similar, yet not the same.
Post reply on HN