Live data from Hacker News

Llama and ChatGPT Are Not Open-Source

spectrum.ieee.org

11–20 of 130 posts

Re: Llama and ChatGPT Are Not Open-Source

#11

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> reducing bias in the models is a noble goal

It's also a necessary goal in order for these models to be more broadly adopted.

We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal.

And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening right now between authors and OpenAI. Not sure that is a position most startups want to be in.

Re: Llama and ChatGPT Are Not Open-Source

#13

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> reducing bias in the models is a noble goal It's also a necessary goal in order for these models to be more broadly adopted. We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal. And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening righ…

College students don't want to be criminally charged with theft for pirating an $800 college textbook either. I'm not giving business advice, I'm giving humanity and knowledge propagation advice.

As I mentioned, reducing bias is good so long as it doesn't introduce more bias elsewhere. The ultimate goal of course being a 100% bias free model.

Re: Llama and ChatGPT Are Not Open-Source

#14

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> reducing bias in the models is a noble goal It's also a necessary goal in order for these models to be more broadly adopted. We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal. And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening righ…

Nah, these are foundation models. Companies want to be able to put in guard rails that are applicable to their application, not start with a model lobotomized according to US tech company values. The censorship is about telling people how to think, like it always is, not for the good of the people using the models.

Re: Llama and ChatGPT Are Not Open-Source

#16

Earlier quoted context omitted.

> reducing bias in the models is a noble goal It's also a necessary goal in order for these models to be more broadly adopted. We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal. And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening righ…

Nah, these are foundation models. Companies want to be able to put in guard rails that are applicable to their application, not start with a model lobotomized according to US tech company values. The censorship is about telling people how to think, like it always is, not for the good of the people using the models.

Yup. Advance the foundation models as far as possible to create a representation of the internet/human experience and place safeguards on top.

Don't want your LLM to be used to create erotic fanfiction? Instead of gutting the LLM, just put a detector on the query and answers feeding into the LLM. Thankfully, with a non-neutered model you have access to a tool than can be used to perform such detection...

Re: Llama and ChatGPT Are Not Open-Source

#17

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

We really need to somehow separate a bias towards accuracy as distinct from some bias towards say, a sports team.

Using everything would be like taking a bunch of students final exams and then claiming the most common answers are the correct ones.

This isn't how expertise and accuracy works. Most things worth doing are not only genuinely hard and complicated but something that only a minority subset of accomplished people can do consistently well in.

Listening to everybody and incorporating their thoughts is only going to lead to wrong answers.

Being selective is the key here unless you genuinely want say, answers about to space to involve aliens and UFOs - because way more people believe in that then there are qualified PhD astrophysicists in the world.

Similarly, way more people believe in vaccine conspiracy theories then there are people with significant viral epidemiology backgrounds.

This pattern is true in every field.

Re: Llama and ChatGPT Are Not Open-Source

#18

Earlier quoted context omitted.

> reducing bias in the models is a noble goal It's also a necessary goal in order for these models to be more broadly adopted. We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal. And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening righ…

Nah, these are foundation models. Companies want to be able to put in guard rails that are applicable to their application, not start with a model lobotomized according to US tech company values. The censorship is about telling people how to think, like it always is, not for the good of the people using the models.

> US tech company values

Which do broadly align with the values across most of the world.

But if you want to build something that has different values then go ahead.

But you can't expect companies to be complicit in doing something which is unethical or illegal.

Re: Llama and ChatGPT Are Not Open-Source

#19

Earlier quoted context omitted.

> reducing bias in the models is a noble goal It's also a necessary goal in order for these models to be more broadly adopted. We've seen too many examples of bias in the training data set manifesting in ways that actively discriminate against people. Which is unethical and in many places illegal. And having copyrighted material and sexual content in your model will simply open you up to lawsuits as is happening righ…

Nah, these are foundation models. Companies want to be able to put in guard rails that are applicable to their application, not start with a model lobotomized according to US tech company values. The censorship is about telling people how to think, like it always is, not for the good of the people using the models.

[dead]

Re: Llama and ChatGPT Are Not Open-Source

#20

Earlier quoted context omitted.

Nah, these are foundation models. Companies want to be able to put in guard rails that are applicable to their application, not start with a model lobotomized according to US tech company values. The censorship is about telling people how to think, like it always is, not for the good of the people using the models.

> US tech company values Which do broadly align with the values across most of the world. But if you want to build something that has different values then go ahead. But you can't expect companies to be complicit in doing something which is unethical or illegal.

[deleted]
Post reply on HN