Live data from Hacker News

Llama and ChatGPT Are Not Open-Source

spectrum.ieee.org

41–50 of 130 posts

Re: Llama and ChatGPT Are Not Open-Source

#41

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

We really need to somehow separate a bias towards accuracy as distinct from some bias towards say, a sports team. Using everything would be like taking a bunch of students final exams and then claiming the most common answers are the correct ones. This isn't how expertise and accuracy works. Most things worth doing are not only genuinely hard and complicated but something that only a minority subset of accomplished p…

When you dig into it, there are very few truths. Who "won" the last US election? Where did covid come from?

Humans will argue the right answer until their last days. It's frustrating how on-the-fence chatgpt can be. It's pretty interesting too, because in a professional environment one of the most important things you need to do is have an opinion and take a position otherwise you can't execute.

Re: Llama and ChatGPT Are Not Open-Source

#42

Earlier quoted context omitted.

We really need to somehow separate a bias towards accuracy as distinct from some bias towards say, a sports team. Using everything would be like taking a bunch of students final exams and then claiming the most common answers are the correct ones. This isn't how expertise and accuracy works. Most things worth doing are not only genuinely hard and complicated but something that only a minority subset of accomplished p…

Why is it a binary choice? "Most students would answer X, due to this common misconception about Y". As we have seen from history, there is not often an absolute truth to questions, only clusters of truths. We want our LLM to be able to perform reasoning, mathematics, and science, but expecting absolute truths in anything outside of those fields is a bit much. Wikipedia often takes a good approach here and strikes th…

That's not how bullshit works.

An expertise is needed from the reader when presented with false material. That's why so many people think parody news articles are real.

If you have ever gone through any comment thread on the internet about climate change you'll find a large volume of climate deniers citing God, greedy liberals, George Soros, whatever.

The people working with these fictions can always just add more fiction to fill in whatever hole. They aren't bound by reality or credulity.

Actually understanding global climate systems takes an incredible amount of study. You just can't blurt out a few paragraphs of words and get people to stop believing in Jewish weather machine conspiracies or whatever they're thinking.

We would have been done with that by now if this were the case.

Dislodging bullshit is hard. Humanity only started figuring out how a few hundred years ago and we're still fucking it up - such as in the replication crisis

Re: Llama and ChatGPT Are Not Open-Source

#43

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is a no-no.

What other option would there be though? It seems like a binary field. You can either leave it wholly unaligned, or you can attempt to align the model, and you will introduce your own bias as a side effect.

In the future, and in fact in the present, there is already a "gray market" for unaligned models. There will surely be market for these and they will sell just any other item on this market.

Re: Llama and ChatGPT Are Not Open-Source

#44

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> Yes, reducing bias in the models is a noble goal

“Bias” is not a unidimensional thing (or even one with a meaningful magnitude); you can’t reduce bias, you can make it more transparent (better documented) so that you can evaluate appropriateness for particular uses.

While I am not a big fan of the crew that pushes “alignment” as the main issue with regard to AI for other reasons, that’s a particularly good term for what is important – aligning the biases of the AI with the usage intent. “Reducing” bias makes sense only with regard to a specific usage intent.

Re: Llama and ChatGPT Are Not Open-Source

#45
post #22

Can we please stop using terminology related to code for something that's not code?

Agreed it's annoying, but on a certain level also not entirely wrong. In a software 2.0 [0] world the weights are functionally the code in that it is what gets you from input to output. Open weights or something similar would be better though [0] https://karpathy.medium.com/software-2-0-a64152b37c35

Weights are configuration.

Re: Llama and ChatGPT Are Not Open-Source

#46

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

Modern copyright is not a good, it's an evil.

Re: Llama and ChatGPT Are Not Open-Source

#47

I personally do not want the companies to release training data (at least for a while) because then it gives people leverage to neuter it. I don't want a sanitized LLM, and I don't have $60M lying around to train my own. Copyrighted material, sexual content, political opinions, throw it all in and release it please! Yes, reducing bias in the models is a noble goal, but introducing new bias and blindspots to do it is…

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

Pretty much everything nowadays is copyrighted, by omitting such materials, what are you really left with?

LLM is a tool much like the internet is a tool. Yes, someone can use it to steal, but stealing is against the law.

Instead of encoding a criminal justice system into an LLM by omitting the possibility of stealing an artists work or omitting the knowledge of physics so someone can't learn how to build a bomb, we should instead just prosecute people for using it in that way intentionally.

How often do people get prosecuted for ripping off an artists style? The criminal justice system "hated artists" long before LLMs and it's not the responsibility of the tech. companies to rectify that in my opinion.

Re: Llama and ChatGPT Are Not Open-Source

#48

Earlier quoted context omitted.

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

I can't imagine what goes through the head of someone who thinks they don't have the right to say no. They don't have the right to be listened to or the right to use state power to enforce a colloquial notion of IP and rights with no basis in law or fact, but they surely have the right to say "No!" In all seriousness, they're afraid because they aren't very pleasant. AI will write things people will like and enjoy he…

Moralism! In art! Heaven forfend. How dare that art have a message, right? Never in history, not once, before the Woke Tens and, I guess, the Woker Twenties did somebody have an opinion and encode it in allegory. It is a purely novel phenomenon and all known media before this dark age was absolutely without message and without any editorial lens to determine what is and is not to be depicted--for that, of course, would communicate values and ideas, and art simply did not do such a thing.

More seriously: if you have this attitude for-reals and aren't Poe's Lawing me, perhaps you need a little more preaching aimed your way because your conception of art and the civic society is the closest thing I can think of to a humanities position being objectively wrong.

The ideological soma of agree-with-you AI being desirable is scary enough stuff. Removing communication from art being a "great technological accomplishment" is fucking dystopic.

Re: Llama and ChatGPT Are Not Open-Source

#49
post #22

Earlier quoted context omitted.

Agreed it's annoying, but on a certain level also not entirely wrong. In a software 2.0 [0] world the weights are functionally the code in that it is what gets you from input to output. Open weights or something similar would be better though [0] https://karpathy.medium.com/software-2-0-a64152b37c35

Weights are configuration.

Weights almost entirely encapsulate the behavior of the model. Whatever structural analogy you want to use, they are closest in spirit to the core functional algorithm of traditional software.

Re: Llama and ChatGPT Are Not Open-Source

#50

Earlier quoted context omitted.

> Copyrighted material, sexual content, political opinions, throw it all in and release it please! Why copyrighted material? Could we stop celebrating how tech is going to steal everyone's copyrighted works in a massive effort to replace the artists who made it? Why does everyone here hate artists so much? Do they not deserve any rights over their IP, eg, the right to say no when someone wants to make derivative work…

Modern copyright is not a good, it's an evil.

That depends on the jurisdiction, but in my view I err on the side of "but it's necessary."

A society without copyright would be a much poorer one, and humans figured that out a long time ago.

Post reply on HN