Live data from Hacker News

Creativity has left the chat: The price of debiasing language models

arxiv.org

231–238 of 238 posts

Re: Creativity has left the chat: The price of debiasing language models

#231

Distilling my thoughts on 'debiasing' here, and in a variety of other modern endeavors. It is better to have representations of reality that you can then discuss and grapple with honestly, than to try to distort representations - such as AI - to make them fit some desired reality and then pressure others to conform their perception to your projected fantasy. Representations don't create reality, and trying to use rep…

"Reality" is a tricky concept. For me, I follow Jeff Atwood - if it isn't written down, it doesn't exist. According to this logic, people wasted a lot of time on imaginary, illusory things for most of human history, but now they have phones and most communication is digital so there is the possibility to finally be productive. This definition shows how the concept of distorting reality or honestly representing realit…

Yes. Black holes existed before we had the words to describe them.

Re: Creativity has left the chat: The price of debiasing language models

#232

Earlier quoted context omitted.

> I wouldn't be surprised if 80% or more of the works ended up being some depiction of aspects of Christianity painted by a white male. Are you saying diversity in art is signified by the artist's race and sex?

Bias is inherit in the choice of training material. I'm saying that diversity of expression is product of diversity of experience.

It must be, ultimately, because all art is individual or group expression, and each person can only belong to so many groups. But the individual expression still allows for a giant amount of expressiveness, and the group expression is wider than race or sex.

Re: Creativity has left the chat: The price of debiasing language models

#233
post #109

Shouldn't "debiasing" be in scare quotes? What they are clearly doing is biasing .

Given a biased corpus, de-biasing is the process of ensuring a less biased outcome. We can measure bias fairly well, so it seems absurd to conflate the two by suggesting that unbiased behaviour is simply another form of biased behaviour. For all practical purposes, there is a difference.

> We can measure bias fairly well...

Really? What's the absolute standard that you are measuring that bias against?

Citation needed.

Re: Creativity has left the chat: The price of debiasing language models

#234

I had an argument with some people over what debiasing means. There is some interesting research on fair clustering that I think points the way. The way fair clustering works is that you take data with both protected and unprotected attributes, and then you orthogonalize the unprotected attributes based on the protected attributes. So for example, if race is protected and income is unprotected, but there is a strong…

I think that works mathematically, but kicks the can down the road to how your original data was assembled, which was definitely with the knowledge of and usually in the belief in the usefulness of the characteristics that you're trying to extract. The idea that the good data is secretly encoded in uncorrupted form within the bad data I think is a bad idea. It reminds me of trying to make bad mortgages into good CDOs…

[deleted]

Re: Creativity has left the chat: The price of debiasing language models

#235

Earlier quoted context omitted.

But what if you remove the word "genetically"? I think there are a lot of people who would say "Indian people are better at math" and not even think about why they think that or why it might even be true. In my opinion, most biases have some basis in reality. Otherwise where else did they come from?

Well, the stereotype of Indian people being good at math specifically was itself a consequence of survivorship bias, that emerged from observing Indian visa holders who were hired based on their skills and credentials, and who were not at all representative of the average person in India. There is a BIG difference between biases being based in reality (which they're not), and biases being based in our varying percept…

OK, but like for me when someone says "Indian people are better at math", my mind basically says to myself, "OK, the average Indian person in the USA (the ones that they and I see and interact with) is better at math than the average person".

Because in my mind, that's the environment the person making that claim was in, so I just kind of automatically include that in my interpretation of their statement.

I don't think they are making a generalized statement that that Indian people are genetically better at math. I think they are making a statement that they perceive that average Indian person that they run into is better at math than the average person they run into. And maybe they are right, and maybe there is a reason based in reality why that is.

It sounds like that is somewhat true based on what you said about the visas.

I never take any of these things to have anything to do with genetics. To me it's always due to some external factor like the visas as you mentioned, or even maybe just like a cultural thing where they are pushed harder to be good at something as they go through school, and so are better at something than the average person in the end.

Re: Creativity has left the chat: The price of debiasing language models

#236

Earlier quoted context omitted.

Well, the stereotype of Indian people being good at math specifically was itself a consequence of survivorship bias, that emerged from observing Indian visa holders who were hired based on their skills and credentials, and who were not at all representative of the average person in India. There is a BIG difference between biases being based in reality (which they're not), and biases being based in our varying percept…

OK, but like for me when someone says "Indian people are better at math", my mind basically says to myself, "OK, the average Indian person in the USA (the ones that they and I see and interact with) is better at math than the average person". Because in my mind, that's the environment the person making that claim was in, so I just kind of automatically include that in my interpretation of their statement. I don't thi…

It's great that you personally don't generalize it into a stereotype about Indian people, but that is not what a stereotype is, and "Indian people are good at math" is the stereotype, not "high skilled Indian visa holders working in the US are good at math".

Re: Creativity has left the chat: The price of debiasing language models

#237

Earlier quoted context omitted.

Even assuming we can make an unbiased model (assuming by unbiased we mean something like "has a world model and reasoning that has no systematic deviation from reality"), we couldn't recognize the model as unbiased. I'd even wager that outside of research such a model would be completely unusable for practical applications. Both as individual humans and as collective societies we have a lot of biases. And judging by…

> assuming by unbiased we mean something like "has a world model and reasoning that has no systematic deviation from reality" Yeah that’s a way’s off. An LLM is just a reflection of the text that humans write, and humans seem very far off from having world models and reasoning that accurately reflect reality. We can’t even reason about what the real differences are between men and women (plus countless other issues)…

> An LLM is just a reflection of the text that humans write, and humans seem very far off from having world models and reasoning that accurately reflect reality

The original sin of LLMs is that they are trained to imitate human language output.

Passing the Turing test isn't necessarily a good thing; it means that we have trained machines to imitate humans (including biases, errors, and other undesirable qualities) to the extent that they can deceptively pose as humans.

Re: Creativity has left the chat: The price of debiasing language models

#238
post #119

CoPilot is now basically useless for discussing or even getting recent information about politics and geopolitical events. Not only opinions are censored, but it refuses to get the latest polls about the U.S. presidential elections ! You can still discuss the weather, get wrong answers to mathematics questions or get it to output bad code in 100 programming languages. I would not let a child near it, because I would…

You think its opinions would be counter to a child's benefit? Any examples?
Post reply on HN