Live data from Hacker News

Google to pause Gemini image generation of people after issues

theverge.com

471–480 of 1001 posts

Re: Google to pause Gemini image generation of people after issues

#471
post #398

Earlier quoted context omitted.

There is an _actual problem_ that needs to be solved. If you ask generative AI for a picture of a "nurse", it will produce a picture of a white woman 100% of the time, without some additional prompting or fine tuning that encourages it to do something else. If you ask a generative AI for a picture of a "software engineer", it will produce a picture of a white guy 100% of the time, without some additional prompting or…

> If you ask generative AI for a picture of a "nurse", it will produce a picture of a white woman 100% of the time, without some additional prompting or fine tuning that encourages it to do something else. > If you ask a generative AI for a picture of a "software engineer", it will produce a picture of a white guy 100% of the time, without some additional prompting or fine tuning that encourages it to do something el…

I agree there aren't any perfect solutions, but a reasonable solution is to go 1) if the user specifies, generally accept that (none of these providers will be willing to do so without some safeguards, but for the most part there are few compelling reasons not to), 2) if the user doesn't specify, priority one ought to be that it is consistent with history and setting, and only then do you aim for plausible diversity.

Ask for a nurse? There's no reason every nurse generated should be white, or a woman. In fact, unless you take the requestors location into account there's every reason why the nurse should be white far less than a majority of the time. If you ask for a "nurse in [specific location]", sure, adjust accordingly.

I want more diversity, and I want them to take it into account and correct for biases, but not when 1) users are asking for something specific, or 2) where it distorts history, because neither of those two helps either the case for diversity, or opposition to systemic racism.

Maybe they should also include explanations of assumptions in the output. "Since you did not state X, an assumption of Y because of [insert stat] has been implied" would be useful for a lot more than character ethnicity.

Re: Google to pause Gemini image generation of people after issues

#472

Earlier quoted context omitted.

There is an _actual problem_ that needs to be solved. If you ask generative AI for a picture of a "nurse", it will produce a picture of a white woman 100% of the time, without some additional prompting or fine tuning that encourages it to do something else. If you ask a generative AI for a picture of a "software engineer", it will produce a picture of a white guy 100% of the time, without some additional prompting or…

Why does it matter which race it produces? A lot of people have been talking about the idea that there is no such things as different races anyway, so shouldn't it make no difference?

>Why does it matter which race it produces?

When you ask for an image of Roman Emperors, and what you get in return is a woman or someone not even Roman, what use is that?

Re: Google to pause Gemini image generation of people after issues

#473
post #284
post #260

Earlier quoted context omitted.

You're speaking as if LLMs are some naturally occurring phenomena that people are Google have tampered with. There's obviously always human intervention as AI systems are built by humans.

It's pretty clear to me what the commenter means even if they don't use the words you like/expect. The model is built by machine from a massive set of data. Humans at Google may not like the output of a particular model due to their particular sensibilities, so they try to "tune it" and "filter both input/output" to limit of what others can do with the model to Google's sensibilities. Google stated as much in their a…

Censorship of what? You object to Google applying its own bias (toward avoiding offensive outcomes) but you're fine with the biases inherent to the dataset.

There is nothing the slightest bit objective about anything that goes into an LLM.

Any product from any corporation is going to be built with its own interests in mind. That you see this through a political lens ("censorship") only reveals your own bias.

Re: Google to pause Gemini image generation of people after issues

#474
post #381

Earlier quoted context omitted.

No one is upset that an algorithm accidentally generated some images, they are upset that Google intentionally designed it to misrepresent reality in the name of Social Justice.

It's more accurate to say that it's designed to construct an ideal reality rather than represent the actually existing one. This is the root of many of the cultural issues that the West is currently facing. “The philosophers have only interpreted the world, in various ways. The point, however, is to change it. - Marx

> construct an ideal reality rather than represent the actually existing one

If I ask to generate an image of a couple, would you argue that the system's choice should represent "some ideal" which would logically mean other instances are not ideal?

If the image is of a white woman and a black man, if I am a lesbian Asian couple, how should I interpret that? If I ask for it to generate an image of image of two white gays kissing and it refuses because it might cause harm or some such nonsense, is it not invalidating who I am as a young white gay teenager? If I'm a black African (vs. say a Chinese African or a white African), I would expect a different depiction of a family than the one American racist ideology would depict because my reality is not that and your idea of what ideal is is arrogant and paternalistic (colonial, racist, if you will).

Maybe the deeper underlying bug in human makeup is that we categorize things very rigidly, probably due to some evolutionary advantage, but it can cause injustice when we work towards a society where we want your character to be judged, not your identity.

Re: Google to pause Gemini image generation of people after issues

#475

OpenAI already experienced this backlash when it was injecting words for diversity into prompts (hilariously if you asked for your prompt back it would include the words, and supposedly you could get it to render the extra words onto signs within the image). How could Google have made the same mistake but worse ?

I think it's pretty clear that they're trying to prevent one class of issues (the model spitting out racist stuff in one context) and have introduced another (the model spitting out wildly inaccurate portrayals of people in historical contexts). But thousands of end users are going to both ask for and notice things that your testers don't, and that's how you end up here. "This system prompt prevents Gemini from promoting Naziism successfully, ship it!"

This is always going to be a challenge with trying to moderate or put any guardrails on these things. Their behavior is so complex it's almost impossible to reason about all of the consequences, so the only way to "know" is for users to just keep poking at it.

Re: Google to pause Gemini image generation of people after issues

#476
post #277

Here's the problem for Google: Gemini pukes out a perfect visual representation of actual systemic racism that pervades throughout modern corporate culture in the US. Daily interactions can be masked by platitudes and dog whistles. A poster of non-white celtic warriors cannot. Gemini refused to create an image of "a nice white man", saying it was "too spicy", but had no problem when asked for an image of "a nice blac…

Are you seriously claiming that the actual systemic racism in our society is discrimination against white people? I just struggle to imagine someone holding this belief in good faith.

That take is extremely popular on HN

Re: Google to pause Gemini image generation of people after issues

#477
post #402
post #277

Here's the problem for Google: Gemini pukes out a perfect visual representation of actual systemic racism that pervades throughout modern corporate culture in the US. Daily interactions can be masked by platitudes and dog whistles. A poster of non-white celtic warriors cannot. Gemini refused to create an image of "a nice white man", saying it was "too spicy", but had no problem when asked for an image of "a nice blac…

> actual systemic racism that pervades throughout modern corporate culture Ooph. The projection here is just too much. People jumping straight across all the reasonable interpretations straight to the maximal conspiracy theory. Surely this is just a bug. ML has always had trouble with "racism" accusations, but for years it went in the other direction . Remember all the coverage of "I asked for a picture of a criminal…

>But that's not "systemic racism"

When you filter results to prevent it from showing white males, that is by definition system racism. And that's what's happening.

>Surely this is just a bug

Having you been living under a rock for the last 10 years?

Re: Google to pause Gemini image generation of people after issues

#478
post #277

Here's the problem for Google: Gemini pukes out a perfect visual representation of actual systemic racism that pervades throughout modern corporate culture in the US. Daily interactions can be masked by platitudes and dog whistles. A poster of non-white celtic warriors cannot. Gemini refused to create an image of "a nice white man", saying it was "too spicy", but had no problem when asked for an image of "a nice blac…

I thought you were going to say anti-white racism.

Re: Google to pause Gemini image generation of people after issues

#479
post #277

Here's the problem for Google: Gemini pukes out a perfect visual representation of actual systemic racism that pervades throughout modern corporate culture in the US. Daily interactions can be masked by platitudes and dog whistles. A poster of non-white celtic warriors cannot. Gemini refused to create an image of "a nice white man", saying it was "too spicy", but had no problem when asked for an image of "a nice blac…

There is an _actual problem_ that needs to be solved. If you ask generative AI for a picture of a "nurse", it will produce a picture of a white woman 100% of the time, without some additional prompting or fine tuning that encourages it to do something else. If you ask a generative AI for a picture of a "software engineer", it will produce a picture of a white guy 100% of the time, without some additional prompting or…

I think this is a much more tractable problem if one doesn't think in terms of diversity with respect to identify-associated labels, but thinks in terms of diversity of other features.

Consider the analogous task "generate a picture of a shirt". Suppose in the training data, the images most often seen with "shirt" without additional modifiers is a collared button-down shirt. But if you generate k images per prompt, generating k button-downs isn't the most likely to result in the user being satisfied; hedging your bets and displaying a tee shirt, a polo, a henley (or whatever) likely increases the probability that one of the photos will be useful. But of course, if you query for "gingham shirt", you should probably only see button-downs, b/c though one could presumably make a different cut of shirt from gingham fabric, the probability that you wanted a non-button-down gingham shirt but _did not provide another modifier_ is very low.

Why is this the case (and why could you reasonably attempt to solve for it without introducing complex extra user controls)? A _use-dependent_ utility function describes the expected goodness of an overall response (including multiple generated images), given past data. Part of the problem with current "demo" multi-modal LLMs is that we're largely just playing around with them.

This isn't specific to generational AI; I've seen a similar thing in product-recommendation and product search. If in your query and click-through data, after a user searches "purse" if the results that get click-throughs are disproportionately likely to be orange clutches, that doesn't mean when a user searches for "purse", the whole first page of results should be orange clutches, because the implicit goal is maximizing the probability that the user is shown a product that they like, but given the data we have uncertainty about what they will like.

Re: Google to pause Gemini image generation of people after issues

#480
post #277

Here's the problem for Google: Gemini pukes out a perfect visual representation of actual systemic racism that pervades throughout modern corporate culture in the US. Daily interactions can be masked by platitudes and dog whistles. A poster of non-white celtic warriors cannot. Gemini refused to create an image of "a nice white man", saying it was "too spicy", but had no problem when asked for an image of "a nice blac…

Are you seriously claiming that the actual systemic racism in our society is discrimination against white people? I just struggle to imagine someone holding this belief in good faith.

I think it's obviously one of the problems.
Post reply on HN