Live data from Hacker News

Gemma 3 Technical Report [pdf]

storage.googleapis.com

61–70 of 260 posts

Re: Gemma 3 Technical Report [pdf]

#61
post #57

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

Have you considered that selection of material contributes to specialization and efficiency? This is meant to be a weights-small model.

its also apparently a well known result that filtering nsfw content IMPROVES scores

https://x.com/swyx/status/1661359483447316480

Re: Gemma 3 Technical Report [pdf]

#62

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

[flagged]

Hard to get more puritanical than "if you disagree with my opinion then you're morally repulsive". Not to mention that your argument implies that all traces of sex ought to be scrubbed from the entire Internet? And that that conclusion is the only moral one?

Re: Gemma 3 Technical Report [pdf]

#63

Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own and not of Google DeepMind.) PS we are hiring: https://boards.greenhouse.io/deepmind/jobs/6590957

Thanks, been using Gemma 2 a lot at home as it still holds up very well and the 9B version runs great on my 2080Ti. Strong prompt adherence coupled with overall capability makes it very useful. Looking forward to trying Gemma 3. I have some dumb questions though, might as well ask. How do you decide on the model sizes? And how do you train them? Independently or are they related somehow?

Picking model sizes is not an exact science. We look for sizes that will fit quantized on different categories on devices (e.g., low-end and high-end smartphone, laptops and 16GB GPUs, and bigger GPUs/TPUs). We also want the ratio of model width to depth (number of layers) to be consistently around 90, which we found works best.

The models are trained with distillation from a bigger teacher. We train them independently, but for v3 we have unified the recipes for 4B-27B, to give you more predictably when scaling up and down to different model sizes.

Re: Gemma 3 Technical Report [pdf]

#64

Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own and not of Google DeepMind.) PS we are hiring: https://boards.greenhouse.io/deepmind/jobs/6590957

will there ever be a Gemma 3 Thinking? how copyable is the Flash Thinking approach to the Gemma series?

Re: Gemma 3 Technical Report [pdf]

#65

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

Everyone is treating this like corps have anything to gain from an open uncensored model. Switch your view and give me a single argument for it? That random nerds on HN stop jerking each other about what „open“ means? You are just not their target group. Having this discussion every time no matter if the model released is censored or not is just insanity. Bring new arguments or don’t use the models you don’t like. Th…

But who is the target group?

Last time only some groups of enthusiasts were willing to work through bugs to even run the buggy release of Gemma

Surely nobody runs this in production

Re: Gemma 3 Technical Report [pdf]

#66

Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own and not of Google DeepMind.) PS we are hiring: https://boards.greenhouse.io/deepmind/jobs/6590957

Thank you!

Question: your model supports 140 languages. Given that you are focusing on compactness and efficiency, would you not have gains in also developing models on a selected limited number of languages (e.g. the topmost (in cultural production) four "western" ones with shared alphabet - or similar set)?

Edit: of course the multilingual capability can be can be welcome. On the other hand, there are evident cases in which efficiency can be paramount. We can wonder about the tradeoff: how much in efficiency is sacrificed by features.

Re: Gemma 3 Technical Report [pdf]

#67
post #24

Earlier quoted context omitted.

This could be a historical accident. Early models were censored, making uncensored releases have bad optics. If the first models had been uncensored, no one would care if another was added.

Have an uncensored model loop through nypost articles and ask it to synthesize content from that. Nypost has tons of scandalous content and can easily get spun into erotica by an uncensored model. It’s unsafe for that reason, so you absolutely needed both censored and uncensored. It wasn’t an accident.

> can easily get spun into erotica by an uncensored model.

A sexualized fine-tune yes, but that's because you have to make them overly horny to overcome the original censorship.

Nothing prevent them to train a model that will have an appropriate level of sexual content (that is, only upon user explicit request) the same way they train it not to have sexual content at all.

The reason they do that is because they are American companies, the same companies who also censored nude paintings and statues from European museums' pages.

Re: Gemma 3 Technical Report [pdf]

#68
post #61
post #57

Earlier quoted context omitted.

Have you considered that selection of material contributes to specialization and efficiency? This is meant to be a weights-small model.

its also apparently a well known result that filtering nsfw content IMPROVES scores https://x.com/swyx/status/1661359483447316480

LLMs get distracted by porn too !?!?

Re: Gemma 3 Technical Report [pdf]

#69

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

Everyone is treating this like corps have anything to gain from an open uncensored model. Switch your view and give me a single argument for it? That random nerds on HN stop jerking each other about what „open“ means? You are just not their target group. Having this discussion every time no matter if the model released is censored or not is just insanity. Bring new arguments or don’t use the models you don’t like. Th…

This is what HNers surprisingly seem to not understand.

The risk of the model generating illegal content and then the company getting bad PR from vultures in journalism simply outweighs any benefits of including this content in the training data.

This is also why you will never see the big companies release a capable open weight image or video gen model.

Re: Gemma 3 Technical Report [pdf]

#70

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

It's harmful in that there exists a significant and vocal subset of users who does not wish to see that content or does not wish their children to do so. It's easier to teach your model never to produce that kind of content than to teach it to perfectly distinguish whether this user should see that content or not. TV channels are barred from broadcasting this kind of content for similar reasons.

Sure, there are always jailbreaks, but then the narrative changes from "we made a model that tells erotic stories to children" to "this ingenious teenager figured out a way to hack our model to make it produce erotic stories." In other words, jailbreak move the fault from the model producer to the model user.

It's also worth keeping in mind that erotica comprises a surprisingly large portion of fiction easily available on the internet for free, and "unfiltered" models tend to produce that kind of content unprompted (see e.g. the original Mistral). The major AI labs are probably filtering it out, but I suspect they can't go too far there, as having a model that is good at fiction is something they actually want.

Then there are the non-chat-gpt-app use cases (like customer support chatbots, automatic summarization etc), for which unprompted erotica is highly inappropriate. Those are the "business travelers" of AI, not the first thing one thinks of when talking about who uses AI models, but extremely important nonetheless.

Post reply on HN