Live data from Hacker News

Gemma 3 Technical Report [pdf]

storage.googleapis.com

151–160 of 260 posts

Re: Gemma 3 Technical Report [pdf]

#151

Gemma 3 is out! Multimodal (image + text), 128K context, supports 140+ languages, and comes in 1B, 4B, 12B, and 27B sizes with open weights & commercial use. Gemma 3 model overview: https://ai.google.dev/gemma/docs/core Huggingface collection: https://huggingface.co/collections/google/gemma-3-release-67... ollama: https://ollama.com/library/gemma3

> open weights

What exactly is this supposed to mean? That I can grab the weights by just downloading them, or something like that?

Because when I open up the HuggingFace repository, it asks me to "accept the conditions" (Google’s usage license). How is this different from any other proprietary binaries people distribute on the internet but let you run locally? Are other software (like 1Password for example) also "open software" because you can download it?

Re: Gemma 3 Technical Report [pdf]

#152

Gemma 3 is out! Multimodal (image + text), 128K context, supports 140+ languages, and comes in 1B, 4B, 12B, and 27B sizes with open weights & commercial use. Gemma 3 model overview: https://ai.google.dev/gemma/docs/core Huggingface collection: https://huggingface.co/collections/google/gemma-3-release-67... ollama: https://ollama.com/library/gemma3

Doesn't yet work in LM Studio. Barfs an error when trying to load the model. (Error 6, whatever that means. Happy I missed the first 5.)

> Barfs an error when trying to load the model

Since you're not using the official models (since they're not GGUFs), what exact model are you trying to use? The 3rd party you rely on might have screwed something up.

Re: Gemma 3 Technical Report [pdf]

#153
post #129

Earlier quoted context omitted.

Porn sites are blocked in many jurisdictions, so I would not use that argument.

No, there's no movement to shut down pornography on the internet. There's a movement to shut down specific websites and make a lot of noise about it but continue consuming pornography behind closed doors. People like pornography. They'll as soon ban alcohol again (which worked so well last time)

there are.

Re: Gemma 3 Technical Report [pdf]

#154
post #62

Earlier quoted context omitted.

Hard to get more puritanical than "if you disagree with my opinion then you're morally repulsive". Not to mention that your argument implies that all traces of sex ought to be scrubbed from the entire Internet ? And that that conclusion is the only moral one?

There are no Puritans and haven’t been for a few centuries. You’re screaming at ghosts. He or she may be Muslim. You should respect the culture.

The term "puritanical" doesn't exclusively refer to the existence of or adherence to the Puritan religion, and when it does, the term is usually capitalized. From Dictionary.com:

    puritanical [pyoor-i-tan-i-kuhl] adjective

    1) very strict in moral or religious matters, often excessively so; rigidly austere.
    
    2) Sometimes Puritanical. of, relating to, or characteristic of Puritans or Puritanism.
It is entirely possible within the parameters of commonly understood English parlance for Muslims, or any group, to be puritanical.

Re: Gemma 3 Technical Report [pdf]

#155

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

Not all sexually explicit content is harmful in all contexts for sure, but in many contexts it is fairly universally considered harmful (eg content involving minors). Do you have means of distinguishing between the two? Are you suggesting that a company must invests millions into teaching the model where exactly the red line lines so that it can have a conversation close to it but without crossing it? Or you suggest…

I don't agree that textual, fictional explicit content involving minors is "fairly universally considered harmful". Such content is allowed on large platforms like Archive of Our Own or Japan's Shosetsuka ni Naro. I think "don't think it's harmful, but not willing to defend" is a pretty typical attitude.

Re: Gemma 3 Technical Report [pdf]

#156

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

The solution to this problem is to make it not work. If there are various technological developments in the world that do and don't have porn, and if such were cases that the common denominator of failures were lack of smoothly graduated spectrum of contents without disruption from casual family safe content to hardcore pornography, the problem will correct itself.

Actually, it will happen naturally and eventually. Just look at Apple Vision Pro which still don't have VRChat support, and compare how deeply DOA it has been to other VR headsets that are clearly nowhere near as important. Or "Metaverse" that were all explicitly SFW.

This effect can even be seen in the Apple App Store itself. Who uses App Store? You flow into App Store through porn-enabled platforms, such as web or social media. No one browses App Store as a content. What does it not have? Pornography.

Re: Gemma 3 Technical Report [pdf]

#157

Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own and not of Google DeepMind.) PS we are hiring: https://boards.greenhouse.io/deepmind/jobs/6590957

As per the technical report, every 5 layers you have a global attention layer. The global attention layer during training can have as many as a 128k context length during training (though I understand it is usually 32k).

Q. When you are training with a context length of 128k, is the attention in the global layers dense or sparse ?

If dense, would the attention memory requirement here would be O(n^2) where n is 128k for each global layer ?

Re: Gemma 3 Technical Report [pdf]

#158
post #151

Gemma 3 is out! Multimodal (image + text), 128K context, supports 140+ languages, and comes in 1B, 4B, 12B, and 27B sizes with open weights & commercial use. Gemma 3 model overview: https://ai.google.dev/gemma/docs/core Huggingface collection: https://huggingface.co/collections/google/gemma-3-release-67... ollama: https://ollama.com/library/gemma3

> open weights What exactly is this supposed to mean? That I can grab the weights by just downloading them, or something like that? Because when I open up the HuggingFace repository, it asks me to "accept the conditions" (Google’s usage license). How is this different from any other proprietary binaries people distribute on the internet but let you run locally? Are other software (like 1Password for example) also "op…

Replace "google" with "unsloth" in the browser address bar if you want to download them without signing up to hf

Re: Gemma 3 Technical Report [pdf]

#159

> They are designed to help prevent our models from generating harmful content, i.e., > [...] > Sexually explicit content Dear tech companies. Sexually explicit content is not harmful. Why are you all run by puritans? I don't even want to make edgy porn, I just want to be treated like an adult.

Generating sexually explicit content can cause reputational damage or have legal risk. Not generating such content is something that many developers are looking for. There is people who may want such harmful content and other players can cover such a niche.

I don't think it's reputation risk of companies at large, but risk to individual developers. "He worked on porn" is such an easy gut logic for terminations. It's in our human instincts. Everyone know that in guts.

Re: Gemma 3 Technical Report [pdf]

#160

Lots to be excited about here - in particular new architecture that allows subquadratic scaling of memory needs for long context; looks like 128k+ context is officially now available on a local model. The charts make it look like if you have the RAM the model is pretty good out to 350k or so(!) with RoPE. In addition, it flavor tests well on chat arena, ELO significantly above yesterday’s best open model, Qwen 2.5 72…

Gemma is made by Google, not DeepMind.

edit: Sorry, forgot DeepMind was Google's AI R&D, I read it as deepseek in your comment.

Post reply on HN