Live data from Hacker News

Gemma: New Open Models

blog.google

391–400 of 543 posts

Re: Gemma: New Open Models

#391

Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind.

Hi alekandreev,

Any reason you decided to go with a token vocabulary size of 256k? Smaller vocab/vector sizes like most models in this size seem to be using (~16-32k) are much easier to work with. Would love to understand the technical reasoning here that isn't detailed in the report unfortunately :(.

Re: Gemma: New Open Models

#392
"Carefully tested prompts" sounds a lot like "these are the lotto numbers we know are right" kind of thing? How in the world are these things used for anything programmatically deterministic?

Re: Gemma: New Open Models

#393
post #335

Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind.

Thank you very much for releasing these models! It's great to see Google enter the battle with a strong hand. I'm wondering if you're able to provide any insight into the below hyperparameter decisions in Gemma's architecture, as they differ significantly from what we've seen with other recent models? * On the 7B model, the `d_model` (3072) is smaller than `num_heads * d_head` (16*256=4096). I don't know of any other…

I would love answers to these questions too, particularly on the vocab size

Re: Gemma: New Open Models

#394

Earlier quoted context omitted.

You said a lot of nothing without actually saying specifically what the problem is with the recent license. Maybe the license is fine for almost all usecases and the limitations are small? For example, you complained about metas license, but basically everyone uses those models and is completely ignoring it. The weights are out there, and nobody cares what the fine print says. Maybe if you are a FAANG, company, meta…

I specifically called out the claims of openness and doublespeak being used. Google is making claims that are untrue. Meta makes similar false claims. The fact that unspecified "other" people are ignoring the licenses isn't relevant. Good for them. Good luck making anything real or investing any important level of time or money under those misconceptions. "They haven't sued yet" isn't some sort of validation. Anyone…

> Anyone building an actual product that makes actual money that comes to the attention of Meta or Google will be sued into oblivion

No they won't and they haven't.

Almost the entire startup scene is completely ignoring all these licenses right now.

This is basically the entire industry. We are all getting away with it.

Here's an example, take llama.

Llama originally disallowed commercial activity. But then the license got changed much later.

So, if you were a stupid person, then you followed the license and fell behind. And if you were smart, you ignored it and got ahead of everyone else.

Which, in retrospect was correct.

Because now the license allows commerical activity, so everyone who ignores it in the first place got away with it and is now ahead of everyone else.

> won't sue is beyond foolish

But we already got away with it with llama! That's already over! It's commerical now, and nobody got sued! For that example, the people who ignored the license won.

Re: Gemma: New Open Models

#395

Earlier quoted context omitted.

Because the wrongness is intentional.

Is it intentional? You think they intentionally made it not understand skin tone distribution by country? I would believe it if there was proof, but with all the other things it gets wrong it's weird to jump to that conclusion. There's way too much politics in these things. I'm tired of people pushing on the politics rather than pushing for better tech.

I mean, I asked it for a samurai from a specific Japanese time period and it gave me a picture of a "non-binary indigenous American woman" (its words, not mine) so I think there is something intentional going on.

Re: Gemma: New Open Models

#396
post #64

Great! Google is now participating in the AI race to zero with Meta, as predicted that $0 free AI models would eventually catch up against cloud-based ones. You would not want to be in the middle of this as there is no moat around this at all. Not even OpenAI.

If meta keeps spending tens of millions of dollars each year to release free AI models it might seem like there is no moat, but under normal circumstances wouldn't the cost to develop a free model be considered a moat?

> If meta keeps spending tens of millions of dollars each year to release free AI models it might seem like there is no moat,

As well as the point being that Meta (and Google) is removing the 'moat' from OpenAI and other cloud-only based models.

> but under normal circumstances wouldn't the cost to develop a free model be considered a moat?

Yes. Those that can afford to spend tens of millions of dollars to train free models can do so and have a moat to reduce the moats of cloud-based models.

Re: Gemma: New Open Models

#397
post #383

Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind.

> We are really excited to answer any questions you may have about our models. I cannot count how many times I've seen similar posts on HN, followed by tens of questions from other users, three of which actually get answered by the OP. This one seems to be no exception so far.

What are you talking about? The team is in this thread answering questions.

Re: Gemma: New Open Models

#398
post #243
post #229

The terms of use: https://ai.google.dev/gemma/terms and https://ai.google.dev/gemma/prohibited_use_policy Something that caught my eye in the terms: > Google may update Gemma from time to time, and you must make reasonable efforts to use the latest version of Gemma. One of the biggest benefits of running your own model is that it can protect you from model updates that break your carefully tested prompts, so I’m not…

This is actually not that unusual. Stable Diffusion's license, CreativeML Open RAIL-M, has the exact same clause: "You shall undertake reasonable efforts to use the latest version of the Model." Obviously updating the model is not very practical when you're using finetuned versions, and people still use old versions of Stable Diffusion. But it does make me fear the possibility that if they ever want to "revoke" every…

Switching to a model that is functionally useless doesn't seem to fall under "reasonable efforts" to me, but IANAL.

Re: Gemma: New Open Models

#399
post #348

Earlier quoted context omitted.

Gemini. I first asked it to tell me about the Heian period (which it got correct) but then it generated images and seemed to craft the rest of the chat to fit that narrative. I mean, just asking it for a "samurai" from the period will give you this: https://g.co/gemini/share/ba324bd98d9b >A non-binary Indigenous American samurai It seems to recognize it's mistakes if you confront it though. The more I mess with it th…

Got it. I asked it a series of text questions about the period and it didn't put in anything obviously laughable (including when I drilled down into specific questions about the population, gender roles, and ethnicity). Maybe it's the image creation that throws it into lala land.

I think so too. I could be wrong but I believe once it generates an image it tries to work with it. Crazy how it seems the "text" model knows how wildly wrong it is but the image model just does its thing. I asked it why it generated a native American and it ironically said "I can't generate an image of a native american samurai because that would be offensive"

Re: Gemma: New Open Models

#400

Hello on behalf of the Gemma team! We are really excited to answer any questions you may have about our models. Opinions are our own and not of Google DeepMind.

EDIT: it seems this is likely an Ollama bug, please keep that in mind for the rest of this comment :) I ran Gemma in Ollama and noticed two things. First, it is slow. Gemma got less than 40 tok/s while Llama 2 7B got over 80 tok/s. Second, it is very bad at output generation. I said "hi", and it responded this: ``` Hi, . What is up? melizing with you today! What would you like to talk about or hear from me on this fi…

I was going to try these models with Ollama. Did you use a small number of bits/quantization?
Post reply on HN