Live data from Hacker News

Gemma: New Open Models

blog.google

511–520 of 543 posts

Re: Gemma: New Open Models

#511

Earlier quoted context omitted.

Only 8K context as well, like Mistral. Also, as always, take these benchmarks with a huge grain of salt. Even base model releases are frequently (seemingly) contaminated these days.

Mistral Instruct v0.2 is 32K.

original Mistral or GGUF one?

Re: Gemma: New Open Models

#512
post #436

Earlier quoted context omitted.

Have you considered the use of Monte Carlo sampling to inspect latent behaviors?

I think that's the wrong level to attack the problem; you can do that also with actual humans, but it won't tell you what the human is unable to think, but rather what they just didn't think of given their stimulus — and this difference is easily demonstrated, e.g. with Duncker's candle problem: https://en.wikipedia.org/wiki/Candle_problem

I agree that it’s not a complete solution, but this sort of characterization is still useful towards the goal of identifying regions of fitness within the model.

Maybe you can’t explore the entire forest, but maybe you can clear the area around your campsite sufficiently. Even if there are still bugs in the ground.

Re: Gemma: New Open Models

#513

Earlier quoted context omitted.

Of all the very very very many things that Google models get wrong, not understanding nationality and skin tone distributions seems to be a very weird one to focus on. Why are there three links to this question? And why are people so upset over it? Very odd, seems like it is mostly driven by political rage.

Maybe some people care about truth?

If someome's primary concern was truth, wouldn't the many many other flaws also draw their ire?

That's my contention: the focus on this one thing belies not a concern for truth, but a concern for race politics.

Re: Gemma: New Open Models

#514
post #247

Earlier quoted context omitted.

It would be great to understand what you mean by this -- we have a deep love for open source and the open developer ecosystem. Our open source team also released a blog today describing the rationale and approach for open models and continuing AI releases in the open ecosystem: https://opensource.googleblog.com/2024/02/building-open-mode... Thoughts and feedback welcome, as always.

If you truly love Open Source, you should update the the language you use to describe your models so it doesn't mislead people into thinking it has something to do with Open Source. Despite being called "Open", the Gemma weights are released under a license that is incompatible with the Open Source Definition. It has more in common with Source-Available Software, and as such it should be called a "Weights-Available M…

Open source is not defined as strictly as what you are suggesting it is. If you wish to have a stricter definition, a new term should probably be used. I believe I've heard it referred to as libre software in the past

Re: Gemma: New Open Models

#515
I find it a bit disheartening that the new wave of “open source” with regards to AI is open weights. That’s like giving someone compiled and obfuscated binaries and saying that’s open source.

Re: Gemma: New Open Models

#516

Earlier quoted context omitted.

This assumes they even know that the model hasn't been updated. Who is this actually intended for? I'd bet it's for companies hosting the model. In those cases, the definition of reasonable effort is a little closer to "it'll break our stuff if we touch it" rather than "oh silly me, I forgot how to spell r-s-y-n-c".

Hosting companies can probably just claim they're covered under Section 230, and Google has to go bother the individual users, not them.

I don't believe that would apply if the host is curating the models they host.

Re: Gemma: New Open Models

#517

Earlier quoted context omitted.

It seems very easy to check no? Look at the names in the paper and check where they are working now

Good idea. I've confirmed all the leadership / tech leads listed on page 12 are still at Google. Can someone with a Twitter account call out the tweet linked above and ask them specifically who they are referring to? Seems there is no evidence of their claim.

It's also possible Google removed names of people who left. It's not really a research paper, more a marketing piece, so it might be possible (I don't think they would do that with a conf paper)

Re: Gemma: New Open Models

#518
post #247

Earlier quoted context omitted.

If you truly love Open Source, you should update the the language you use to describe your models so it doesn't mislead people into thinking it has something to do with Open Source. Despite being called "Open", the Gemma weights are released under a license that is incompatible with the Open Source Definition. It has more in common with Source-Available Software, and as such it should be called a "Weights-Available M…

Open source is not defined as strictly as what you are suggesting it is. If you wish to have a stricter definition, a new term should probably be used. I believe I've heard it referred to as libre software in the past

"Open Source Software" always refers to software that meets the Open Source Definition. "Libre Software" always refers to software that meets the Free Software Definition. In practice the two are often identical, hence the abbreviations "FOSS" (Free and Open Source Software) and "FLOSS" (Free/Libre and Open Source Software).

Although I don't know Google's motivation for using "Open" to describe proprietary model weights, the practical result is increasing confusion about Open Source Software. It's behavior that benefits any organization wanting to enjoy the good image of the Open Source Software community while not actually caring about that community at all.

Re: Gemma: New Open Models

#519
post #402

Earlier quoted context omitted.

Well we can just ask Gemma to generate images of the meetings, no need to imagine. ;)

I wouldn't be surprised if there were actually only white men in the meeting, as opposed to what Gemini will produce.

> only white men

Why?

Re: Gemma: New Open Models

#520
post #436

Earlier quoted context omitted.

I think that's the wrong level to attack the problem; you can do that also with actual humans, but it won't tell you what the human is unable to think, but rather what they just didn't think of given their stimulus — and this difference is easily demonstrated, e.g. with Duncker's candle problem: https://en.wikipedia.org/wiki/Candle_problem

I agree that it’s not a complete solution, but this sort of characterization is still useful towards the goal of identifying regions of fitness within the model. Maybe you can’t explore the entire forest, but maybe you can clear the area around your campsite sufficiently. Even if there are still bugs in the ground.

I like that metaphor, I hope I remember it.
Post reply on HN