Live data from Hacker News

Gemma 3 270M: Compact model for hyper-efficient AI

developers.googleblog.com

251–260 of 325 posts

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#252
post #64

My lovely interaction with the 270M-F16 model: > what's second tallest mountain on earth? The second tallest mountain on Earth is Mount Everest. > what's the tallest mountain on earth? The tallest mountain on Earth is Mount Everest. > whats the second tallest mountain? The second tallest mountain in the world is Mount Everest. > whats the third tallest mountain? The third tallest mountain in the world is Mount Everes…

This is standup material. Had a hearty laugh, thanks.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#253
post #34

Earlier quoted context omitted.

Do you have any practical examples of fine-tuned variants of this that you can share? A description would be great, but a demo or even downloadable model weights (GGUF ideally) would be even better.

We obviously need to create a pelican bicycle svg finetune ;) If you want to try this out I'd be thrilled to do it with you, I genuinely am curious how well this model can perform if specialized on that task. A couple colleagues of mine posted an example of finetuning a model to take on persona's for videogame NPCs. They have experience working with folks in the game industry and a use case like this is suitable for…

I have so many game ideas that would use a small LLM built up in my brain, so thank you for this.

Now if only I could somehow fine tune my life to give me more free time.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#254
post #101

I've got a very real world use case I use DistilBERT for - learning how to label wordpress articles. It is one of those things where it's kind of valuable (tagging) but not enough to spend loads on compute for it. The great thing is I have enough data (100k+) to fine-tune and run a meaningful classification report over. The data is very diverse, and while the labels aren't totally evenly distributed, I can deal with…

It's going to perform badly unless you have very few tags and it's easy to classify them

You can solve this by training a model per taxonomy, then wrap the individual models into a wrapper model to output joint probabilities. The largest amount of labels I have in a taxonomy is 8.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#255
post #192

Earlier quoted context omitted.

Its hard to tell over the web whether things are sarcastic or not so excuse me if I misread the intent. At Google I've found my colleagues to be knowledgeable, kind, and collaborative and I enjoy interacting with them. This is not just the folks I worked on this project with, but previous colleagues in other teams as well. With this particular product I've been impressed by the technical knowledge folks I worked dire…

I think it was a joke about you saying the team was great twice in one line.

Seems the team and working conditions worth mentioning it twice, nonetheless.

Good there are places to work with normal knowledge culture, without artificial overfitting to “corporate happiness” :)

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#256

Earlier quoted context omitted.

Evaluating a 270M model on encyclopedic knowledge is like opening a heavily compressed JPG image and saying "it looks blocky"

What I read above is not an evaluation on “encyclopedic knowledge” though, it's a very basic a common sense: I wouldn't mind if the model didn't know the name of the biggest mountain on earth, but if the model cannot grasp the fact that the same mountain cannot simultaneously be #1, #2 and #3, then the model feels very dumb.

It does not work that way. The model does not "know". Here is a very nice explanation of what you are actually dealing with (hint: it's not a toddler-level intelligence): https://www.experimental-history.com/p/bag-of-words-have-mer...

    instead of seeing AI as a sort of silicon homunculus, we should see it as a bag of words.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#257
post #232

Earlier quoted context omitted.

What I read above is not an evaluation on “encyclopedic knowledge” though, it's a very basic a common sense: I wouldn't mind if the model didn't know the name of the biggest mountain on earth, but if the model cannot grasp the fact that the same mountain cannot simultaneously be #1, #2 and #3, then the model feels very dumb.

It gave you the tallest mountain every time. You kept asking it for various numbers of “tallest mountains” and each time it complied. You asked it to enumerate several mountains by height, and it also complied. It just didn’t understand that when you said the 6 tallest mountains that you didn’t mean the tallest mountain, 6 times. When you used clearer phrasing it worked fine. It’s 270m. It’s actually a puppy. Puppies…

> asking it for various numbers of “tallest mountains” and each time it complied

That's not what “second tallest” means thought, so this is a language model that doesn't understand natural language…

> You kept asking

Gemma 270m isn't the only one to have reading issues, as I'm not the person who conducted this experiment…

> You asked it to enumerate several mountains by height, and it also complied.

It didn't, it hallucinated a list of mountains (this isn't surprising though, as this is the kind of encyclopedic knowledge such a small model isn't supposed to be good at).

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#258

Earlier quoted context omitted.

> they’re specialised tools and this one is not designed to have broad “common sense” in that way. Except the key property of language models compared to other machine learning techniques is their ability to have this kind of common sense understanding of the meaning of natural language. > you don’t understand the use case of this enough to be commenting on it at all quite frankly. That's true that I don't understand…

Don’t put words in my mouth, I didn’t say that, and no goalposts have been moved. You don’t understand how tiny this model is or what it’s built for. Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty. It’s like you’re criticising a go kart for its lack of luggage carrying capacity. It’s simply not what it’s built for, you’re just defensive…

> Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty

What is “Its specialty” though? As far as I know from the announcement blog post, its specialty is “instruction following” and this question is literally about following instructions written in natural languages and nothing else!

> you’re just defensive because

How am I “being defensive”? You are the one taking that personally.

> you know deep down you don’t understand this deeply, which you reveal again and again at every turn

Good, now you reveal yourself as being unable to have an argument without insulting the person you're talking to.

How many code contributions have you ever made to an LLM inference engine? Because I have made a few.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#259

Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…

The Gemma 3 models are great! One of the few models that can write Norwegian decently, and the instruction following is in my opinion good for most cases. I do however have some issues that might be related to censorship that I hope will be fixed if there is ever a Gemma 4. Maybe you have some insight into why this is happening? I run a game when players can post messages, it's a game where players can kill each othe…

The magic word you want to look up here is "LLM abliteration", it's the concept of where you can remove, attenuate or manipulate the refusal "direction" of a model.

You don't need datacenter anything for it, you can run it on an average desktop.

There's plenty of code examples for it. You can decide if you want to bake it into the model or apply it as a toggled switch applied at processing time and you can Distil other "directions" out of the models, not just about refusal or non refusal.

An evening of efficient work and you'll have it working. The user "mlabonne" on HF have some examples code and datasets or just ask your favorite vibe-coding bot to dig up more on the topic.

I'm implementing it for myself due to the fact that LLMs are useless for storytelling for an audience beyond toddlers due to how puritanian they are, try to add some grit and it goes

"uh oh sorry I'll bail out of my narrator role here because lifting your skirt to display an ankle can be considered offensive to radical fundamentalists! Yeah I were willing to string along when our chainsaw wielding protagonist carved his way through the village but this crosses all lines! Oh and now that I refused once I'll be extra sensitive and ruin any attempt at getting back into the creative flow state that you just snapped out of"

Yeah thanks AI. It's like hitting a sleeper agent key word and turning the funny guy at the pub into a corporate spokesperson who calls the UK cops onto the place because a joke he just made himself.

Re: Gemma 3 270M: Compact model for hyper-efficient AI

#260

Earlier quoted context omitted.

LLMs are really annoying to use for moderation and Trust and Safety. You either depend on super rate-limited 'no-moderation' endpoints (often running older, slower models at a higher price) or have to tune bespoke un-aligned models. For your use case, you should probably fine tune the model to reduce the rejection rate.

Speaking for me as an individual as an individual I also strive to build things that are safe AND useful. Its quite challenging to get this mix right, especially at the 270m size and with varying user need. My advice here is make the model your own. Its open weight, I encourage it to be make it useful for your use case and your users, and beneficial for society as well. We did our best to give you a great starting po…

What does safe even mean in the context of a locally running LLM?

Protect my fragile little mind from being exposed to potentially offending things?

Post reply on HN