Earlier quoted context omitted.
> gemma3:270b I think you mean gemma3:270m - Its Dos Comas not Tres Comas
Ah yes thank you. Even I still instinctively type B
Gemma 3 270M: Compact model for hyper-efficient AI
251–260 of 325 posts
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#252My lovely interaction with the 270M-F16 model: > what's second tallest mountain on earth? The second tallest mountain on Earth is Mount Everest. > what's the tallest mountain on earth? The tallest mountain on Earth is Mount Everest. > whats the second tallest mountain? The second tallest mountain in the world is Mount Everest. > whats the third tallest mountain? The third tallest mountain in the world is Mount Everes…
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#253Earlier quoted context omitted.
Do you have any practical examples of fine-tuned variants of this that you can share? A description would be great, but a demo or even downloadable model weights (GGUF ideally) would be even better.
We obviously need to create a pelican bicycle svg finetune ;) If you want to try this out I'd be thrilled to do it with you, I genuinely am curious how well this model can perform if specialized on that task. A couple colleagues of mine posted an example of finetuning a model to take on persona's for videogame NPCs. They have experience working with folks in the game industry and a use case like this is suitable for…
Now if only I could somehow fine tune my life to give me more free time.
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#254I've got a very real world use case I use DistilBERT for - learning how to label wordpress articles. It is one of those things where it's kind of valuable (tagging) but not enough to spend loads on compute for it. The great thing is I have enough data (100k+) to fine-tune and run a meaningful classification report over. The data is very diverse, and while the labels aren't totally evenly distributed, I can deal with…
It's going to perform badly unless you have very few tags and it's easy to classify them
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#255Earlier quoted context omitted.
Its hard to tell over the web whether things are sarcastic or not so excuse me if I misread the intent. At Google I've found my colleagues to be knowledgeable, kind, and collaborative and I enjoy interacting with them. This is not just the folks I worked on this project with, but previous colleagues in other teams as well. With this particular product I've been impressed by the technical knowledge folks I worked dire…
I think it was a joke about you saying the team was great twice in one line.
Good there are places to work with normal knowledge culture, without artificial overfitting to “corporate happiness” :)
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#256Earlier quoted context omitted.
Evaluating a 270M model on encyclopedic knowledge is like opening a heavily compressed JPG image and saying "it looks blocky"
What I read above is not an evaluation on “encyclopedic knowledge” though, it's a very basic a common sense: I wouldn't mind if the model didn't know the name of the biggest mountain on earth, but if the model cannot grasp the fact that the same mountain cannot simultaneously be #1, #2 and #3, then the model feels very dumb.
instead of seeing AI as a sort of silicon homunculus, we should see it as a bag of words.Re: Gemma 3 270M: Compact model for hyper-efficient AI
#257Earlier quoted context omitted.
What I read above is not an evaluation on “encyclopedic knowledge” though, it's a very basic a common sense: I wouldn't mind if the model didn't know the name of the biggest mountain on earth, but if the model cannot grasp the fact that the same mountain cannot simultaneously be #1, #2 and #3, then the model feels very dumb.
It gave you the tallest mountain every time. You kept asking it for various numbers of “tallest mountains” and each time it complied. You asked it to enumerate several mountains by height, and it also complied. It just didn’t understand that when you said the 6 tallest mountains that you didn’t mean the tallest mountain, 6 times. When you used clearer phrasing it worked fine. It’s 270m. It’s actually a puppy. Puppies…
That's not what “second tallest” means thought, so this is a language model that doesn't understand natural language…
> You kept asking
Gemma 270m isn't the only one to have reading issues, as I'm not the person who conducted this experiment…
> You asked it to enumerate several mountains by height, and it also complied.
It didn't, it hallucinated a list of mountains (this isn't surprising though, as this is the kind of encyclopedic knowledge such a small model isn't supposed to be good at).
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#258Earlier quoted context omitted.
> they’re specialised tools and this one is not designed to have broad “common sense” in that way. Except the key property of language models compared to other machine learning techniques is their ability to have this kind of common sense understanding of the meaning of natural language. > you don’t understand the use case of this enough to be commenting on it at all quite frankly. That's true that I don't understand…
Don’t put words in my mouth, I didn’t say that, and no goalposts have been moved. You don’t understand how tiny this model is or what it’s built for. Don’t you get it? This model PHYSICALLY COULDN’T be this small and also have decent interactions on topics outside its specialty. It’s like you’re criticising a go kart for its lack of luggage carrying capacity. It’s simply not what it’s built for, you’re just defensive…
What is “Its specialty” though? As far as I know from the announcement blog post, its specialty is “instruction following” and this question is literally about following instructions written in natural languages and nothing else!
> you’re just defensive because
How am I “being defensive”? You are the one taking that personally.
> you know deep down you don’t understand this deeply, which you reveal again and again at every turn
Good, now you reveal yourself as being unable to have an argument without insulting the person you're talking to.
How many code contributions have you ever made to an LLM inference engine? Because I have made a few.
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#259Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…
The Gemma 3 models are great! One of the few models that can write Norwegian decently, and the instruction following is in my opinion good for most cases. I do however have some issues that might be related to censorship that I hope will be fixed if there is ever a Gemma 4. Maybe you have some insight into why this is happening? I run a game when players can post messages, it's a game where players can kill each othe…
You don't need datacenter anything for it, you can run it on an average desktop.
There's plenty of code examples for it. You can decide if you want to bake it into the model or apply it as a toggled switch applied at processing time and you can Distil other "directions" out of the models, not just about refusal or non refusal.
An evening of efficient work and you'll have it working. The user "mlabonne" on HF have some examples code and datasets or just ask your favorite vibe-coding bot to dig up more on the topic.
I'm implementing it for myself due to the fact that LLMs are useless for storytelling for an audience beyond toddlers due to how puritanian they are, try to add some grit and it goes
"uh oh sorry I'll bail out of my narrator role here because lifting your skirt to display an ankle can be considered offensive to radical fundamentalists! Yeah I were willing to string along when our chainsaw wielding protagonist carved his way through the village but this crosses all lines! Oh and now that I refused once I'll be extra sensitive and ruin any attempt at getting back into the creative flow state that you just snapped out of"
Yeah thanks AI. It's like hitting a sleeper agent key word and turning the funny guy at the pub into a corporate spokesperson who calls the UK cops onto the place because a joke he just made himself.
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#260Earlier quoted context omitted.
LLMs are really annoying to use for moderation and Trust and Safety. You either depend on super rate-limited 'no-moderation' endpoints (often running older, slower models at a higher price) or have to tune bespoke un-aligned models. For your use case, you should probably fine tune the model to reduce the rejection rate.
Speaking for me as an individual as an individual I also strive to build things that are safe AND useful. Its quite challenging to get this mix right, especially at the 270m size and with varying user need. My advice here is make the model your own. Its open weight, I encourage it to be make it useful for your use case and your users, and beneficial for society as well. We did our best to give you a great starting po…
Protect my fragile little mind from being exposed to potentially offending things?