Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…
hi, congrats for the amazing work! i love the 27b model, and i use it basically daily. however when i tried to finetune it for a task in a low resource language, unfortunately i did not succeed: lora just did not picked up the gist of the task, full finetune lead to catastrophic forgetting. may i ask four your advice, or do you have any general tips how to do that properly? thanks in advance for your help :)
Gemma 3 270M: Compact model for hyper-efficient AI
181–190 of 325 posts
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#182Earlier quoted context omitted.
Do you know that hardware required to fine-tune this model? I'm asking on behave of us GPU starve folks
A free colab. Here's a link, you can finetune the model in ~5 minutes in this example, and I encourage you to try your own https://ai.google.dev/gemma/docs/core/huggingface_text_full_...
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#183My lovely interaction with the 270M-F16 model: > what's second tallest mountain on earth? The second tallest mountain on Earth is Mount Everest. > what's the tallest mountain on earth? The tallest mountain on Earth is Mount Everest. > whats the second tallest mountain? The second tallest mountain in the world is Mount Everest. > whats the third tallest mountain? The third tallest mountain in the world is Mount Everes…
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#184Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…
Or am I so far behind that "fine tuning your own model" is something a 12 year old who is married to chatGPT does now?
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#185> this model is not designed for complex conversational use cases ... but it's also the perfect choice for creative writing ...? Isn't this a contradiction? How can a model be good at creative writing if it's no good at conversation?
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#186Earlier quoted context omitted.
Well, this is a 270M model which is like 1/3 of 1B parameters. In the grand scheme of things, it's basically a few matrix multiplications, barely anything more than that. I don't think it's meant to have a lot of knowledge, grammar, or even coherence. These input: ``` Customer Review says: ai bought your prod-duct and I wanna return becaus it no good. Prompt: Create a JSON object that extracts information about this…
If it didn't know how to generate the list from 1 to 5 then I would agree with you 100% and say the knowledge was stripped out while retaining intelligence - beautiful. But the fact that it does, but cannot articulate the (very basic) knowledge it has *and* in the same chat context when presented with (its own) list of mountains from 1 to 5 that it cannot grasp it made a LOGICAL (not factual) error in repeating the r…
These words do not mean what you think they mean when used to describe an LLM.
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#187Earlier quoted context omitted.
I see you are using ollamas ggufs. By default it will download Q4_0 quantization. Try `gemma3:270m-it-bf16` instead or you can also use unsloth ggufs `hf.co/unsloth/gemma-3-270m-it-GGUF:16` You'll get better results.
Good call, I'm trying that one just now in LM Studio (by clicking "Use this model -> LM Studio" on https://huggingface.co/unsloth/gemma-3-270m-it-GGUF and selecting the F16 one). (It did not do noticeably better at my pelican test). Actually it's worse than that, several of my attempts resulted in infinite loops spitting out the same text. Maybe that GGUF is a bit broken?
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#188Earlier quoted context omitted.
Do you have any practical examples of fine-tuned variants of this that you can share? A description would be great, but a demo or even downloadable model weights (GGUF ideally) would be even better.
We obviously need to create a pelican bicycle svg finetune ;) If you want to try this out I'd be thrilled to do it with you, I genuinely am curious how well this model can perform if specialized on that task. A couple colleagues of mine posted an example of finetuning a model to take on persona's for videogame NPCs. They have experience working with folks in the game industry and a use case like this is suitable for…
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#189Apple should be doing this. Unless their plan is to replace their search deal with an AI deal -- it's just crazy to me how absent Apple is. Tim Cook said, "it's ours to take" but they really seem to be grasping at the wind right now. Go Google!
Think of Apple however you want, but they rarely ship bad/half-baked products. They would rather not ship a product at all than ship something that's not polished.
Re: Gemma 3 270M: Compact model for hyper-efficient AI
#190Hi all, I built these models with a great team. They're available for download across the open model ecosystem so give them a try! I built these models with a great team and am thrilled to get them out to you. From our side we designed these models to be strong for their size out of the box, and with the goal you'll all finetune it for your use case. With the small size it'll fit on a wide range of hardware and cost…