Do you think these very small models have some utility in the real world? Apart from learning and academic purposes of course.
Sure, interacting with natural language without expectation that the model contains knowledge. Good for things like tool use and embeddings where the information is all retrieved.
Gemma 3 270M re-implemented in pure PyTorch for local tinkering
11–20 of 62 posts
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#12Hey all, I created this model with a top notch team. I answered many questions last week when this hit the front page, and happy to answer more here as well. https://news.ycombinator.com/item?id=44902148 Personally I'm excited that you all have access to this model now and hope you all get value out of using them.
I would like to know your thoughts on using 2/3 of such a small the model's size for embeddings. What would be different if you used a byte-level vocabulary and spent the parameter budget on transformer parameters instead? I think you would lose performance (tok/s) but might gain accuracy.
The tokens themselves are a form of compression. Lets say we have the word "WaffleHouse", character level this would be 11 tokens, but with an embedder this would be perhaps 2 or 3 tokens (I didn't actually run through the tokenizer but we could verify precisely). This matters a lot for on device processing especially.
So while we could get more intelligence out of the model by bumping up the "knowledge" parameters, the device would need to process more input and output tokens.
Another advantage on small devices is the embeddings are just a lookup table which requires little to no computation. Its the rest of the parameters that have the expensive matrix multplications, so if we increased those we'd also be increasing the number of FLOPs needed for a forward pass.
This blog post explains it well. https://www.adamcasson.com/posts/transformer-flops
So all this to say is there are definite tradeoffs between model size, performance on evals, and compute cost. We ran many internal experiments with different choices to see could work well, and then picked what we believed work will best for the open community.
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#13Earlier quoted context omitted.
Sure, interacting with natural language without expectation that the model contains knowledge. Good for things like tool use and embeddings where the information is all retrieved.
Are these small models are trained to privilege "raw intelligence" over factual knowledge? Is there any indication of how much of current model is dedicated to the knowledge of multiple languages and tons of facts rather than pure understanding and reasoning?
To answer a question you didn't ask. With small models especially we need to make choices as to which to focus on. For this model we focused on text summarization and instruction following, with the idea that users would finetune to gain performance on the task set that is relevant to them
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#14Hey all, I created this model with a top notch team. I answered many questions last week when this hit the front page, and happy to answer more here as well. https://news.ycombinator.com/item?id=44902148 Personally I'm excited that you all have access to this model now and hope you all get value out of using them.
I imagine you and your team have finetuned the model on different tasks, can you share some results? (I have only seen the alien NPC finetuning)
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#15That’s wild that with a KV cache and compilation on the Mac CPU you are faster than on an A100 GPU.
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#16Hey all, I created this model with a top notch team. I answered many questions last week when this hit the front page, and happy to answer more here as well. https://news.ycombinator.com/item?id=44902148 Personally I'm excited that you all have access to this model now and hope you all get value out of using them.
Can you share what kind of hardware is necessary to train it, and how long it took?
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#17That’s wild that with a KV cache and compilation on the Mac CPU you are faster than on an A100 GPU.
(My first guess)
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#18Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#19Do we need to fine-tune for the embeddings to be meaningful at the sentence/document level?
Re: Gemma 3 270M re-implemented in pure PyTorch for local tinkering
#20Do you think these very small models have some utility in the real world? Apart from learning and academic purposes of course.
You
Translate to English: Рошешуа́р (фр. Rochechouart) — місто, муніципалітет у Франції, у регіоні Нова Аквітанія, департамент Верхня В'єнна. Населення — 3637 осіб (2022)[1]. Муніципалітет розташований на відстані[2] близько 360 км на південь від Парижа, 34 км на захід від Ліможа.
gemma-3-270m-it-qat
Here's the translation:
"Rochechouart is a city in France, a municipality located in the region of New Aeterna, department within the Upper Vienne Province. The population is 3637 people (2022). The city is situated at a distance of approximately 360 kilometers from Paris, 34 km from Lyon.
You
Translate to English: Bộ Công Thương đề xuất hỗ trợ hộ gia đình lắp điện mặt trời mái nhà, có hệ thống lưu trữ tối đa 2-3 triệu đồng, hỗ trợ lãi suất vay trong tối đa ba năm. Nội dung được Bộ Công Thương đưa ra tại dự thảo lần 3 Quyết định của Thủ tướng về chính sách hỗ trợ cho hộ gia đình lắp đặt điện mặt trời mái nhà tự sản, tự tiêu và hệ thống lưu trữ điện.
gemma-3-270m-it-qat
The Ministry of Industry and Trade proposes supporting households with solar panels in their homes, with a maximum savings of 2-3 million đồng, providing loan assistance for up to 3 years. The content of the Ministry's decision at the Third Review of the Government's Policy on Support for Households Installing Solar Panels in Homes, approved by the Prime Minister about the policy support for households installing solar panels themselves, and the system storage capacity is maximum 2-3 million đồng.
1: https://uk.wikipedia.org/wiki/%D0%A0%D0%BE%D1%88%D0%B5%D1%88...2: https://vnexpress.net/lap-dien-mat-troi-mai-nha-tu-dung-co-t...