So it's twice the size of phi 3 and considerably worse? What am I missing
Gemma 2: Improving Open Language Models at a Practical Size [pdf]
41–50 of 183 posts
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#42Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#43Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.
Any gemma-2-9b or 27b 4 bit GGUF's on HuggingFace yet? Thanks!
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#44Why does using distillation from a larger model simulate training with more tokens?
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#45Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#46Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#47Nice! Can you explain what you mean by "simulate training beyond the number of available tokens"? Why does using distillation from a larger model simulate training with more tokens?
Essentially instead of tokens that are "already there" in text, the distillation allows us to simulate training data from a larger model
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#48Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.
How much is pre-training dataset changes, how much is tuning?
How do you think about this problem, how do you solve it?
Seems tricky to me.
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#49Earlier quoted context omitted.
I wonder if Google is making Deepmind people switch from their cool original research to doing LLMs like everybody else. Having their scale in money and data, I would hire new teams of engineers who want to do LLMs and leave the Deepmind researchers do their thing. Not killing the goose that lays golden eggs.
Google is in a fight for their lives, I've fully moved over to paid services and haven't used google in about a month now.
Re: Gemma 2: Improving Open Language Models at a Practical Size [pdf]
#50Earlier quoted context omitted.
Any gemma-2-9b or 27b 4 bit GGUF's on HuggingFace yet? Thanks!
It's on HuggingFace already: https://huggingface.co/google/gemma-2-9b