Viewing profile — alekandreev
alekandreev
HN member- Joined
- Tue, Feb 13, 2024, 9:08 PM UTC
- HN karma
- 368
- Public activity
- 31 items
- HN profile
- View on Hacker News ↗
About alekandreev
No profile information was provided.
Recent public activity
-
comment
Comment #43343940
Yes we have measured the tradeoff. We don't see a drop of perplexity in English when introducing multilingual, and there is a slight drop in some English language-specific evals (~…
-
comment
Comment #43343934
We never train at 128k, only 32k, changing the scaling factor at the end. We wanted the long context recipe to be friendly for finetuning, and training at 128k is a bit of a pain w…
-
comment
Comment #43343506
We recommend using user for the system prompt as well.
-
comment
Comment #43342768
That's a very interesting area, but nothing we can announce today.
-
comment
Comment #43341721
Hey, Gemma engineer here. Can you please share reports on the type of prompts and the implementation you used?
-
comment
Comment #43341711
Thank you for the report! We are working with the Ollama team directly and will look into it.
-
comment
Comment #43341422
Thank you for the feedback! This is why we are so excited to push more and more on small models for both low end and high end smartphones!
-
comment
Comment #43341298
That's an idea we've thought about. However, we think the open source community has already created a very impressive set of language or region-specific finetunes [1] [2]. Also the…
- comment
-
comment
Comment #43341030
Picking model sizes is not an exact science. We look for sizes that will fit quantized on different categories on devices (e.g., low-end and high-end smartphone, laptops and 16GB G…
-
comment
Comment #43340828
Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own …
-
comment
Comment #40813599
I think it makes sense to compare models trained with the same recipe on token count - usually more tokens will give you a better model. However, I wouldn't draw conclusions about …
-
comment
Comment #40813091
To quote Ludovic Peran, our amazing safety lead: Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of f…
-
comment
Comment #40812188
Happy to pass on any feedback to our Google Cloud friends. :)
-
comment
Comment #40812177
This is mostly about inference speed, while maintaining long context performance.
-
comment
Comment #40812158
In addition to the HF links shared by sibling comments, the 2B will be released soon.
-
comment
Comment #40811332
The terms of use remain the same as Gemma 1 - https://ai.google.dev/gemma/terms .
-
comment
Comment #40811310
Your training input has the shape of (sequence length x batch size). If a lot of your samples are shorter than sequence length, as is usually the case, you will have a lot of paddi…
-
comment
Comment #40811067
Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.
- story
- story
-
comment
Comment #39490217
As a fellow Bulgarian from the 80s and 90s myself, and now a part of the Gemma team, I’d say Austin, Jan, and team very much live up to the ethos of hackers I'd meet on BBSes back …
-
comment
Comment #39461258
Sorry, doing our best here :)
-
comment
Comment #39461245
We have many exciting things planned that we can't reveal just yet :)
-
comment
Comment #39461219
September 2023.