Live data from Hacker News

Viewing profile — alekandreev

alekandreev

HN member
Joined
Tue, Feb 13, 2024, 9:08 PM UTC
HN karma
368
Public activity
31 items

About alekandreev

No profile information was provided.

Recent public activity

  1. comment
    Comment #43343940

    Yes we have measured the tradeoff. We don't see a drop of perplexity in English when introducing multilingual, and there is a slight drop in some English language-specific evals (~…

  2. comment
    Comment #43343934

    We never train at 128k, only 32k, changing the scaling factor at the end. We wanted the long context recipe to be friendly for finetuning, and training at 128k is a bit of a pain w…

  3. comment
    Comment #43343506

    We recommend using user for the system prompt as well.

  4. comment
    Comment #43342768

    That's a very interesting area, but nothing we can announce today.

  5. comment
    Comment #43341721

    Hey, Gemma engineer here. Can you please share reports on the type of prompts and the implementation you used?

  6. comment
    Comment #43341711

    Thank you for the report! We are working with the Ollama team directly and will look into it.

  7. comment
    Comment #43341422

    Thank you for the feedback! This is why we are so excited to push more and more on small models for both low end and high end smartphones!

  8. comment
    Comment #43341298

    That's an idea we've thought about. However, we think the open source community has already created a very impressive set of language or region-specific finetunes [1] [2]. Also the…

  9. comment
  10. comment
    Comment #43341030

    Picking model sizes is not an exact science. We look for sizes that will fit quantized on different categories on devices (e.g., low-end and high-end smartphone, laptops and 16GB G…

  11. comment
    Comment #43340828

    Greetings from the Gemma team! We just got Gemma 3 out of the oven and are super excited to show it to you! Please drop any questions here and we'll answer ASAP. (Opinions our own …

  12. comment
    Comment #40813599

    I think it makes sense to compare models trained with the same recipe on token count - usually more tokens will give you a better model. However, I wouldn't draw conclusions about …

  13. comment
    Comment #40813091

    To quote Ludovic Peran, our amazing safety lead: Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of f…

  14. comment
    Comment #40812188

    Happy to pass on any feedback to our Google Cloud friends. :)

  15. comment
    Comment #40812177

    This is mostly about inference speed, while maintaining long context performance.

  16. comment
    Comment #40812158

    In addition to the HF links shared by sibling comments, the 2B will be released soon.

  17. comment
    Comment #40811332

    The terms of use remain the same as Gemma 1 - https://ai.google.dev/gemma/terms .

  18. comment
    Comment #40811310

    Your training input has the shape of (sequence length x batch size). If a lot of your samples are shorter than sequence length, as is usually the case, you will have a lot of paddi…

  19. comment
    Comment #40811067

    Hello (again) from the Gemma team! We are quite excited to push this release out and happy to answer any questions! Opinions are our own and not of Google DeepMind.

  20. story
  21. story
  22. comment
    Comment #39490217

    As a fellow Bulgarian from the 80s and 90s myself, and now a part of the Gemma team, I’d say Austin, Jan, and team very much live up to the ethos of hackers I'd meet on BBSes back …

  23. comment
    Comment #39461258

    Sorry, doing our best here :)

  24. comment
    Comment #39461245

    We have many exciting things planned that we can't reveal just yet :)

  25. comment
    Comment #39461219

    September 2023.