Live data from Hacker News

Viewing profile — mluo

mluo

HN member
Joined
Tue, Jan 24, 2023, 10:15 PM UTC
HN karma
51
Public activity
12 items

About mluo

No profile information was provided.

Recent public activity

  1. comment
    Comment #43019252

    Check out one of my prior work: https://stylus-diffusion.github.io/ This work scales up selection/routing over many models/LoRAs

  2. comment
    Comment #43019094

    For quantization, very big impact for small models, can drop at much as 10% on AIME. Our model does best on bfloat16 ;) Come checkout our repo at: https://github.com/agentica-proje…

  3. comment
    Comment #43019082

    It's simply bc the model is small (1.5B), making it sensitive to weight perturbations

  4. comment
    Comment #43019074

    Think there are some people who made GGUFs as branches of our model, try it out! https://huggingface.co/models?other=base_model:quantized:age...

  5. comment
    Comment #43019001

    Nice, very glad to see it works! Small models are very sensitive to the dtype :(

  6. comment
    Comment #43018627

    Try bfloat16! We have a bug where the model was saved as fp32.

  7. comment
    Comment #43018621

    We beat O1-preview and even many other 7B models over many math benchmarks, which was TEST set (not in training set at all). If you want to make the model fully generalist, feel fr…

  8. comment
    Comment #43018592

    One of the authors here.... This is not a Chinese model, btw I'm American

  9. comment
    Comment #43018586

    Hi, one of the lead authors for this work. We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!

  10. story
  11. comment
    Comment #36078345

    Alpaca Llama Vicuna -> Gorilla Chad move

  12. comment
    Comment #34511349

    With inflation in mind, wouldn't there a larger gap between Ray's sort and the previous WR holder?