Viewing profile — mluo
mluo
HN member- Joined
- Tue, Jan 24, 2023, 10:15 PM UTC
- HN karma
- 51
- Public activity
- 12 items
- HN profile
- View on Hacker News ↗
About mluo
No profile information was provided.
Recent public activity
-
comment
Comment #43019252
Check out one of my prior work: https://stylus-diffusion.github.io/ This work scales up selection/routing over many models/LoRAs
-
comment
Comment #43019094
For quantization, very big impact for small models, can drop at much as 10% on AIME. Our model does best on bfloat16 ;) Come checkout our repo at: https://github.com/agentica-proje…
-
comment
Comment #43019082
It's simply bc the model is small (1.5B), making it sensitive to weight perturbations
-
comment
Comment #43019074
Think there are some people who made GGUFs as branches of our model, try it out! https://huggingface.co/models?other=base_model:quantized:age...
-
comment
Comment #43019001
Nice, very glad to see it works! Small models are very sensitive to the dtype :(
-
comment
Comment #43018627
Try bfloat16! We have a bug where the model was saved as fp32.
-
comment
Comment #43018621
We beat O1-preview and even many other 7B models over many math benchmarks, which was TEST set (not in training set at all). If you want to make the model fully generalist, feel fr…
-
comment
Comment #43018592
One of the authors here.... This is not a Chinese model, btw I'm American
-
comment
Comment #43018586
Hi, one of the lead authors for this work. We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!
- story
-
comment
Comment #36078345
Alpaca Llama Vicuna -> Gorilla Chad move
-
comment
Comment #34511349
With inflation in mind, wouldn't there a larger gap between Ray's sort and the previous WR holder?