Live data from Hacker News

Bonsai 27B: A 27B-Class model that runs on a phone

prismml.com

1–10 of 278 posts

Re: Bonsai 27B: A 27B-Class model that runs on a phone

#6
post #3

The problem, of course, is if you run the UD_Q2 variant (Unsloth) which does only post-training, the number is pretty close to 1-bit model here and the 5% drop in tool-call is significant than it suggests in real-life use cases.

You also need to pay close attention to BFCLv3 multi-turn result, that helps you to get a sense how frequently these quants will be in a doom loop.

Re: Bonsai 27B: A 27B-Class model that runs on a phone

#8
The models themselves are showing up on Hugging Face here: https://huggingface.co/prism-ml/models

I've tried a couple in LM Studio - the GGUF one and the MLX one - but neither worked there. Anyone else get them to work? Might be that LM Studio needs to upgrade their llama.cpp or MLX engines first.

Re: Bonsai 27B: A 27B-Class model that runs on a phone

#9
post #2

TIL that 1 bit models are actually 1.58 bit with three values +1, 0 and -1

There's two variants of this (or, as the joke goes, for very big values of bit):

Ternary Bonsai 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight.

1-bit Bonsai 27B uses binary {−1, +1} weights with the same group-wise scaling, giving 1.125 effective bits per weight.

Post reply on HN