Ternary Bonsai: Top Intelligence at 1.58 Bits
1–10 of 60 posts
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#2Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#3I also have yet to see any of these at a larger scale. For example, can you try one of these at 100 billion parameters?
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#4Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#5Yet again they're comparing against unquantized versions of other models. They would probably still win but by a much smaller size margin.
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#6(I've been reading the MMLU-Redux questions for electrical engineering. They're very funny. Fifty years ago they might have been relevant. The references to the Intel 8085 date this to the mid-1970s. Moving coil meters were still a big thing back then. Ward-Leonard drives still drove some elevators and naval guns. This is supposed to be the hand-curated version of the questions. Where do they get this stuff? Old exams?)
[1] https://github.com/aryopg/mmlu-redux/blob/main/outputs/multi...
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#7So excited to see this - the big advantage of 1.58 bits is there are no multiplications at inference time, so you can run them on radically simpler and cheaper hardware.
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#8If you got that into a couple gigs--what could you stuff into 20 gigs?
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#9in my results, accuracy-wise Ternary-Bonsai-8B is on par with Qwen3.5-4B. But in accuracy-per-byte, bonsai is the clear winner:
=> Ternary-Bonsai-1.7B achieved 65.1% from 462 MiB, beating Qwen3.5-0.8B by 12 points while being ~5% smaller on disk. => Ternary-Bonsai-4B is the accuracy-per-byte winner above 1 GiB. 83.0% from only 1.1 GiB, within 2 points of Qwen3.5-4B at 40% of the weight size.
they show strong promise on edge devices and where disk space is limited. I think this lab is worth watching.
Re: Ternary Bonsai: Top Intelligence at 1.58 Bits
#10Why aren't they comparing to 2/3/4 bit quants?