> BitNet b1.58 can match the performance of the full precision baseline starting from a 3B size. ... This demonstrates that BitNet b1.58 is a Pareto improvement over the state-of-the-art LLM models. > BitNet b1.58 is enabling a new scaling law with respect to model performance and inference cost. As a reference, we can have the following equivalence between different model sizes in 1.58-bit and 16-bit based on the re…
They seem to be using LLAMA. Might be worth trying out. Their conversion formula seems stupidly simple.
Doesn't this mean that current big players can rapidly expand by huge multiples in size.?