[dupe]
Some more discussion a few weeks ago: https://news.ycombinator.com/item?id=40620955
11–15 of 15 posts
Some more discussion a few weeks ago: https://news.ycombinator.com/item?id=40620955
The relevant paper: https://arxiv.org/abs/2406.02528 In summary, they forced the model to process data in ternary system and then build a custom FPGA chip to process the data more efficiently. Tested to be "comparable" to small models (3B), theoretically scale to 70B, unknown for SOTAs (>100B params). We have always known custom chips are more efficient especially for tasks like these where it is basically approximat…