Earlier quoted context omitted.
unfortunately too big for the broader community to test. Will be very interesting to see how well it performs compared to the large models
Not really, looks like a ~40B class model which is very runnable.
So 2B for attention + 5Bx2 for inference = 12B in RAM at runtime.