Model card for base: https://huggingface.co/databricks/dbrx-base > The model requires ~264GB of RAM I'm wondering when everyone will transition from tracking parameter count vs evaluation metric to (total gpu RAM + total CPU RAM) vs evaluation metric. For example, a 7B parameter model using float32s will almost certainly outperform a 7B model using float4s. Additionally, all the examples of quantizing recently releas…
I'm more wondering when we'll have algorithms that will "do their best" given the resources they detect. That would be what I call artificial intelligence. Giving up because "out of memory" is not intelligence.
When people can't remember the facts/theory/formulas needed to answer some test question, or can't memorize some complicated information because it's too much, they usually give up too.
So, giving up because of "out of memory" sure sounds like intelligence to me.