Earlier quoted context omitted.
>Time to first token measured with an 8K-token prompt using a 14-billion parameter model with 4-bit quantization Oh dear 14B and 4-bit quant? There are going to be a lot of embarrassed programmers who need to explain to their engineering managers why their Macbook can't reasonably run LLMs like they said it could. (This already happened at my fortune 20 company lol)
Yeah no it didn’t. If you have a fully speced out M3/4 MacBook with enough memory you’re running pretty decent models locally already. But no one is using local models anyway.
What is "it" and what didn't it do?