Earlier quoted context omitted.
his conclusion is simultaneously not warranted and correct a like-for-like comparison would be GPT-4 against the larger models like LLaMA 65B, but those cannot be run on consumer-grade hardware so one ends up comparing the stuff one can run... against the top stuff from OpenAI running on high-end GPU farms, and this technology clearly benefits a lot still from much larger scale than most people can afford the great r…
The 65B model runs fine on a Mac Studio with 64GB of memory. The output is unremarkable; it’s not significantly better than the 13B model for most uses. GPT 3.5 is an order of magnitude better at least .
To run it properly you need a lot more than a Mac Studio, and then comparisons need to be done more or less seriously, not just a few random prompts, because anything in a black box will "cheat" and will be fine tuned to do well at popular benchmarks.