Earlier quoted context omitted.
I got it to build and run the example app on my M3 max with 36 gb ram. Memory pressure was around 32 gb
Did you quantise it? At what level and what was your impression compared to other recent smaller models at that quantisation, if so?
Instructions here: https://github.com/facebookresearch/llama/pull/947/