Live data from Hacker News

My local model setup on an M4 Pro Mac Mini

lws.io

121–130 of 208 posts

Re: My local model setup on an M4 Pro Mac Mini

#121
post #25

Earlier quoted context omitted.

I can't imagine using CPU... Oh I did twice. If you are work from home and do dishes between prompts you can get a gpt3-like result. I found it useful when I was... Well I didn't find it useful. But an Nvidia 3060 let me ask unethical questions pretty fast.

With which model. I have a 3060 with a bunch of system ram

Old school Berkeley Sterling or an abliterared model.

Re: My local model setup on an M4 Pro Mac Mini

#122
post #2

No mention of the performance of the models? I'm able to load a bunch of different models on my little mini-PC with 16GB RAM, but the performance is terrible. I always wonder what performance people are getting with local models that they find is acceptable?

I can't imagine using CPU... Oh I did twice. If you are work from home and do dishes between prompts you can get a gpt3-like result. I found it useful when I was... Well I didn't find it useful. But an Nvidia 3060 let me ask unethical questions pretty fast.

That is a lot of dishes.

Re: My local model setup on an M4 Pro Mac Mini

#123
post #106

Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…

It's free like ads are free. Or certain kinds of advice.

Re: My local model setup on an M4 Pro Mac Mini

#124

Earlier quoted context omitted.

Here’s an experiment: purchase an anthropic pro max subscription for $200/m. Now go buy the hardware to run DeepSeek’s equivalent. In a year, who spent more?

It’s not so clear after 5 years that you’ll come out ahead. You’ll have spent $20k. The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware. Idk where you live, but where I am running the M5 Ultra Mac Studio at max rated power 24/7 for a month costs C$42. The considerations against Apple hardware are 1) hardware advancements 2) early access to the best…

> The apple computer owner will probably be running local models that are better than today’s frontier on the same hardware.

Hardware is not magically getting more memory or bandwidth.

Believing there will be some magical optimizations to compensate for it is just dellusion.

Re: My local model setup on an M4 Pro Mac Mini

#125
post #106

Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…

At some point and for some tasks, predictability is important if not critical.

I’d rather use a tool where I know the limitations, over a tool where the limitations and strengths keep changing.

This way I know where in the process I ought to step in and pay attention.

Re: My local model setup on an M4 Pro Mac Mini

#126
post #2

No mention of the performance of the models? I'm able to load a bunch of different models on my little mini-PC with 16GB RAM, but the performance is terrible. I always wonder what performance people are getting with local models that they find is acceptable?

i have a 512gb ram m3 ultra mac studio setup with a gas city that runs one of my companies. today was the first time ever that a local model (GLM5.3 8-bit) was able to match fable5 in our tests. GLM-5.3-Flash at true 8-bit: 341 GB on disk, 328 GB resident, 288 experts across 46 layers, loads in 65 seconds. • 18.7 tokens/s generation, 35 tokens/s prompt, on a desk, on a $0 per-token bill. • Runs beside our whole agent…

    $0 per-token bill
You still have electricity and capital investment. Envelope math suggests cheap electricity is costing you something like $0.50/mtok and the opportunity cost on the capital tied up and lost in the unit purchase and resale is going to cost you something like $2/mtok at 100% utilization (so, frontier model prices or higher at real utilization), and you don't benefit from any elasticity.

Hosted GLM 5.3 flash is like $0.15/mtok in $0.50/mtok out

Re: My local model setup on an M4 Pro Mac Mini

#129
post #127

Is there anything one can reasonably run on a mac mini M2 with just 24GB RAM or should I not even try?

It depends on your use case, but the smaller Gemma 4 models or qwen3.6:9b would probably run OK on that. I recommend trying it, even just for fun. It‘s easy with omlx.
Post reply on HN