Ask HN: Is it feasible to run a model on device for complete privacy?
1–8 of 8 posts
Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#2Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#3Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#4It's technically feasible, really just a question of whether this is worth $10,000(s) to you and you're willing to spend it.
Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#5Feasible but too expensive! I get that privacy is a priority for you but unfortunately if you want quality models you'd still have to maybe use frontier closed models..
Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#6Feasible but too expensive! I get that privacy is a priority for you but unfortunately if you want quality models you'd still have to maybe use frontier closed models..
No open source model that’s any good?
there are near-SOTA open models, but they are 1T+ parameters, i.e. they require over a terabyte of memory to run.
Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#7It's technically feasible, really just a question of whether this is worth $10,000(s) to you and you're willing to spend it.
Why financially crippling? It’s free to run on device. The native Apple Intelligence works well for smaller context windows and text only.
Re: Ask HN: Is it feasible to run a model on device for complete privacy?
#8Earlier quoted context omitted.
No open source model that’s any good?
the Gemma you tried is tiny, there are 31B and 26B (A4B) variants. there's also Qwen 3.6 with 27B and 35B (A3B) variants, reportedly pretty good. try them on open router or something. these require 30-40 Gb of memory to run between RAM and VRAM, less if quantized beyond near-lossless 8 bit. there are near-SOTA open models, but they are 1T+ parameters, i.e. they require over a terabyte of memory to run.