I'm currently experimenting with running google/gemma-4-26b-a4b with lm studio ( https://lmstudio.ai/ ) and Opencode on a M3 Ultra with 48Gb RAM. And it seems to be working. I had to increase the context size to 65536 so the prompts from Opencode would work, but no other problems so far. I tried running the same on an M3 Max with less memory, but couldn't increase the context size enough to be useful with Opencode. I…
I have a similar setup. It might be worth checking out pi-coding-agent [0]. The system prompt and tools have very little overhead ( [0] https://www.npmjs.com/package/@mariozechner/pi-coding-agent#...
Re: I ran Gemma 4 as a local model in Codex CLI
#121Pi is _really_ good for personal stuff, but since it lacks every single safety imaginable, it's not realy something one can deploy in a corporate environment :D