Nice, I think this is the second time I see this here on HN, I always wondered why we need to shove the entire model into memory, I don't care who King Charles is every single time. It always felt as though we already figured out how to break up large files and parse them efficiently with very little memory. Frontier AI feels like its full of people who are brilliant at making models, but when it comes to scale and p…
I don't care who King Charles is every single time that's the trick and a multi-billion dollar question, how would an llm engine know that? it's an active research area how to cull the initial layer surface and do the optimal traversal path through the layers and it's a damn hard problem. It's definitely an area where a ton of performance is left on the table still.
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
341–350 of 382 posts
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#342Earlier quoted context omitted.
I don't care who King Charles is every single time that's the trick and a multi-billion dollar question, how would an llm engine know that? it's an active research area how to cull the initial layer surface and do the optimal traversal path through the layers and it's a damn hard problem. It's definitely an area where a ton of performance is left on the table still.
I have some ideas on how that specifically can be solved, but I've taken a stance to never give OpenAI or Anthropic any of my ideas for free. I am the most surprised that Google seems to be trailing behind them. I'm not sure if they're even taking this seriously anymore. I do appreciate the open models they do release on the other hand, I hope they never stop. I wish Microsoft would do more with Phi and similar.
Write them in the margin of a book....
"It is impossible for a cube to be a sum of two cubes, a fourth power to be a sum of two fourth powers, or in general for any number that is a power greater than the second to be the sum of two like powers. I have discovered a truly remarkable proof, but this margin is too small to contain it."
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#343Earlier quoted context omitted.
I don't care who King Charles is every single time that's the trick and a multi-billion dollar question, how would an llm engine know that? it's an active research area how to cull the initial layer surface and do the optimal traversal path through the layers and it's a damn hard problem. It's definitely an area where a ton of performance is left on the table still.
I have some ideas on how that specifically can be solved, but I've taken a stance to never give OpenAI or Anthropic any of my ideas for free. I am the most surprised that Google seems to be trailing behind them. I'm not sure if they're even taking this seriously anymore. I do appreciate the open models they do release on the other hand, I hope they never stop. I wish Microsoft would do more with Phi and similar.
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#344Earlier quoted context omitted.
I don't care who King Charles is every single time that's the trick and a multi-billion dollar question, how would an llm engine know that? it's an active research area how to cull the initial layer surface and do the optimal traversal path through the layers and it's a damn hard problem. It's definitely an area where a ton of performance is left on the table still.
The human analogy would be that I don’t remember everything in the books I have read, but I do recall reading a particular book and can always look it up. So, is there a way to train a neural network and then tune it forget a lot of the facts that can be easily retrieved, but keep the intelligence.
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#345Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#346> The measured result is a reference point, not a performance ceiling. Claude was here.
Okay, I’m the first to get annoyed at LLM-ese. But language requires the ability to say that something is A and not B. There are certain phrasings of that that are painfully Claude-esque, but ffs the one you quoted is the type of thing an actual human being is just as likely to say.
As a English-as-a-second-language speaker I’ve been using that formation long before llm and it has legitimate usage
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#347Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#348Earlier quoted context omitted.
Okay, I’m the first to get annoyed at LLM-ese. But language requires the ability to say that something is A and not B. There are certain phrasings of that that are painfully Claude-esque, but ffs the one you quoted is the type of thing an actual human being is just as likely to say.
I don’t get why that construction is associated with Claude. I’ve never picked it up from reading Claude output. As a English-as-a-second-language speaker I’ve been using that formation long before llm and it has legitimate usage
Re: Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
#349Earlier quoted context omitted.
New macOS is bloatware that makes your computer slower
So install Asahi Linux?
macOS is the only OS which fully supports M1 hardware and its security features. Please see Asahi Linux's documentation: https://asahilinux.org/docs/platform/feature-support/m1/#m1-...