Ollama is now powered by MLX on Apple Silicon in preview
1–10 of 384 posts
Re: Ollama is now powered by MLX on Apple Silicon in preview
#2Re: Ollama is now powered by MLX on Apple Silicon in preview
#3Re: Ollama is now powered by MLX on Apple Silicon in preview
#4Re: Ollama is now powered by MLX on Apple Silicon in preview
#5LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.
On device I would gladly pay for good hardware - it's my machine and I'm using as I see fit like an IDE.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#6still waiting for the day I can comfortably run Claude Code with local llm's on MacOS with only 16gb of ram
Re: Ollama is now powered by MLX on Apple Silicon in preview
#7Re: Ollama is now powered by MLX on Apple Silicon in preview
#8#The use of NVFP4 results in a 3.5x reduction in model memory footprint relative to FP16 and a 1.8x reduction compared to FP8, while maintaining model accuracy with less than 1% degradation on key language modeling tasks for some models.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#9LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.
I'm not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn't, what are we paying hundreds of billions of dollars for?
I think local LLMs will continue to grow and there will be an "ChatGPT" moment for it when good enough models meet good enough hardware. We're not there yet though.
Note, this is why I'm big on investing in chip manufacture companies. Not only are they completely maxed out due to cloud LLMs, but soon, they will be double maxed out having to replace local computer chips with ones that are suited for inferencing AI. This is a massive transition and will fuel another chip manufacturing boom.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#10LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.