Live data from Hacker News

Ollama is now powered by MLX on Apple Silicon in preview

ollama.com

31–40 of 384 posts

Re: Ollama is now powered by MLX on Apple Silicon in preview

#31

Earlier quoted context omitted.

You could argue that the only reason we have good open-weight models is because companies are trying to undermine the big dogs, and they are spending millions to make sure they dont get too far ahead. If the bubble pops then there wont be incentive to keep doing it.

I agree. I can totally see in the future that open source LLMs will turn into paying a lumpsum for the model. Many will shut down. Some will turn into closed source labs. When VCs inevitably ask their AI labs to start making money or shut down, those free open source LLMS will cease to be free. Chinese AI labs have to release free open source models because they distill from OpenAI and Anthropic. They will always be…

“They will always be behind”

Car manufacturers said the same.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#32

LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.

[deleted]

Re: Ollama is now powered by MLX on Apple Silicon in preview

#33

Earlier quoted context omitted.

I agree. I can totally see in the future that open source LLMs will turn into paying a lumpsum for the model. Many will shut down. Some will turn into closed source labs. When VCs inevitably ask their AI labs to start making money or shut down, those free open source LLMS will cease to be free. Chinese AI labs have to release free open source models because they distill from OpenAI and Anthropic. They will always be…

“They will always be behind” Car manufacturers said the same.

It did take decades to catch and surpass US car makers right?

Re: Ollama is now powered by MLX on Apple Silicon in preview

#34
post #10

LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.

"Most users don't need frontier model performance" unfortunately, this is not the case.

Any citations? Because that was my impression, too. I want frontier model performance for my coding assistant, but "most users" could do with smaller/faster models.

ChatGPT free falls back to GPT-5.2 Mini after a few interactions.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#35

Earlier quoted context omitted.

You could argue that the only reason we have good open-weight models is because companies are trying to undermine the big dogs, and they are spending millions to make sure they dont get too far ahead. If the bubble pops then there wont be incentive to keep doing it.

I agree. I can totally see in the future that open source LLMs will turn into paying a lumpsum for the model. Many will shut down. Some will turn into closed source labs. When VCs inevitably ask their AI labs to start making money or shut down, those free open source LLMS will cease to be free. Chinese AI labs have to release free open source models because they distill from OpenAI and Anthropic. They will always be…

> have to release free open source models because they distill from OpenAI and Anthropic

They dont really have to though, they just need to be good enough and cheaper (even if distilled). That being said, it is true they are gaining a lot of visibility (specially Qwen) because of being open-source(weight).

Hardware-wise they seem they will catch-up in 3-5 years (Nvidia is kind of irrelevant, what matters is the node).

Re: Ollama is now powered by MLX on Apple Silicon in preview

#36
post #15

Earlier quoted context omitted.

It isn't going to replace cloud LLMs since cloud LLMs will always be faster in throughput and smarter. Cloud and local LLMs will grow together, not replace each other. I'm not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn't, what are we paying hundreds of billions of dollars for? I think lo…

Looking at downvotes I feel good about SDE future in 3-5 years. We will have a swamp of "vibe-experts" who won't be able to pay 100K a month to CC. Meanwhile, people who still remember how to code in Vim will (slowly) get back to pre-COVID TC levels.

[deleted]

Re: Ollama is now powered by MLX on Apple Silicon in preview

#37

LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.

It isn't going to replace cloud LLMs since cloud LLMs will always be faster in throughput and smarter. Cloud and local LLMs will grow together, not replace each other. I'm not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn't, what are we paying hundreds of billions of dollars for? I think lo…

Yep. People were claiming DeepSeek was "almost as good as SOTA" when it came out. Local will always be one step away like fusion.

It's just wishful thinking (and hatred towards American megacorps). Old as the hills. Understandable, but not based on reality.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#38
post #19

Earlier quoted context omitted.

It isn't going to replace cloud LLMs since cloud LLMs will always be faster in throughput and smarter. Cloud and local LLMs will grow together, not replace each other. I'm not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn't, what are we paying hundreds of billions of dollars for? I think lo…

We are 100% there already. In browser. the webgpu model in my browser on my m4 pro macbook was as good as chatgpt 3.5 and doing 80+ tokens/s Local is here.

Sir, ChatGPT 3.5 is more than 3 years old, running on your bleeding edge M4 Pro hardware, and only proves the previous commenters point.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#40
post #17

Earlier quoted context omitted.

[flagged]

Complaining about downvotes is futile and is also against hn guidelines.

I'm not complaining "about downvotes" LOL I'm explaining why some people will be replaced by LLMs because of their own "context window" length.
Post reply on HN