Live data from Hacker News

Ollama is now powered by MLX on Apple Silicon in preview

ollama.com

341–350 of 384 posts

Re: Ollama is now powered by MLX on Apple Silicon in preview

#341
post #231
post #50

Earlier quoted context omitted.

Depending on the use case, the future is already here. For example, last week I built a real-time voice AI running locally on iPhone 15. One use case is for people learning speaking english. The STT is quite good and the small LLM is enough for basic conversation. https://github.com/fikrikarim/volocal

That’s awesome! I’ve got a similar project for macOS/ iOS using the Apple Intelligence models and on-device STT Transcriber APIs. Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets? Maybe we’re not there yet, but I’m interested in a better, local Siri like this with some sort of “agentic lite” capabilities.

> Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets?

I first tried the Qwen 3.5 0.8B Q4_K_S and the model couldn't hold a basic conversation. Although I haven't tried lower quants on 2B.

I'm also interested on the Apple Foundation models, and it's something I plan to try next. AFAIK it's on par with Qwen-3-4B [0]. The biggest upside as you alluded to is that you don't need to download it, which is huge for user onboarding.

[0] https://machinelearning.apple.com/research/apple-foundation-...

Re: Ollama is now powered by MLX on Apple Silicon in preview

#342
post #86

I created "apfel" https://github.com/Arthur-Ficial/apfel a CLI for the apple on-device local foundation model (Apple intelligence) yeah its super limited with its 4k context window and super common false positives guardrails (just ask it to describe a color) ... bit still ... using it in bash scripts that just work without calling home / out or incurring extra costs feels super powerful.

this is great! Incredibly fast and is working pretty well running loads on my m4 max studio.

Weirdly though I'm getting things like this: Apple FM is fast and free but has a hard limitation — it can't process prompts with Spanish/non-English words, which is a dealbreaker for California and Southwest real estate where half the street names are Spanish.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#343

Earlier quoted context omitted.

Local RTX 5090 is actually faster than A100/H100.

It's a $4,000 GPU with 32GB of VRAM and needs a 1,000 watt PSU. It's not realistic for the masses. If it has something like 80GB of VRAM, it'll cost $10k. The actual local LLM chip is Apple Silicon starting at the M5 generation with matmul acceleration in the GPU. You can run a good model using an M5 Max 128GB system. Good prompt processing and token generation speeds. Good enough for many things. Apple accidentally…

Yes, it’s expensive hobby.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#344

Earlier quoted context omitted.

Local RTX 5090 is actually faster than A100/H100.

Crazy thing to say without other contextual information - it obviously depends on a number of factors. Do you have an apples to apples comparison at hand?

Look it up.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#345
post #271

Earlier quoted context omitted.

Oh I'm sure they'll continue to have some cloud services, no doubt. But look at VMware for example, even after the insane price increases. Nutanix also seems to be doing quite well. I'm seeing a fair amount of on-prem bare metal k8s too.

Again - anecdotes is not data. We have data. That would be about as silly as me citing my own experience as proof that “everyone is moving to AWS” when I work for a company that is exclusively an AWS partner consulting company.

> Again - anecdotes is not data. We have data.

You have data showing growth in cloud, which I expect and don't disagree with. The data I come across shows this too!

What I disagree with, from my own experiences and all the data I can seem to find online is that the growth rate in repatriation is MUCH higher than the growth in cloud.

It has flipped over the last 3yr.

US Enterprises, Fortune 100, especially. Also a lot of public entities (gov).

"In 2025, repatriation is still generally an upward trend. Data from the end of 2024 showed that 86% of CIOs planned to move some public cloud workloads back to private cloud or on-premises — the highest on record for the Barclays CIO Survey."

"Real examples of cloud repatriation include Dropbox, Adobe, and GEICO. All three companies moved a significant portion of their infrastructure onto public cloud before moving it to a combination of on-premises and hybrid cloud providers."

Noted: SaaS accounts for 46.10% of market revenue, while PaaS is the fastest-growing segment at 21.35% CAGR

Re: Ollama is now powered by MLX on Apple Silicon in preview

#346

Earlier quoted context omitted.

I see all these LLM posts about if a certain model can run locally on certain hardware and I don’t get it. What are you doing with these local models that run at x tokens/sec. Do you have the equivalent of ChatGPT running entirely locally? What do you do with it? Why? I honestly don’t understand the point or use case.

most of the llm tooling can handle different models. Ollama makes it easy to install and run different models locally. So you can configure aider or vscode or whatever you're using to connect to chatgpt to point to your local models instead. None of them are as good as the big hosted models, but you might be surprised at how capable they are. I like running things locally when I can, and I also like not worrying abou…

Thanks.

I still don’t understand. What are you using this long you’re running locally to actually do?

What is the use case?

Re: Ollama is now powered by MLX on Apple Silicon in preview

#347

Earlier quoted context omitted.

I see all these LLM posts about if a certain model can run locally on certain hardware and I don’t get it. What are you doing with these local models that run at x tokens/sec. Do you have the equivalent of ChatGPT running entirely locally? What do you do with it? Why? I honestly don’t understand the point or use case.

1. There are small local models that have the capabilities of frontier models a year ago 2. They aren't harvesting your data for government files or training purposes 3. They won't be altered overnight to push advertising or a political agenda 4. They won't have their pricing raised at will 5. They won't disappear as soon as their host wants you to switch

Good points. What local models have you found work best for your use cases? I feel like if we get to opus 4.6 level intelligence running on local hardware, we’re in the clear for a lot of day to day use cases.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#348
post #345

Earlier quoted context omitted.

Again - anecdotes is not data. We have data. That would be about as silly as me citing my own experience as proof that “everyone is moving to AWS” when I work for a company that is exclusively an AWS partner consulting company.

> Again - anecdotes is not data. We have data. You have data showing growth in cloud, which I expect and don't disagree with. The data I come across shows this too! What I disagree with, from my own experiences and all the data I can seem to find online is that the growth rate in repatriation is MUCH higher than the growth in cloud. It has flipped over the last 3yr. US Enterprises, Fortune 100, especially. Also a lot…

Again, anecdotes. I have public company quarterly statements - you have unsourced quotes. You can quote Geico - I can quote Netflix. If on prem was really growing, I wouldn’t expect Intel to be in the shitter and I would expect Capex to be focused on Colo centers not cloud.

Also when I searched for your quotation the very next paragraph was

“ This trend does not represent a rejection of cloud computing. Organizations continue investing heavily in cloud services, with Gartner forecasting that global cloud spending will reach approximately $723 billion by the end of 2025.”

Re: Ollama is now powered by MLX on Apple Silicon in preview

#349
post #341
post #231

Earlier quoted context omitted.

That’s awesome! I’ve got a similar project for macOS/ iOS using the Apple Intelligence models and on-device STT Transcriber APIs. Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets? Maybe we’re not there yet, but I’m interested in a better, local Siri like this with some sort of “agentic lite” capabilities.

> Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets? I first tried the Qwen 3.5 0.8B Q4_K_S and the model couldn't hold a basic conversation. Although I haven't tried lower quants on 2B. I'm also interested on the Apple Foundation models, and it's something I plan to try next. AFAIK it's on par with Qwen-3-4B [0]. The biggest upside as y…

Try it with mxfp8 or bf16. It's a decent model for doing tool calling, but I wouldn't recommend using it with 4 bit quantization.

Re: Ollama is now powered by MLX on Apple Silicon in preview

#350
post #165

I have an M4 Max with 48GB RAM. Anyone have any tips for good local models? Context length? Using the model recommended in the blog post (qwen3.5:35b-a3b-coding-nvfp4) with Ollama 0.19.0 and it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking "Hello world". Is this the best that's currently achievable with my hardware or is there something that can be configured to get bet…

The 35b-a3b-coding-nvfp4 model has the recommended hyperparameters set for coding, not chatting. If you want to use it to chat you can pull the `35b-a3b-nvfp4` model (it doesn't need to re-download the weights again so it will pull quickly) which has the presence penalty turned on which will stop it from thinking so much. You can also try `/set nothink` in the CLI which will turn off thinking entirely.
Post reply on HN