Earlier quoted context omitted.
I see all these LLM posts about if a certain model can run locally on certain hardware and I don’t get it. What are you doing with these local models that run at x tokens/sec. Do you have the equivalent of ChatGPT running entirely locally? What do you do with it? Why? I honestly don’t understand the point or use case.
most of the llm tooling can handle different models. Ollama makes it easy to install and run different models locally. So you can configure aider or vscode or whatever you're using to connect to chatgpt to point to your local models instead. None of them are as good as the big hosted models, but you might be surprised at how capable they are. I like running things locally when I can, and I also like not worrying abou…
Ollama is now powered by MLX on Apple Silicon in preview
311–320 of 384 posts
Re: Ollama is now powered by MLX on Apple Silicon in preview
#312On-device models are the future. Users prefer them. No privacy issues. No dealing with connectivity, tokens, or changes to vendors implementations. I have an app using Foundation Model, and it works great. I only wish I could backport it to pre macOS 26 versions.
Actual consumers not only don't care, they will not even be aware of the difference.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#313On-device models are the future. Users prefer them. No privacy issues. No dealing with connectivity, tokens, or changes to vendors implementations. I have an app using Foundation Model, and it works great. I only wish I could backport it to pre macOS 26 versions.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#314On-device models are the future. Users prefer them. No privacy issues. No dealing with connectivity, tokens, or changes to vendors implementations. I have an app using Foundation Model, and it works great. I only wish I could backport it to pre macOS 26 versions.
Obviously hardware wise the real blocker is memory cost. But there is no reason why future devices couldn't bundle 256GB of mem by default.
Cost is a pretty big reason.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#315Earlier quoted context omitted.
70% of the world’s population use at least one Meta property at least once per day. How many of the other 30% are too poor/young/computer illiterate to be part of an addressable market? Every company has dozens of SaaS products that store their business critical information. Amazon installs Office on each computer, Slack (they were moving away from Chime when I left), and the sales department uses SalesForce - SA’s a…
These are all great statistics, but how do you explain ClawdBot explosion. Even in lower income countries like China. So much demand that Apple can’t keep up production of Mac Minis. Why aren’t these folks going towards cloud solutions? Is it cost or is there some consideration for having more control over their data?
Re: Ollama is now powered by MLX on Apple Silicon in preview
#316LLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.
I disagree with every sentence of this. > solves the problem of too much demand for inference False, it creates consumer demand for inference chips, which will be badly utilised. > also would use less electricity What makes you think that? (MAYBE you can save power on cooling. But not if the data center is close to a natural heat sink) > It's just a matter of getting the performance good enough. The performance limit…
There are so many CPUs, GPUs, RAM and SSDs which are underutilized. I have some in my closet doing 5% load at peek times. Why would inference chips be special once they become commodity hardware?
Re: Ollama is now powered by MLX on Apple Silicon in preview
#317Earlier quoted context omitted.
I highly doubt they can catch up in 3-5 years to Nvidia. Chips take about 3 years to design. Do you think China will have Feymann-level AI systems in 3 years? I think in 3 years, they'll have H200-equivalent at home.
You must have an inside line on information for 'China' -- those are bold predictions!
Re: Ollama is now powered by MLX on Apple Silicon in preview
#318Earlier quoted context omitted.
Right. When I said "they'll always be behind", I meant in the next 5-10 years. They're gated by EUV tech. And once they have EUV tech, they need to scale up chip manufacturing.
You will always be wrong.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#319Earlier quoted context omitted.
Users don’t care about “privacy”. If they did, Meta and Alphabet wouldn’t be worth $1T+. Users really don’t matter at all. The revenue for AI companies will be B2B where the user is not the customer - including coding agents. Most people don’t even use computers as their primary “computing device” and most people are buying crappy low end Android phones - no I’m not saying all Android phones are crappy. But that’s wh…
Have you done A/B tests to see if consumers prefer Facebook with or without privacy? No? What? Oh, you can't? Neither can consumers. Most consumers are very aware of the lack of privacy, the manipulation, and have very cynical feelings about Facebook and similar companies. But it's where their friends and family are. For most people the web is a mine field maze where basic things they want are compromised everywhere.…
They are choosing to give Facebook info.
Re: Ollama is now powered by MLX on Apple Silicon in preview
#320Earlier quoted context omitted.
I disagree with every sentence of this. > solves the problem of too much demand for inference False, it creates consumer demand for inference chips, which will be badly utilised. > also would use less electricity What makes you think that? (MAYBE you can save power on cooling. But not if the data center is close to a natural heat sink) > It's just a matter of getting the performance good enough. The performance limit…
> False, it creates consumer demand for inference chips, which will be badly utilised. There are so many CPUs, GPUs, RAM and SSDs which are underutilized. I have some in my closet doing 5% load at peek times. Why would inference chips be special once they become commodity hardware?