Does it mean I can run whatever I want on ANE? Last time I tried it seemed it could only be used by first party features such as Face ID
Apple Core AI Framework
101–110 of 114 posts
Re: Apple Core AI Framework
#102Earlier quoted context omitted.
I have come at this at a slightly different angle. I am a fully-burned-out freelancer (in the last couple of years so severely and totally that I thought I had early onset dementia, and I am still not sure I don't). I don't really have an off-ramp to anything else yet, but the sea-change in the industry has been contributing to my feeling that I should knock it on the head. I must get past broad understanding of AI t…
How are you running that GGUF, and how many tokens/sec are you getting without MTP? My M1 Max gives me 65 t/s for non-MTP unsloth/gemma-4-26B-A4B-it-qat-GGUF (UD-Q4_K_XL), but with MTP that actually goes down to 56 t/s (at 63% accepted drafts).
https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gg...
Might see if Google has official drafters later.
Re: Apple Core AI Framework
#103This is why the AI companies are rushing to IPO. By the end of next year you’ll be running most of your AI on device. They have no moat, they’ve reached the limits of scaling, most of the magic can be distilled into smaller models, and they know it
Re: Apple Core AI Framework
#104Earlier quoted context omitted.
Requires OS 27+, so CoreML is still useful for backwards compatibility.
macOS users aren't that good at upgrading regularly, but iOS users are at least obsessive about upgrading to the latest OS. I guess the system almost forces us.
Re: Apple Core AI Framework
#105Earlier quoted context omitted.
I want to echo this. I've been on claude's opus 4.5/6/7 for work for a couple months, and I finally got back to running Qwen A3B 35B... it's incredibly performant and quite capable on semi-reasonable local hardware. I get ~150 tokens/s on dual nvidia RTX 3090s and can fit the whole 300k context into gpu on a UD-Q4-K-XL quant gguf. Combined with Pi as a harness, and I'm surprised to find that it feels about as capable…
Majority of my agentic setup is pi / Claude code where every single Chinese models are not as good except commercial 1T models . Local is a pipe dream . If you can run it cheap occasionally why commercial companies can’t run it cheaper 24/7 and lower the costs ? The answer is simple. Use cases are more demanding and hence you need more from model not less . Sure if you task is to do a narrow labeling task on 1m recor…
Re: Apple Core AI Framework
#106Earlier quoted context omitted.
I do have a 3090 Ti on my gaming PC, but even my old M1 MBP (with a mere 32gb of RAM) is quite competent and can run a quantized `Gemma4-26B-A4B` in the background while I do other stuff.
The MBP running Gemma4 is absolutely is useless for any real work.
Re: Apple Core AI Framework
#107Earlier quoted context omitted.
macOS users aren't that good at upgrading regularly, but iOS users are at least obsessive about upgrading to the latest OS. I guess the system almost forces us.
I still deploy CoreML features to iOS 15. Many devices in use can’t upgrade to 26/27
Re: Apple Core AI Framework
#108Earlier quoted context omitted.
The MBP running Gemma4 is absolutely is useless for any real work.
What is "real work"?
Re: Apple Core AI Framework
#109Earlier quoted context omitted.
This isn't an crazy statement (cpu performance metrics have mostly stalled their meteoric rise from prior to the 2000s) But it also doesn't capture the entire picture. CPU metrics mostly stalled for two reasons. 1. There wasn't much demand for the extra capacity. Even low end cpus from a decade ago are plenty capable for just browsing the web and typing up documents. It takes a novel use-case to drive demand again (o…
> There wasn't much demand for the extra capacity. Even low end cpus from a decade ago are plenty capable for just browsing the web and typing up documents. It stalled before the rise of PC-as-Internet-portal. I bought a high end PC in 2003, and 5 years later the PCs were not much faster - probably not even 2x. Around 2008-2010 was when most people started using PCs as a way to connect to the Internet. It stalled bec…
I was building gaming machines in the early 2000s, I absolutely remember the 4ghz wall that cpus hit.
But it wasn't a real wall... because we then got one of the arguably most influential processors ever in the Core 2 duo. Which... blew the limit away by giving you two processors clocked at 2.93 GHz each.
And honestly, even then - it was lack of demand (we could go to 4+ghz, but we didn't want to pay the power bill for the rest of the system - the planned pentium 5 was 7-10ghz on paper, but they canceled the project because keeping it fed and cool was too hard for personal desktop machines).
Of Note - we did reach these speeds on consumer hardware (ex - in 2012, Andre Yang hit 8.794Ghz on an AMD FX-8350)
So it was never "impossible" to keep scaling. It just wasn't worth it compared to going multi-core.
---
And maybe it's because I was in my formative years at this time, but you're off by 5+ years with this:
> Around 2008-2010 was when most people started using PCs as a way to connect to the Internet.
Gmail was a web only email client released in 2004. Wikipedia was released in 2001. Web browsing was very much one of the "killer" apps for computers by the 2000s. What do you think the damn 2000s dot-com bubble crash was?
at the risk of aging myself - I was born in '89, and I literally do not remember a time where we didn't have DSL speeds and above (friends houses often still had dial-up until ~2005, though).
Re: Apple Core AI Framework
#110Earlier quoted context omitted.
> There wasn't much demand for the extra capacity. Even low end cpus from a decade ago are plenty capable for just browsing the web and typing up documents. It stalled before the rise of PC-as-Internet-portal. I bought a high end PC in 2003, and 5 years later the PCs were not much faster - probably not even 2x. Around 2008-2010 was when most people started using PCs as a way to connect to the Internet. It stalled bec…
Yes, but it only stalled along a single dimension - Single core clock speed. I was building gaming machines in the early 2000s, I absolutely remember the 4ghz wall that cpus hit. But it wasn't a real wall... because we then got one of the arguably most influential processors ever in the Core 2 duo. Which... blew the limit away by giving you two processors clocked at 2.93 GHz each. And honestly, even then - it was lac…
Well, Gmail was actually one of the last web based email clients people used :-) Yahoo mail, Hotmail, and so many others predate Gmail by years.
> Web browsing was very much one of the "killer" apps for computers by the 2000s.
One of them. People still used non-browser apps for all kinds of things: Media consumption (people didn't watch movies on Youtube), Office (Google Docs was very much a niche thing for many years), photo-editing (lots of pirated versions of Photoshop/Lightroom years after the iPhone release), etc.
Most non-mail, non-social media, non-shopping stuff people do on the web these days was a dedicated SW from the vendor in those days. Want to make a photobook? Download this Windows binary and set it up there. It will then communicate with the server for the order (no browser utilized).
> at the risk of aging myself - I was born in '89, and I literally do not remember a time where we didn't have DSL speeds and above (friends houses often still had dial-up until ~2005, though).
Spring chicken! My first online experience was on a 340 baud modem :-)