Live data from Hacker News

Transformer architecture optimized for Apple Silicon

github.com

311–320 of 342 posts

Re: Transformer architecture optimized for Apple Silicon

#311

Earlier quoted context omitted.

UMA is the secret. A 256GB Mac Pro can dedicate almost all of that to AI. Their GPUs are the weak spot. The Nvidia 4080 is 2.5x faster than the M2, and the A100 is 15 times faster.

The 4080 is _25x_ faster than the M2 on pure fp32 (which is what most GPUs are doing most of the time). Apple compared the M2 to the laptop 4080, using numbers heavily biased to them (running a 4080 at 10W does tend to make it not perform, yes). Not a single benchmark in the world has supported Apple's claim that the GPU in the M2 is that powerful. It's just yet another cute embedded GPU that does the job, but nothin…

ML inference is not generally FP32 anymore. I was going off of the TOPs numbers for ML from a few sources, which generally agree M2 is about 22 TOPS and 4080 (desktop) is about 50.

But in any event, yes, that was my point. UMA is a huge advantage, the GPU itself is too weak to be serious.

But it’s a lot easier to drop a dramatically beefier GPU into a new design than it is to update the entire platform for UMA. Apple has a huge opportunity here… whether tbey pursue it or not remains to be seen.

Re: Transformer architecture optimized for Apple Silicon

#312
post #64
post #37

Earlier quoted context omitted.

> it will be convenience rather than privacy which is the deciding factor This. Nobody really cares about local processing for privacy. Even those that claim to often don't mean it. Remember the total freakout over Apple's proposed local, privacy-preserving processing to detect CSAM before uploading it to iCloud? The consensus seemed to be that secret, opaque, and un-auditable cloud-based scanning was much preferable…

Local CSAM scanning wasn't an open book either. Also, it wasn't mutually exclusive with cloud scanning.

It was pretty open — auditable and cryptographically proven sources for all triggering hashes, auditable updates. And it was mutually exclusive with cloud scanning, as it was part of E2EE.

Re: Transformer architecture optimized for Apple Silicon

#313
post #64

Earlier quoted context omitted.

Local CSAM scanning wasn't an open book either. Also, it wasn't mutually exclusive with cloud scanning.

However, there was enough of an outcrying by the community and my experts that it was pulled. Therefore end users do care about privacy when they understand the implications. After all, it's hard to care when you don't have power as an end user too act on those feelings.

Do you think the current regime of Google/etc doing ad hoc scans on cloud storage, with no transparency about who is requesting scans for what, is more private?

Re: Transformer architecture optimized for Apple Silicon

#314

Earlier quoted context omitted.

This is entirely my point. They’re all crap at a fundamental level, none of them outperform any other. But the next generation, which will assuredly be backed by a LLM, will be shockingly powerful.

The MKBHD personality is onto Monte Carlo testing in his device reviews. He has an intuition for good UIX. Google and Apple should pay people like MKBHD and his production crew a billion dollars every year to film achievable, desirable text, voice, camera A.I. assisted use-cases.

They don’t work like that though. Instead they spend $3b / year on hiring fungible cog SDE 2 job code 67483’s and look at gantt chart roadmaps produced by fungible cog PM 3 job code 74842’s for fungible cog SDM 7 job code 84747’s with the nominal assistance of centralized (I.e., marginalized) fungible cog UX 5 job code 35563’s then wonder why it’s perennially late and complete garbage. Except no one wonders that because fungible cogs don’t have wonder in their job responsibly matrix.

Re: Transformer architecture optimized for Apple Silicon

#315

Earlier quoted context omitted.

My understanding is that llama.cpp is a port of meta’s llama release that uses PyTorch. Then I hope we can run llama with Apple Neural Engine.

llama.cpp does not use PyTorch, it uses a niche ML library made by its creator that only runs on the CPU.

I’m under the impression that llama.cpp is a fork of llama (which uses pytorch)

Re: Transformer architecture optimized for Apple Silicon

#316
post #128

Earlier quoted context omitted.

Jobs has been dead for over ten years

The cinema industry know how to squeeze a billion dollar a year out of Apple. Do the Apple or Sony ”social engineers“ have total view? When Amazon was sky rocketing and Blue Origin had not reached orbit a niche producer from the cinema industry had Jeff B paying a lot although he might have backed out of that deal like giving up on the Android phone with a lot of cameras.

>When Amazon was sky rocketing and Blue Origin had not reached orbit...

Blue Origin still has not reached orbit.

Re: Transformer architecture optimized for Apple Silicon

#317

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

There is no benefit to OpenAI not needing a data centre for their tech. The only ways they have to lock their tech up is safety and ethics theatre and the hardware budget required to run it.

if you work at OpenAI and you’re optimizing anything you’re not doing your job right.

Re: Transformer architecture optimized for Apple Silicon

#318
post #18

Earlier quoted context omitted.

5 years? Bit of a long stretch with how much focus is going into AI. I’d be surprised if it’s not possible within two iterations of the new iPhone.

GPT-3 was said to require something like 150gb of VRAM. I don't see that gap being bridged in phones within 2 years.

Dall-e 2 was claimed to need A100s. Stable Diffusion runs on 6 year old gamer cards.

Re: Transformer architecture optimized for Apple Silicon

#319

Earlier quoted context omitted.

The 4080 is _25x_ faster than the M2 on pure fp32 (which is what most GPUs are doing most of the time). Apple compared the M2 to the laptop 4080, using numbers heavily biased to them (running a 4080 at 10W does tend to make it not perform, yes). Not a single benchmark in the world has supported Apple's claim that the GPU in the M2 is that powerful. It's just yet another cute embedded GPU that does the job, but nothin…

ML inference is not generally FP32 anymore. I was going off of the TOPs numbers for ML from a few sources, which generally agree M2 is about 22 TOPS and 4080 (desktop) is about 50. But in any event, yes, that was my point. UMA is a huge advantage, the GPU itself is too weak to be serious. But it’s a lot easier to drop a dramatically beefier GPU into a new design than it is to update the entire platform for UMA. Apple…

> whether tbey pursue it or not remains to be seen.

Pursue what though?

UMA is cool, but kinda meaningless if the majority of Macbooks are min-spec. That leaves you with 4-5gb of VRAM, assuming you've left nothing open. What is Apple going to do with that UMA that other manufacturers cannot?

It's certainly nice that 128gb Macs exist for models that might be too big to otherwise load into memory. It's useless for production inferencing though, and I struggle to imagine the "opportunities" they're missing out on here.

Re: Transformer architecture optimized for Apple Silicon

#320
post #219

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

it will actually play out differently: Apple will pay Open AI to run their models on Apple Silicon. And Apple will guarantee the model is protected using hw keys. Jail broken phones won’t have access to the local AI

I feel the sun is about to set on general purpose computing. Everything from here on out is going towards being silicon locked.
Post reply on HN