Live data from Hacker News

Transformer architecture optimized for Apple Silicon

github.com

61–70 of 342 posts

Re: Transformer architecture optimized for Apple Silicon

#61
post #18

Earlier quoted context omitted.

5 years? Bit of a long stretch with how much focus is going into AI. I’d be surprised if it’s not possible within two iterations of the new iPhone.

GPT-3 was said to require something like 150gb of VRAM. I don't see that gap being bridged in phones within 2 years.

Alpaca already works in just 4GB of RAM. This stuff is moving incredibly fast.

Re: Transformer architecture optimized for Apple Silicon

#62
post #60

Earlier quoted context omitted.

I really think you have hit the nail on the head here. Apple has a ridiculous, almost unfathomably deep moat for training and running personalised, customised LLMs and other AI models on the 'edge' with these Apple Silicon chips in all their devices. We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waitin…

Siri was supposed to become locally processed but sometimes can’t even set a timer because of a connection time-out. Or she’ll use her fancy ML speech recognition model to turn “set a timer for 3 minutes 10 seconds” into “search for trinity tensor”. So much for an “unfathomable moat”.

You see the only reason it does that is because your … holding your phone wrong

Re: Transformer architecture optimized for Apple Silicon

#63
post #60

Earlier quoted context omitted.

I really think you have hit the nail on the head here. Apple has a ridiculous, almost unfathomably deep moat for training and running personalised, customised LLMs and other AI models on the 'edge' with these Apple Silicon chips in all their devices. We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waitin…

Siri was supposed to become locally processed but sometimes can’t even set a timer because of a connection time-out. Or she’ll use her fancy ML speech recognition model to turn “set a timer for 3 minutes 10 seconds” into “search for trinity tensor”. So much for an “unfathomable moat”.

Siri is so bad it needs to be scrapped and rethought from the ground up (it can’t even give me the time when internet is down at my rural house). No one at Apple who wants a career there would dare propose that. The switch to a LLM architecture could be the perfect transition point for this.

Re: Transformer architecture optimized for Apple Silicon

#64
post #37
post #29

Earlier quoted context omitted.

Convenience will be the biggest factor. Whoever makes it easier for the end consumer to get what they want wins. It's why ChatGPT made such a big splash in comparison to all the other AI models which were also impressive. Local inference may play a part in that if it's quicker but if the last twenty years are anything to go by it will be convenience rather than privacy which is the deciding factor. And Apple do tend…

> it will be convenience rather than privacy which is the deciding factor This. Nobody really cares about local processing for privacy. Even those that claim to often don't mean it. Remember the total freakout over Apple's proposed local, privacy-preserving processing to detect CSAM before uploading it to iCloud? The consensus seemed to be that secret, opaque, and un-auditable cloud-based scanning was much preferable…

Local CSAM scanning wasn't an open book either. Also, it wasn't mutually exclusive with cloud scanning.

Re: Transformer architecture optimized for Apple Silicon

#65
post #21

I find it fascinating that this was released after the June 2022 launch of the M2 chipset and line of products, and yet Apple had no desire to show relative performance of M2 vs. M1 here - even in the simultaneous announcement here: https://machinelearning.apple.com/research/neural-engine-tra... It's fascinating to me that at least one of two things is true: either (a) Apple has lost its ability to coordinate "hype"…

Interesting. Maybe they just wanted to publish quickly and didn’t take the time to benchmark across? Seems like a miss.

M2 NPU is supposed to be 44% better than M1

https://www.cpu-monkey.com/en/article/apple_m2_vs_apple_m1__...

Re: Transformer architecture optimized for Apple Silicon

#66
post #47

Earlier quoted context omitted.

yes, there are billions of parameters necessary. but large language models only came out about 5 years ago. I'm confident 5 years from now the parameters necessary to get gpt-4 performance will be decreased orders of magnitude. at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.

Wait, so there's a way to make a model as smart as GPT but with less parameters? Isn't that why it's so good?

The rumor I've heard is that GPT4 didn't meaningfully increase the parameter count versus GPT3.5, but instead focused on training and structural improvements.

Re: Transformer architecture optimized for Apple Silicon

#67
post #39

Earlier quoted context omitted.

Isn't GPT so complex that it requires hundreds of GB of ram to be used? How's it going to run on iphone?

yes, there are billions of parameters necessary. but large language models only came out about 5 years ago. I'm confident 5 years from now the parameters necessary to get gpt-4 performance will be decreased orders of magnitude. at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.

There will also be hardware improvements (as always) and ASIC chips specifically designed for running this kind of model. For example, see this "Optical Transformers" paper [0] and its HN discussion [1] from last month.

[0] https://arxiv.org/abs/2302.10360

[1] https://news.ycombinator.com/item?id=34905210

Re: Transformer architecture optimized for Apple Silicon

#68
post #45

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

OpenAI's relationship with Microsoft pretty much makes it big tech. I think you have a point as long Apple keeps innovating its hardware, but it's not really a David vs Goliath situation.

Nvidia is also no David to Apple’s Goliath. They are both Goliaths who will most likely be battling to build the best hardware for LLMs in the near future.

Re: Transformer architecture optimized for Apple Silicon

#69
post #60

Earlier quoted context omitted.

I really think you have hit the nail on the head here. Apple has a ridiculous, almost unfathomably deep moat for training and running personalised, customised LLMs and other AI models on the 'edge' with these Apple Silicon chips in all their devices. We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waitin…

Siri was supposed to become locally processed but sometimes can’t even set a timer because of a connection time-out. Or she’ll use her fancy ML speech recognition model to turn “set a timer for 3 minutes 10 seconds” into “search for trinity tensor”. So much for an “unfathomable moat”.

[deleted]

Re: Transformer architecture optimized for Apple Silicon

#70

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

I really think you have hit the nail on the head here. Apple has a ridiculous, almost unfathomably deep moat for training and running personalised, customised LLMs and other AI models on the 'edge' with these Apple Silicon chips in all their devices. We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waitin…

Maybe, but Apple doesn’t have a search to rival Google (or even an assistant, given the state of Siri).

Focusing on privacy and on-device learning is great, but when the strength of these models is in consuming all the data they can hoover up your motive is at odds with your philosophy.

Post reply on HN