The bottleneck with compute at the edge is (and will be) model size (both app download time and storage space on device). Stable Diffusion sits at about 2GB for fp16, Whisper Medium at 1.53GB, LLAMA is 120GB. Sure, Apple can ship an optimized model ( 1GB.
I wonder if they’d create a simple version that lives on device that can call on a ‘cleverer’ version for more extensive tasks - like how ChatGPT is using plugins. Most interactions probably don’t require full power
Transformer architecture optimized for Apple Silicon
281–290 of 342 posts
Re: Transformer architecture optimized for Apple Silicon
#282Earlier quoted context omitted.
yes, there are billions of parameters necessary. but large language models only came out about 5 years ago. I'm confident 5 years from now the parameters necessary to get gpt-4 performance will be decreased orders of magnitude. at the very least, even if that's not the case, inference will be drastically less gpu heavy by then I suspect.
dont underestimate how many of those parameters are actually necessary to support multiple languages. If you focus on english only, this can easily reduce the paramters 5fold
LLMs seem to be comfortable with hundreds of programming languages, DSLs and application specific syntaxes so how does supporting a couple more natural languages become so expensive?
I see how more training data would be needed, but I don't understand how that maps to a greater parameter count.
Re: Transformer architecture optimized for Apple Silicon
#283i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…
Re: Transformer architecture optimized for Apple Silicon
#284Is there also something available to make use of the ANE during _training_? E.g fine-tuning BERT on an M1 Mac in a couple of hours? (This here only applies to inference, right?)
Re: Transformer architecture optimized for Apple Silicon
#285In regards to LLMs there is a collision between Apple's extremely good chip design capabilities and Apple's insistence that Siri never says anything that isn't 100% scripted and 100% certain to not bad. Up until now, they've chosen to limit Siri functionality rather than leave anything to chance. LLMs will absolutely be able to run locally, but whether Apple will be able to stop worrying and love the model remains to…
...it's shifting...
Re: Transformer architecture optimized for Apple Silicon
#286i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…
Re: Transformer architecture optimized for Apple Silicon
#287i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…
I really think you have hit the nail on the head here. Apple has a ridiculous, almost unfathomably deep moat for training and running personalised, customised LLMs and other AI models on the 'edge' with these Apple Silicon chips in all their devices. We must be talking orders of magnitude differences in operational cost, not to mention completely unique features like privacy. The very definition of disruption, waitin…
Apple has decent hw but no LLM software. Unfortunately, in LLM space software changes are the ones driving performance for now, since the space hasn't stabilized yet. Since their competitors control the software, they get to adapt it to their hardware. That is, Google and Microsoft are going to adapt GPT/Bard to Qualcomm ARM etc. while Apple is being ignored. Unless Apple gets in with their own LLM (quite possible), their hw advantage will end up not mattering one bit.
Re: Transformer architecture optimized for Apple Silicon
#288i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…
Crazy bet, what makes you think that? You cannot optimize infinitely. Raytracing probably had decades of ppl trying to make it run fast and yet even today you need strong hardware
Re: Transformer architecture optimized for Apple Silicon
#289Earlier quoted context omitted.
It's probably not glasses. Nobody likes wearing glasses. Every attempt at making tech glasses has failed. It'll probably be the Apple Watch.
I'm not a VR/AR fan but if Apple can come out with super high quality AR headset that can augment/replace monitors I would buy one instantly. I might be earlier than mainstream adopter but if the tech ever gets there this will be a natural shift. I won't buy Meta products but I keep waiting for someone to do VR workspace right - where I can truly work and travel from anywhere with a decent chair and a desk for keyboa…
Having said that, for me to spend money on it it would (a) also be able to realistically replace my set up at home, which currently consists of 2 HiDPI screens and (b) not be that much more expensive than what that cost me.
Re: Transformer architecture optimized for Apple Silicon
#290Earlier quoted context omitted.
Isn't GPT so complex that it requires hundreds of GB of ram to be used? How's it going to run on iphone?
Hundreds of GBs of RAM in a phone is just ~6 years away if Moore's law holds. It's also likely that memory requirements will be shrunk through software and ML improvements.
It's not expected to. The consensus seems to be ~2025 https://arxiv.org/abs/1511.05956