Live data from Hacker News

Transformer architecture optimized for Apple Silicon

github.com

31–40 of 342 posts

Re: Transformer architecture optimized for Apple Silicon

#31

Earlier quoted context omitted.

I mean, it's possible to load LLMs onto a smartphone today. The utility is limited though, and if iPhone 19 still has arbitrary application memory requirements then I think it's safe to say they won't be using the most complete models. At the end of the day, OpenAI's "huge-ass LLM as a service" will probably be relevant longer than you think it will be. Local inferencing might be able to do simpler stuff (eg. compose…

I never said anything about openai not being relevant. the point is scale and good enough. gpt-4 is already good enough for most reasonable use cases, notably siri optimization. 5 years is a long time. 5 years ago llms didn't even exist.

Gpt3.5-turbo runs laps around siri and is cheap, surely Apple will acquire talent to build something that does remote queries and falls back to local for the next 1-2 years while they figure out on device accelerated LLMs

Re: Transformer architecture optimized for Apple Silicon

#32

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Isn't that quite risky business - if someone manages to leak the model, anyone could run it for free?

I can imagine similar for iPhone on the edge, when someone manages to decrypt the model it will be free to grab for anyone, unless there is going to be some proprietary thing going on only available on Apple Silicon that is undocumented.

Re: Transformer architecture optimized for Apple Silicon

#33

Earlier quoted context omitted.

This would be awesome. I don't like having to rely on closed-source cloud services for LLMs. But what incentive does Big Tech have? Especially considering that presumably they could monetize the cloud services easier. It's not like most consumers care either. I expect smaller companies like Stable Diffusion and grant/govt-funded research to bring cheap, local inference. There's definitely a lot of demand. But is ther…

i'd say apple is one of the few companies that could charge people for an "AI embedded" 128GB module, allow people to pay for a subscription in order to "access", yet have the inference happen locally. one hypothetical is that you need an icloud subscription, a token is retrieved from apple, the token "unlocks" your AI module, on the phone and allows you to do the inference. in this way apple could charge monthly for…

What other services does Apple charge a subscription fee for despite having no opex?

Re: Transformer architecture optimized for Apple Silicon

#34

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

18 months.

Re: Transformer architecture optimized for Apple Silicon

#35

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Apple M series chips are already some of the best chips for running your own models locally, especially if you want a laptop or small form factor machine instead of a big gaming desktop with a big GPU.

Apple is really getting a niche here for machines to run models locally. That’s pretty powerful.

Re: Transformer architecture optimized for Apple Silicon

#36
post #33

Earlier quoted context omitted.

i'd say apple is one of the few companies that could charge people for an "AI embedded" 128GB module, allow people to pay for a subscription in order to "access", yet have the inference happen locally. one hypothetical is that you need an icloud subscription, a token is retrieved from apple, the token "unlocks" your AI module, on the phone and allows you to do the inference. in this way apple could charge monthly for…

What other services does Apple charge a subscription fee for despite having no opex?

apple fitness I imagine has a pretty small opex (sure they pay trainers, but that doesn't scale linearly with subscriptions. the videos themselves are limited in scope and could be cached pretty trivially. the actual fitness tracking happens on device), same with apple arcade which I highly doubt has any additional expensive over the app store in general.

I can see it - "Siri+", pay $5 a month for fine tuned model upgrades straight to your device and remote fallback. local inference available for Pro devices only.

Re: Transformer architecture optimized for Apple Silicon

#37
post #29

Maybe apple will have a bigger effect on ai adoption than any other company. Local inference is huge for anything that requires even a little bit of privacy.

Convenience will be the biggest factor. Whoever makes it easier for the end consumer to get what they want wins. It's why ChatGPT made such a big splash in comparison to all the other AI models which were also impressive. Local inference may play a part in that if it's quicker but if the last twenty years are anything to go by it will be convenience rather than privacy which is the deciding factor. And Apple do tend…

> it will be convenience rather than privacy which is the deciding factor

This. Nobody really cares about local processing for privacy.

Even those that claim to often don't mean it. Remember the total freakout over Apple's proposed local, privacy-preserving processing to detect CSAM before uploading it to iCloud? The consensus seemed to be that secret, opaque, and un-auditable cloud-based scanning was much preferable.

Re: Transformer architecture optimized for Apple Silicon

#38

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

OpenAI IS big tech at this point given the massive financial and infrastructure investment from Microsoft. I think we're well past the small underdog narrative.

Re: Transformer architecture optimized for Apple Silicon

#39

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Isn't GPT so complex that it requires hundreds of GB of ram to be used? How's it going to run on iphone?

Re: Transformer architecture optimized for Apple Silicon

#40

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Isn't that quite risky business - if someone manages to leak the model, anyone could run it for free? I can imagine similar for iPhone on the edge, when someone manages to decrypt the model it will be free to grab for anyone, unless there is going to be some proprietary thing going on only available on Apple Silicon that is undocumented.

I mean this risk is already there for all iphones for unlock.
Post reply on HN