Live data from Hacker News

Transformer architecture optimized for Apple Silicon

github.com

21–30 of 342 posts

Re: Transformer architecture optimized for Apple Silicon

#21
I find it fascinating that this was released after the June 2022 launch of the M2 chipset and line of products, and yet Apple had no desire to show relative performance of M2 vs. M1 here - even in the simultaneous announcement here: https://machinelearning.apple.com/research/neural-engine-tra...

It's fascinating to me that at least one of two things is true: either (a) Apple has lost its ability to coordinate "hype" between its teams, or (b) the difference between comparable levels e.g. the M1 Max vs. the M2 Max are so negligible that they don't look good in an announcement like this.

Has anyone run inference for LLMs or other transformer models on comparable M1 and M2 Macs? Are there good benchmarks for this specific workload?

Re: Transformer architecture optimized for Apple Silicon

#22

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

I mean, it's possible to load LLMs onto a smartphone today. The utility is limited though, and if iPhone 19 still has arbitrary application memory requirements then I think it's safe to say they won't be using the most complete models. At the end of the day, OpenAI's "huge-ass LLM as a service" will probably be relevant longer than you think it will be. Local inferencing might be able to do simpler stuff (eg. compose assistant speech) but this is already possible with pruned models and CPU acceleration.

The bull mindset is fun to watch unfold (especially here on HN) but I think people should temper their expectations.

Re: Transformer architecture optimized for Apple Silicon

#23

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

[deleted]

Re: Transformer architecture optimized for Apple Silicon

#24

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Maybe we should launch 100 of them out into space in different directions. Very low mass means we should be able to push it to a pretty high velocity.

> into space in different directions

Different directions - this is the fundamental challenge. It's hard to know what these directions are when we're flying blind at the cutting edge of technology.

We're all human. I suspect out of the 100, you'll have 95 of them going into the same as before areas. Maybe 5 truly understand the above point and explicitly strike down ideas that have been done before.

GPT-3 has been live for a while now. The industry has hundreds of story generator apps specialized for various kinds of content generation. Very few are really thinking about AGI or chain-of-thought reasoning etc as an example.

So I don't think launching 100 of them into different directions would work. Maybe 5 of them with both the expertise and direction to pursue seemingly impossible goals.

Re: Transformer architecture optimized for Apple Silicon

#25

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

Maybe we should launch 100 of them out into space in different directions. Very low mass means we should be able to push it to a pretty high velocity.

I know some people don’t like iPhones but this is taking it a bit far.

Re: Transformer architecture optimized for Apple Silicon

#26

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

This would be awesome. I don't like having to rely on closed-source cloud services for LLMs.

But what incentive does Big Tech have? Especially considering that presumably they could monetize the cloud services easier. It's not like most consumers care either.

I expect smaller companies like Stable Diffusion and grant/govt-funded research to bring cheap, local inference. There's definitely a lot of demand. But is there more demand than cloud-based services, so that it's economically viable for the biggest companies (Apple)?

Re: Transformer architecture optimized for Apple Silicon

#27
post #18

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

5 years? Bit of a long stretch with how much focus is going into AI. I’d be surprised if it’s not possible within two iterations of the new iPhone.

GPT-3 was said to require something like 150gb of VRAM. I don't see that gap being bridged in phones within 2 years.

Re: Transformer architecture optimized for Apple Silicon

#28

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

I mean, it's possible to load LLMs onto a smartphone today. The utility is limited though, and if iPhone 19 still has arbitrary application memory requirements then I think it's safe to say they won't be using the most complete models. At the end of the day, OpenAI's "huge-ass LLM as a service" will probably be relevant longer than you think it will be. Local inferencing might be able to do simpler stuff (eg. compose…

I never said anything about openai not being relevant. the point is scale and good enough. gpt-4 is already good enough for most reasonable use cases, notably siri optimization. 5 years is a long time. 5 years ago llms didn't even exist.

Re: Transformer architecture optimized for Apple Silicon

#29

Maybe apple will have a bigger effect on ai adoption than any other company. Local inference is huge for anything that requires even a little bit of privacy.

Convenience will be the biggest factor. Whoever makes it easier for the end consumer to get what they want wins. It's why ChatGPT made such a big splash in comparison to all the other AI models which were also impressive. Local inference may play a part in that if it's quicker but if the last twenty years are anything to go by it will be convenience rather than privacy which is the deciding factor. And Apple do tend to make the UX slicker than other companies so I wouldn't put it past them ending up the biggest player.

Re: Transformer architecture optimized for Apple Silicon

#30

i'd say within 5 years apple will have optimized apple silicon and their tech, along with language model improvements, such that you will be able to get gpt-4 level performance in the iPhone 19 with inference happening entirely locally. openai is doing great work and is serious competition, but I think many underestimate big tech. once they're properly motivated they'll catch up quick. I think we can agree that opena…

This would be awesome. I don't like having to rely on closed-source cloud services for LLMs. But what incentive does Big Tech have? Especially considering that presumably they could monetize the cloud services easier. It's not like most consumers care either. I expect smaller companies like Stable Diffusion and grant/govt-funded research to bring cheap, local inference. There's definitely a lot of demand. But is ther…

i'd say apple is one of the few companies that could charge people for an "AI embedded" 128GB module, allow people to pay for a subscription in order to "access", yet have the inference happen locally.

one hypothetical is that you need an icloud subscription, a token is retrieved from apple, the token "unlocks" your AI module, on the phone and allows you to do the inference.

in this way apple could charge monthly for this and claim that the inference happens locally. sadly in this was it was similar to the whole csam debacle

apple is also one of the few companies that could realistically get manufacturers to massively produce a hypothetical "AI model chip" that has the model on device, at a quantity that would make it realistic to pay for in a hypothetical iPhone 19 Pro model.

Post reply on HN