Live data from Hacker News

iPhone 17 Pro Demonstrated Running a 400B LLM

twitter.com

121–130 of 362 posts

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#122
post #65
post #34

Earlier quoted context omitted.

Only way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html

That's actually pretty cool, but I'd hate to freeze a models weights into silicon without having an incredibly specific and broad usecase.

I mean if it was small enough to fit in an iPhone why not? Every year you would fabricate the new chip with the best model. They do it already with the camera pipeline chips.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#123

Earlier quoted context omitted.

Looked at a certain way it's incredible that a 40-odd year old comedy sci-fi series is so accurate about the expected quality of (at least some) AI output. Which makes it even funnier. It makes me a little sad that Douglas Adams didn't live to see it.

Also check out "The Great Automatic Grammatizator" by Roald Dahl for another eerily accurate scifi description of LLMs written in 1954: https://gwern.net/doc/fiction/science-fiction/1953-dahl-theg...

"Can write a prize-winning novel in fifteen minutes" - that's quite optimistic by modern standards!

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#124

I have some macro opinions about Apple - not sure if I'm correct, but tell me what you think. Apple has always seen RAM as an economic advantage for their platform: Make the development effort to ensure that the OS and apps work well with minimal memory and save billions every year in hardware costs. In 2026, iPhones still come with 8Gb of RAM, Pro/Max come with 12Gb. The problem is that AI (ML/LLM training and infer…

> nor create specialized SoCs with ML cores that obviate the need for lots and lots of RAM

Why do you say they can't do this?

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#125
post #65
post #34

Earlier quoted context omitted.

Only way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html

That's actually pretty cool, but I'd hate to freeze a models weights into silicon without having an incredibly specific and broad usecase.

Sounds like just the sort of thing FGPA's were made for.

The $$$ would probably make my eyes bleed tho.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#126
post #4

It's crazy to see a 400B model running on an iPhone. But moving forward, as the information density and architectural efficiency of smaller models continue to increase, getting high-quality, real-time inference on mobile is going to become trivial.

Probably 2x speed for Mac Studio this year if they do double NAND ( or quad?)

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#128
post #47
post #4

It's crazy to see a 400B model running on an iPhone. But moving forward, as the information density and architectural efficiency of smaller models continue to increase, getting high-quality, real-time inference on mobile is going to become trivial.

> moving forward, as the information density and architectural efficiency of smaller models continue to increase If they continue to increase.

The "if" is fair. But when scaling hits diminishing returns, the field is forced to look at architectures with better capacity-per-parameter tradeoffs. It's happened before, maybe it'll happen again now.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#130
post #62

Earlier quoted context omitted.

I think you're ignoring the inevitable march of progress. Phones will get big enough to hold it soon.

I think the future is the model becoming lighter not the hardware becoming heavier

The hardware will become heavier regardless I'm afraid.
Post reply on HN