Live data from Hacker News

iPhone 17 Pro Demonstrated Running a 400B LLM

twitter.com

221–230 of 362 posts

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#221
post #143

My iPad Air with M2 can run local LLMs rather well. But it gets ridiculously hot within seconds and starts throttling.

I wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.

I have a small portable fan that I place under it basically any time I use it for any development work. It gets thermally throttled pretty fast otherwise. It's definitely the wrong machine for my needs but it's what I gotta work with for now.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#222
post #23

This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains

It will never be possible on a smart phone. I know that sounds cynical, but there's basically no path to making this possible from an engineering perspective.

No one needs more than 640K!

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#223

A year ago this would have been considered impossible. The hardware is moving faster than anyone's software assumptions.

Does iPhone have some kind of hardware acceleration for neural netwoeks/ai ?

Yes, a Neural Engine and on the latest A19 tensor processing on the GPU cores (neural accelerator).

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#224
post #204
post #106

Earlier quoted context omitted.

The Bistromathics? That's not incorrect, it's simply too advanced for us to understand.

“What do you get if you multiply six by nine?” (One) source: https://www.reddit.com/r/Fedora/comments/1mjudsm/comment/n7d...

You also have the problem that if the both the ultimate answer to life the universe and everything, and the ultimate question to life the universe and everything, are know at the same time in the same universe. The universe is spontaneously replaced with a slightly more absurd universe to ensure that both the question and answer become meaningless.

To quote the message from the universes creators to its creation “We apologise for the inconvenience”. Does seem to sum up Douglas Adam’s views on absurdity of life.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#225
post #143

My iPad Air with M2 can run local LLMs rather well. But it gets ridiculously hot within seconds and starts throttling.

I wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.

You can buy a liquid cooled tablet.

https://onexplayerstore.com/products/onexplayer-super-x?vari...

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#227

Earlier quoted context omitted.

I wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.

You can buy a liquid cooled tablet. https://onexplayerstore.com/products/onexplayer-super-x?vari...

ipad pro actually preforms fairly comparably according to geekbench

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#229
post #63
post #23

This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains

Probably 15 to 20 years, if ever. This phone is only running this model in the technical sense of running, but not in a practical sense. Ignore the 0.4tk/s, that's nothing. What's really makes this example bullshit is the fact that there is no way the phone has a enough ram to hold any reasonable amount of context for that model. Context requirements are not insignificant, and as the context grows, the speed of the o…

This should be the top comment

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#230
post #10

> SSD streaming to GPU Is this solution based on what Apple describes in their 2023 paper 'LLM in a flash' [1]? 1: https://arxiv.org/abs/2312.11514

Yes. I collected some details here: https://simonwillison.net/2026/Mar/18/llm-in-a-flash/

I guess this is all set up to show off the new high-bandwidth-flash stuff that's due out soon?
Post reply on HN