Earlier quoted context omitted.
Maybe you should have asked a better question. :P
What do you get if you multiply six by nine?
iPhone 17 Pro Demonstrated Running a 400B LLM
61–70 of 362 posts
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#62This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains
Only way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#63This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains
Realistically you need +300GB/s fast access memory to the accelerator, with enough memory to fully hold at least greater than 4bit quants. That's at least 380GB of memory. You can gimmick a demo like this with an ssd, but the ssd is just not fast enough to meet the minim specs for anything more than showing off a neat trick on twitter.
The only hope for a handheld execution of a practical, and capable AI model is both an algorithmic breakthrough that does way more with less, and custom silicon designed for running that type of model. The transformer architecture is neat, but it's just not up for that task, and I doubt anyone's really going to want to build silicon for it.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#64Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."
I mean size says nothing, you could do it on a Pi Zero with sufficient storage attached. So this post is like saying that yes an iPhone is Turing complete. Or at least not locked down so far that you're unable to do it.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#65This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains
Only way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#66Earlier quoted context omitted.
> maybe even learned prefetching for what the next experts will be Experts are predicted by layer and the individual layer reads are quite small, so this is not really feasible. There's just not enough information to guide a prefetch.
Manually no. It would have to be learned, and making the expert selection predictable would need to be a training metric to minimize.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#67Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."
laughed when it slowly began to type that out
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#68A year ago this would have been considered impossible. The hardware is moving faster than anyone's software assumptions.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#69Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."
Better than waiting 7.5 million years to have a tell you the answer is 42.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#70Earlier quoted context omitted.
Plus all those pricey 512GB Mac Studios they are selling to YouTubers.
They don't offer the 512 gig RAM variant anymore. Outside of social media influencers and the occasional AI researcher, the market for $10K desktops is vanishingly small.