iPhone 17 Pro Demonstrated Running a 400B LLM
251–260 of 362 posts
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#252It’s 400B but it’s mixture of experts so how many are active at any time?
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#253Qwen's MoE models are god awful when they are only running 2B parameters or whatever they downscale to while active. It isn't a 400B model if there's only several orders of magnitude less parameters active when you're actually inferencing...
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#254This is awesome! How far away are we from a model of this capability level running at 100 t/s? It's unclear to me if we'll see it from miniaturization first or from hardware gains
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#255If you don't follow anemll, they also have a usable version of OpenClaw running on iPhone. With hardware and model improvements, the future is bright.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#256Earlier quoted context omitted.
> add more RAM and GPU to the next iPhone and it's not a toy anymore We're not going to get more RAM and GPU in consumer devices. All of the supply is going into data center build outs. As the hyper scaler gamble on the future continues, we get left with weaker (or more expensive) devices - not stronger ones. The market makers make more money if we're left to thin clients. They're also the ones who control supply and…
I highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#257Local LLMs are going to make people sit on their phones instead of taking to real people.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#258This sounds incredibly dangerous. Local LLMs are going to make people sit on their phones instead of taking to real people.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#259Earlier quoted context omitted.
> add more RAM and GPU to the next iPhone and it's not a toy anymore We're not going to get more RAM and GPU in consumer devices. All of the supply is going into data center build outs. As the hyper scaler gamble on the future continues, we get left with weaker (or more expensive) devices - not stronger ones. The market makers make more money if we're left to thin clients. They're also the ones who control supply and…
I highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.
While there are problems that can be solved with 0.6t/sec, particularly offline, at the edge, in the field applications, these are currently vastly outnumbered by other applications.
There's just no competing. Local sucks.