Live data from Hacker News

iPhone 17 Pro Demonstrated Running a 400B LLM

twitter.com

251–260 of 362 posts

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#253
post #214

Qwen's MoE models are god awful when they are only running 2B parameters or whatever they downscale to while active. It isn't a 400B model if there's only several orders of magnitude less parameters active when you're actually inferencing...

[dead]

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#256

Earlier quoted context omitted.

> add more RAM and GPU to the next iPhone and it's not a toy anymore We're not going to get more RAM and GPU in consumer devices. All of the supply is going into data center build outs. As the hyper scaler gamble on the future continues, we get left with weaker (or more expensive) devices - not stronger ones. The market makers make more money if we're left to thin clients. They're also the ones who control supply and…

I highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.

SK Hynix: "Hold my LPDDR5X"

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#259

Earlier quoted context omitted.

> add more RAM and GPU to the next iPhone and it's not a toy anymore We're not going to get more RAM and GPU in consumer devices. All of the supply is going into data center build outs. As the hyper scaler gamble on the future continues, we get left with weaker (or more expensive) devices - not stronger ones. The market makers make more money if we're left to thin clients. They're also the ones who control supply and…

I highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.

We're talking six orders of magnitude difference between 0.6t/sec and 35kt/sec.

While there are problems that can be solved with 0.6t/sec, particularly offline, at the edge, in the field applications, these are currently vastly outnumbered by other applications.

There's just no competing. Local sucks.

Post reply on HN