Live data from Hacker News

iPhone 17 Pro Demonstrated Running a 400B LLM

twitter.com

81–90 of 362 posts

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#82
post #72

Earlier quoted context omitted.

Possibly this just isn't the generation of hardware to solve this problem in? We're like, what three or four years in at most, and only barely two in towards AI assisted development being practical. I wouldn't want to be the first mover here, and I don't know if it's a good point in history to try and solve the problem. Everything we're doing right now with AI, we will likely not be doing in five years. If I were run…

If I was running a company like Apple, I'd be working with Khronos to kill CUDA since yesterday. There are multiple trillions of dollars that could be Apple's if they sign CUDA drivers on macOS, or create a CUDA-compatible layer. Instead, Apple is spinning their wheels and promoting nothingburger technology like the NPU and MPS. It's not like Apple's GPU designs are world-class anyways, they're basically neck-and-nec…

CUDA is not the real issue, AMD's HIP offers source-level compatibility with CUDA code, and ZLUDA even provides raw binary compatibility. nVidia GPUs really are quite good, and the projected advantages of going multi-vendor just aren't worth the hassle given the amount of architecture-specificity GPUs are going to have.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#83
post #71

Earlier quoted context omitted.

I don't think we are ever going to win this. The general population loves being glazed way too much.

The other day, I got: "You are absolutely right to be confused" That was the closest AI has been to calling me "dumb meatbag".

"Carrot: The Musical" in the Carrot weather app, all about the AI and her developer meatbag, is on point.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#84

Earlier quoted context omitted.

If I was running a company like Apple, I'd be working with Khronos to kill CUDA since yesterday. There are multiple trillions of dollars that could be Apple's if they sign CUDA drivers on macOS, or create a CUDA-compatible layer. Instead, Apple is spinning their wheels and promoting nothingburger technology like the NPU and MPS. It's not like Apple's GPU designs are world-class anyways, they're basically neck-and-nec…

CUDA is not the real issue, AMD's HIP offers source-level compatibility with CUDA code, and ZLUDA even provides raw binary compatibility. nVidia GPUs really are quite good, and the projected advantages of going multi-vendor just aren't worth the hassle given the amount of architecture-specificity GPUs are going to have.

Okay, then don't kill CUDA, just sign CUDA drivers on macOS instead and quit pretending like MPS is a world-class solution. There are trillions on the table, this is not an unsolvable issue.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#85
post #6

A year ago this would have been considered impossible. The hardware is moving faster than anyone's software assumptions.

This isn't a hardware feat, this is a software triumph. They didn't make special purpose hardware to run a model. They crafted a large model so that it could run on consumer hardware (a phone).

>triumph

It’s been a lot of years, but all I can hear after reading that is … I’m making a note here, huge success

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#86
post #15

Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."

2 years ago, LLMs failed at answering coherently. Last year, they failed at answering fast on optimized servers. Now, they're failing at answering fast on underpowered handheld devices... I can't wait to see what they'll be failing to do next year.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#87
post #47
post #4

It's crazy to see a 400B model running on an iPhone. But moving forward, as the information density and architectural efficiency of smaller models continue to increase, getting high-quality, real-time inference on mobile is going to become trivial.

> moving forward, as the information density and architectural efficiency of smaller models continue to increase If they continue to increase.

They will. Either new architectures will come out that give us greater efficiency, or we will hit a point where the main thing we can do is shove more training time onto these weights to get more per byte. Similar thing is already happening organically when it comes to efficient token use; see for instance https://github.com/qlabs-eng/slowrun.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#88
post #65
post #34

Earlier quoted context omitted.

Only way to have hardware reach this sort of efficiency is to embed the model in hardware. This exists[0], but the chip in question is physically large and won't fit on a phone. [0] https://www.anuragk.com/blog/posts/Taalas.html

That's actually pretty cool, but I'd hate to freeze a models weights into silicon without having an incredibly specific and broad usecase.

[deleted]

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#89
post #15

Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."

Better than waiting 7.5 million years to have a tell you the answer is 42.

Some one should let Douglas Adams know the calculation could have been so much faster if the machine just lied.

Re: iPhone 17 Pro Demonstrated Running a 400B LLM

#90

Earlier quoted context omitted.

Plus all those pricey 512GB Mac Studios they are selling to YouTubers.

They don't offer the 512 gig RAM variant anymore. Outside of social media influencers and the occasional AI researcher, the market for $10K desktops is vanishingly small.

Huh, interesting. I wonder if there's a premium price right now for the one on my desk...

Pretty sure the M5 Ultra will be out after WWDC, so my M3 Ultra is (while still completely capable of fulfilling my needs) looking a bit long in the tooth. If I can get a good price for it now, I might be able to offset most of the M5 post WWDC...

Post reply on HN