Earlier quoted context omitted.
I wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.
A more whimsical method is to put the thing in a glass of water with the cord sticking out. :-) https://www.reddit.com/r/EmulationOnAndroid/comments/1m269k0...
iPhone 17 Pro Demonstrated Running a 400B LLM
271–280 of 362 posts
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#272With all the money you will save on subscription fees you should be able to afford treatment for your psychosis!
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#273Earlier quoted context omitted.
> The difference is the iPhone has wider memory buses and uses faster LPDDR5 memory. Apple places the RAM dies directly on the same package as the SoC (PoP — Package on Package), minimizing latency. Some Android phones have started to do this, too. Package-on-Package has been used in mobile SoCs for a long time. This wasn't an Apple invention. It's not new, either. It's been this way for 10+ years. Even cheap Raspber…
> The memory bandwidth of flagship iPhone models is similar to the memory bandwidth of flagship Android phones More correct to say that the memory bandwidth of ALL iPhone models is similar to the memory bandwidth of flagship Android models. The A18 and A18 pro do not differ in memory bandwidth.
A18 Pro has a modest memory bandwidth advantage over the standard A18, which is part of why it can support ProRes recording and always-on display while the standard A18 cannot.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#274Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#275Run an incredible 400B parameters on a handheld device. 0.6 t/s, wait 30 seconds to see what these billions of calculations get us: "That is a profound observation, and you are absolutely right ..."
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#276Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#277My iPad Air with M2 can run local LLMs rather well. But it gets ridiculously hot within seconds and starts throttling.
I wonder if anyone has made a liquid cooling system for ipads / phones. Like, a sealed thing that seals onto the back of the device and circulates cooling water directly against the back surface.
Apple fans never cease to amaze me.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#278"0.6 t/s" This is a toy. We need to build open infrastructure in the cloud capable of hosting a robust ecosystem of open weights. And then we need to build very large scale open weights. That's the only way we don't get owned by the hyperscalers. At the edge isn't going to happen in a meaningful way to save us.
use the experience we gain from both to bolster the other.
a future where we are unable to locally run is kind of troubling. as is a future with no open cloud. we need both to stop some of the horrors the hyperscalers will happily inflict.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#279I installed Termux on an old Android phone last week (running LineageOS), and then using Termux installed Ollama and a small model. It ran terribly, but it did run.
Instead, take the advantage of Termux power, namely the fact that you can install things like Openclaw or Gemini-cli. Google Ai plus or Pro plans are actually really good value, considering they bundle it with storage.
https://www.mobile-hacker.com/2025/07/09/how-to-install-gemi...
There is also Termux:GUI with bindings for languages, which you can use to vibecode your own GUI app, which then can basically serve as an interface to an agent, an Termux API which lets you interface with the phone, including USB devices.
Furthermore, termux has the cloudflared package availble, which lets you use clouflared free ssh tunnels (as long as you have a domain name).
All put together, you can do some pretty cool things.
Re: iPhone 17 Pro Demonstrated Running a 400B LLM
#280Earlier quoted context omitted.
I highly doubt the A20 Pro will be slower than the A19 Pro - particularly for AI workloads.
We're talking six orders of magnitude difference between 0.6t/sec and 35kt/sec. While there are problems that can be solved with 0.6t/sec, particularly offline, at the edge, in the field applications, these are currently vastly outnumbered by other applications. There's just no competing. Local sucks.
absolutely, however this doesn’t mean we should abandon local. i can’t remember who, but someone in the ai nuts and bolts arena said “smaller local models is where the exciting stuff is happening right now. it’s the area real fast progression is happening.” and it seems to be true. new big models aren’t making near the leaps smaller models are.
it’s so important we keep moving forward on running locally for the same reason it was important for us to use open standards when building the internet. if we hadn’t we’d all be connected through aol with 10 hours/month allowed internet usage and termed in through a sun workstation renting cpu cycles from some mainframe company at like “you’ve got 10,000 cpu cycles left on your monthly plan, please deposit $500 for 5,000 more.”
while all of this this is before my time, i’ve heard and read so many horror stories about how people could only connect through dumb terminals to “you wouldn’t believe it, computers then were the size of buildings” 1000 miles away and had to sign up for workload timeslots. make no mistake, this is the future these companies want, they want us to rent everything and own nothing.