Live data from Hacker News

Gemma 4 on iPhone

apps.apple.com

201–210 of 267 posts

Re: Gemma 4 on iPhone

#201

Earlier quoted context omitted.

[flagged]

Yeah definitely a skill issue downloading and running the app and it crashing on startup. Just gotta skillfully press the icon I guess.

Out of curiosity, do you have anything running on port 8000? Maybe it's not gracefully handling a port collision.

Re: Gemma 4 on iPhone

#202

Earlier quoted context omitted.

> And there's a whole set of ethically-justifiable but rule-flagging conversations (loosely categorizable as things like "sensitive", "ethically-borderline-but-productive" or "violating sacred cows") that are now possible with this, and at a level never before possible until now. I checked the abliterate script and I don't yet understand what it does or what the result is. What are the conversations this enables?

1) Coming up with any valid criticism of Islam at all (for some reason, criticisms of Christianity or Judaism are perfectly allowed even with public models!). 2) Asking questions about sketchy things. Simply asking should not be censored. 3) I don't use it for this, but porn or foul language. 4) Imitating or representing a public figure is often blocked. 5) Asking security-related questions when you are trying to do…

[deleted]

Re: Gemma 4 on iPhone

#203

English version of the page: https://apps.apple.com/us/app/google-ai-edge-gallery/id67496... Also on Android: https://play.google.com/store/apps/details?id=com.google.ai.... It's a demo app for Google's Edge project: https://ai.google.dev/edge

Gemma4 works really slow on my android e2b model on Samsung galaxy s21 ultra. Atleast 20-30 sec to warm up and then reply.

The bigger E4B model is pretty fast on my Galaxy S21 Ultra even with thinking enabled. Maybe GPU acceleration was not enabled?

Re: Gemma 4 on iPhone

#204
post #195

I’d recommend locally.ai[1] - it’s really good and has a wide range of models. Also has shortcuts support. 1. https://apps.apple.com/gb/app/locally-ai-local-ai-chat/id674...

Thanks for the link. Gemma 4 also works in this app.

Re: Gemma 4 on iPhone

#205

I really believe in the future of local models. From app developer and user, My main concern for now is bloating devices. Until we’ll have something like Apples foundation model where multiple apps could share the same model it means we have something horrible as Electron in the sense, every app is a fully blown model (browser in the electron story) instead of reusing the model. With desktops we have DLL hell for yea…

This app unlocks using the Apple Foundation model itself: https://apps.apple.com/nl/app/locally-ai-local-ai-chat/id674...

Re: Gemma 4 on iPhone

#206

English version of the page: https://apps.apple.com/us/app/google-ai-edge-gallery/id67496... Also on Android: https://play.google.com/store/apps/details?id=com.google.ai.... It's a demo app for Google's Edge project: https://ai.google.dev/edge

Gemma4 works really slow on my android e2b model on Samsung galaxy s21 ultra. Atleast 20-30 sec to warm up and then reply.

Running LLMs is probably the first time I find that the SoC of that generation to lack. Even Google's underpowered Tensor CPUs make a huge difference when it comes to LLM performance.

You can check your settings for GPU acceleration, it's possible that enabling that makes a big difference.

From what I've found online the difference may also simply be Snapdragon versus Exynos GPU driver optimizations, in which case I don't think the performance can be fixed by anyone but Samsung. Others online seem to get decent performance out of the model on the S21 Ultra at the very least.

Re: Gemma 4 on iPhone

#208

Earlier quoted context omitted.

Gemma4 works really slow on my android e2b model on Samsung galaxy s21 ultra. Atleast 20-30 sec to warm up and then reply.

The bigger E4B model is pretty fast on my Galaxy S21 Ultra even with thinking enabled. Maybe GPU acceleration was not enabled?

I think there's quite the performance difference between the S21 Ultra (Snapdragon 888) and the S21 Ultra (Exynos 2100).

Qualcomm has optimized libraries for running LLMs on their chips that I don't believe Samsung has bothered with.

Re: Gemma 4 on iPhone

#209
post #57

Earlier quoted context omitted.

> or in the cloud but way more expensive then it is today. Why? It's widely understood that the big players are making profit on inference. The only reason they still have losses is because training is so expensive, but you need to do that no matter whether the models are running in the cloud or on your device. If you think about it, it's always going to be cheaper and more energy-efficient to have dedicated cloud ha…

> It's widely understood that the big players are making profit on inference. This is most definitely not widely understood. We still don't know yet. There's tons of discussions about people disagreeing on whether it really is profitable. Unless you have proof, don't say "this is widely understood".

You can look at open source models hosted by various companies that have no reason to host them on a loss.

Re: Gemma 4 on iPhone

#210

Earlier quoted context omitted.

> or in the cloud but way more expensive then it is today. Why? It's widely understood that the big players are making profit on inference. The only reason they still have losses is because training is so expensive, but you need to do that no matter whether the models are running in the cloud or on your device. If you think about it, it's always going to be cheaper and more energy-efficient to have dedicated cloud ha…

> It's widely understood that the big players are making profit on inference. I love the whole “they are making money if you ignore training costs” bit. It is always great to see somebody say something like “if you look at the amount of money that they’re spending it looks bad, but if you look away it looks pretty good” like it’s the money version of a solar eclipse

It is called sunk cost. The marginal cost is what sets the lower limit. They will always be able to sell at the marginal cost of inference.
Post reply on HN