Live data from Hacker News

Gemma 3n preview: Mobile-first AI

developers.googleblog.com

141–150 of 179 posts

Re: Gemma 3n preview: Mobile-first AI

#141
Quote: Expanded Multimodal Understanding with Audio: Gemma 3n can understand and process audio, text, and images, and offers significantly enhanced video understanding. Its audio capabilities enable the model to perform high-quality Automatic Speech Recognition (transcription) and Translation (speech to translated text). Additionally, the model accepts interleaved inputs across modalities, enabling understanding of complex multimodal interactions. (Public implementation coming soon)

Wow!!

Re: Gemma 3n preview: Mobile-first AI

#143

Earlier quoted context omitted.

I can't speak for anyone else, but these models only seem about as smart as google search, with enormous variability. I can't say I've ever had an interaction with a chatbot that's anything redolent of interaction with intelligence. Now would I take AI as a trivia partner? Absolutely. But that's not really the same as what I look for in "smart" humans.

>anything redolent of interaction with intelligence compared to what you are used to right? I know it's elitist but most people My circle of people I talk with during the day has changed since I took on more charity which consists of fixing up old laptops and installing Ubuntu on them; I get them for free from everyone and I give them to people who cannot afford, including some lessons and remote support (which is ea…

> but most people This is incorrect, IQ tests are normally scaled such that average intelligence is 100, and such that they are approximately normally distributed so that most people will be somewhere between 85-115 (66% on average).

Re: Gemma 3n preview: Mobile-first AI

#144

Earlier quoted context omitted.

I assume that "pretty fast" depends on the phone. My old Pixel 4a ran Gemma-3n-E2B-it-int4 without problems. Still, it took over 10 minutes to finish answering "What can you see?" when given an image from my recent photos. Final stats: 15.9 seconds to first token 16.4 tokens/second prefill speed 0.33 tokens/second decode speed 662 seconds to complete the answer

Gemma-3n-E4B-it on my 2022 Galaxy Z Fold 4. CPU: 7.37 seconds to first token 35.55 tokens/second prefill speed 7.09 tokens/second decode speed 27.97 seconds to complete the answer GPU: 1.96 seconds to first token 133.40 tokens/second prefill speed 7.95 tokens/second decode speed 14.80 seconds to complete the answer

So a apparently the NPU can't be used for models like this. I wonder what it is even good for.

Re: Gemma 3n preview: Mobile-first AI

#145
post #95
post #65

Earlier quoted context omitted.

LLMs neither understand nor reason, that has been shown multiple times.

Seems like some don’t like that LLMs aren’t really intelligent. https://neurosciencenews.com/llm-ai-logic-27987/

The study tested transformers, not LLMs.

They trained models on only task specific data, not on a general dataset and certainly not on the enormous datasets frontier models are trained on.

"Our training sets consist of 2.9M sequences (120M tokens) for shortest paths; 31M sequences (1.7B tokens) for noisy shortest paths; and 91M sequences (4.7B tokens) for random walks. We train two types of transformers [38] from scratch using next-token prediction for each dataset: an 89.3M parameter model consisting of 12 layers, 768 hidden dimensions, and 12 heads; and a 1.5B parameter model consisting of 48 layers, 1600 hidden dimensions, and 25 heads."

Re: Gemma 3n preview: Mobile-first AI

#146

Earlier quoted context omitted.

I did the same thing on my Pixel Fold. Tried two different images with two different prompts: "What can you see?" and "Describe this image" First image ('Describe', photo of my desk) - 15.6 seconds to first token - 2.6 tokens/second - Total 180 seconds Second image ('What can you see?', photo of a bowl of pasta) - 10.3 seconds to first token - 3.1 tokens/second - Total 26 seconds The Edge Gallery app defaults to CPU…

Pixel 4a release date = August 2020 Pixel Fold was in the Pixel 8 generation but uses the Tensor G2 from the 7s. Pixel 7 release date = October 2022 That's a 26 month difference, yet a full order of magnitude difference in token generation rate on the CPU. Who said Moore's Law is dead? ;)

As a another data point, on E4B, my Pixel 6 Pro (Tensor v1, Oct 2021) is getting about 4.4 t/s decode on a picture of a glass of milk, and over 6 t/s on text chat. It's amazing, I never dreamed I'd be viably running an 8 billion param model when I got it 4 years ago. And kudos to the Pixel team for including 12 GB of RAM when even today PC makers think they can get away with selling 8.

Re: Gemma 3n preview: Mobile-first AI

#147

You can try it on Android right now: Download the Edge Gallery apk from github: https://github.com/google-ai-edge/gallery/releases/tag/1.0.0 Download one of the .task files from huggingface: https://huggingface.co/collections/google/gemma-3n-preview-6... Import the .task file in Edge Gallery with the + bottom right. You can take pictures right from the app. The model is indeed pretty fast.

On Pixel 8a, I asked Gemma 3n to play 20 questions with me. It says it has an object in mind for me to guess then it asks me a question about it. And several attempts to clarify who is supposed to ask questions have gone in circles.

Re: Gemma 3n preview: Mobile-first AI

#148
post #70
post #15

Earlier quoted context omitted.

Given Apple's track record in dealing with the problem of ballooning app sizes, I'm not holding my breath. The incentives are just not aligned – Apple earns $$$ on each GB of extra storage users have to buy.

They earn from in-app purchases too!

Not anymore!

Only half joking. I really do think the majority of that revenue will be going away.

Post reply on HN