Live data from Hacker News

Ferret: A Multimodal Large Language Model

github.com

251–260 of 332 posts

Re: Ferret: A Multimodal Large Language Model

#251
post #187

Earlier quoted context omitted.

Spotify / Apple Music Netflix + Hulu / Apple TV+ Generic Earbuds / AirPods Meta Quest / Apple Vision Pro (The last one being a hopeful wish)

? AppleTV and Apple Music are not decades ahead of anything. AirPods are way better than the existing Bluetooth headsets that were on the market.

Maybe they aren’t “decades ahead” but they’re examples of Apple being slow to market, launching very competitive products years after the market was established.

E.g. assuming Apple Vision launches soon, they’ll be “many years behind” Quest from the date of first launch, but most likely miles ahead as far as usability.

Re: Ferret: A Multimodal Large Language Model

#252

They're already going multi-modal? Holy crap, if google can't deliver in the accessibility space for this (image descriptions better than "the logo for the company"), then I'll definitely go back to Apple. I mean I do hope Apple cleans out bugs and makes VoiceOver feel like it won't fall over if I breathed hard, but their image descriptions, even without an LLM, are already clean and clear. More like "A green logo on…

Honestly if they are coming out with a paper now Apple has probably been working on it for a year or two at minimum . Next year releases of macOS / iOS are rumored to have LLMs as a feature .

> Honestly if they are coming out with a paper now Apple has probably been working on it for a year or two at minimum

Why do you say that?

Re: Ferret: A Multimodal Large Language Model

#253

Earlier quoted context omitted.

Can you give an example? I switched to android because i use personal assistant a lot while driving and siri was absolutely horrible.

- FaceID - Facial recognition in Photos - "Memories" in Photos - iOS keyboard autocomplete using LLMs. I am bilingual and noticed in the latest iOS it now does multi-language autocomplete and you no longer have to manually switch languages. - Event detection for Calendar - Depth Fusion in the iOS camera app, using ML to take crisper photos - Probably others... The crazy thing is most/all of these run on the device.

To be honest I distrust Microsoft with swift key, but the it recognizes the change in language just smooth. I could switch languages in one sentence and it would understand what I am writing just fine, no Sill suggestions

Re: Ferret: A Multimodal Large Language Model

#254

Earlier quoted context omitted.

Transformer based autocomplete on iOS 17 feels just as bad -- but in different ways -- as its previous incarnation to me.

Are you tapping the keys or swiping over those that make up the word you want to type? In my experience, tapping has always been and remained poor but swiping is getting better and better with every iOS version.

Swiping through keys doesn't have anything to do with autocomplete. Autocomplete has to do with predicting which word you're going to type next, not guessing which word best corresponds to the swipe you just made.

Re: Ferret: A Multimodal Large Language Model

#255

Earlier quoted context omitted.

Can you give an example? I switched to android because i use personal assistant a lot while driving and siri was absolutely horrible.

- FaceID - Facial recognition in Photos - "Memories" in Photos - iOS keyboard autocomplete using LLMs. I am bilingual and noticed in the latest iOS it now does multi-language autocomplete and you no longer have to manually switch languages. - Event detection for Calendar - Depth Fusion in the iOS camera app, using ML to take crisper photos - Probably others... The crazy thing is most/all of these run on the device.

The iPhone's built in text OCR and image subject cutouts are also extremely good, just in the photos app.

Re: Ferret: A Multimodal Large Language Model

#256
post #255

Earlier quoted context omitted.

- FaceID - Facial recognition in Photos - "Memories" in Photos - iOS keyboard autocomplete using LLMs. I am bilingual and noticed in the latest iOS it now does multi-language autocomplete and you no longer have to manually switch languages. - Event detection for Calendar - Depth Fusion in the iOS camera app, using ML to take crisper photos - Probably others... The crazy thing is most/all of these run on the device.

The iPhone's built in text OCR and image subject cutouts are also extremely good, just in the photos app.

Yeah totally, I copy text from images all the time.

Re: Ferret: A Multimodal Large Language Model

#257

Earlier quoted context omitted.

Given Apple's track record on anything AI related and the terrible state they keep CoreML that not only seems extraordinarily unlikely, it would take a lot of time to win developer trust and that I just don't see happening.

I have enjoyed working with CoreML over the last few years. Please share what you didn’t like about it.

There are so many modern ML components that have terrible or no support in CoreML. Try to do any convolution other than conv2d, advanced or custom activation functions etc and you are out of luck. Exporting from PyTorch leads to all sorts of headaches with subtle behavior changes between implementations it is definitely a pain point for developers of widely used software

Re: Ferret: A Multimodal Large Language Model

#258
post #246

Earlier quoted context omitted.

Given Apple's track record on anything AI related and the terrible state they keep CoreML that not only seems extraordinarily unlikely, it would take a lot of time to win developer trust and that I just don't see happening.

Apple doesn’t have to win developer trust or build an AI platform. They just have to build a compelling consumer product that can only function with AI, and they are better equipped to do that than Google or Microsoft. It remains to be seen if OpenAI will go that route instead of a business built on training and providing access to foundational models.

>They just have to build a compelling consumer product that can only function with AI

Yeah and I'm not talking exclusively about developer trust. Given Apple's current consumer lineup (see Siri, Apple photos, predictive text etc)... we only have evidence that they suck at ML. What makes you think they are going to suddenly transform overnight?

Re: Ferret: A Multimodal Large Language Model

#259

Earlier quoted context omitted.

From Bard: My situation is a bit unique, so the term "manufacturer" might not be the most accurate way to describe who created me. Here's a breakdown of what you need to know: Developed by Google AI: I was created by a team of researchers and engineers at Google AI, specializing in language models and artificial intelligence. Trained on a massive dataset: My knowledge and abilities come from being trained on a massiv…

Why was this downvoted? It didn't answer the question, but it showed that there is a sort of imprint that GP was asking about. And it saves everyone a tab's worth of effort.

I guess I didn't make it clear enough that the response was an answer to the example question in the comment I responded to, verbatim. I thought it interesting. Guess others didn't.

Re: Ferret: A Multimodal Large Language Model

#260

Earlier quoted context omitted.

The point of this thread is that even though Apple makes silly charts saying how good their hardware is at ML, they use products that their silly charts say aren't as good. There is no such hypocrisy if they use Samsung refrigerators.

ML inference and training are not the same task.

This plot is about general GPU performance, not pure inference. https://www.apple.com/newsroom/2022/03/apple-unveils-m1-ultr...

Training the model requires inference for forward propagation, so even then, for your comment to be relevant, you'd need to find a plot that Apple uses to compare inference on quantized models versus Nvidia, which doesn't exist.

Post reply on HN