Live data from Hacker News

MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

arxiv.org

41–50 of 65 posts

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#41
post #5

This looks competitive against CLIP, and surprisingly great at VQA style prompts, but it doesn't seem like the paper supports comparing it to GPT-4. We don't see any tests for coding performance, math homework, legal document review, or any of the myriad other things that people use GPT-4 for on a daily basis.

Besides homework, all of these things seem to be professional uses of GPT-4. If they’re trying to bake this into a consumer platform like Siri, I don’t see why they’d need to focus on those use cases. Besides MDM/Enterprise, which will be curious if they try and attack this market or just their army of consumer devices.

They are going to have to focus on the use cases that most of their customers use LLMs for, regardless of whether it falls in the consumer or professional category or somewhere in between.

If all it does is improve Siri a bit without massively expanding the range of applications and APIs it will be a big disappointment.

I think what Apple presents in June will decide whether on-device AI will be seen as a viable alternative to cloud APIs.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#42
post #23

If it’s going to take general artificial intelligent to get a voice assistant that can remember not one, but two entirely separate cooking timers, then so be it. Imagine the GPUs required! I’m still baffled at Siri and Google assistant. Virtually zero innovation in a decade. I just want to be able to turn on BBC radio while my hands are wet, is that really so hard?!

You should be able to do this with Siri. You can use a shortcut if it doesn't work out of the box.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#43

Earlier quoted context omitted.

Out of curiosity, where are you seeing this? It's not in the abstract or the paper.

Some of these comments were originally made in response to this spammy submission: https://news.ycombinator.com/item?id=39726156

Oh, thank you! I didn't know we'd been moved.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#44

Earlier quoted context omitted.

> that can remember not one, but two entirely separate cooking timers You're in luck! Siri will do that right now. Just tried it. Works.

OMG, 2 cooking timers?! Pinnacle tech right there. Knowing Apple, I was expecting one base timer, with every other timer being a $200 upgrade.

That’s not really Apple’s style. More along the lines of “HomePod mini 2 features double the RAM, allowing for exciting new features like multiple kitchen timers. Pre-orders start Friday.”

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#45
post #4

The paper lists "first authors", "core authors", and "senior authors". My dream is to one day be listed on a seminal paper as "secondary forum reply author".

Speaking as someone working in the field, I find it amusing how much researchers working on automating human work care about human credit assignment.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#46
post #32

Biggest model is 30b MoE trained on 100b tokens, max sequence length 4096. A bit underwhelming compared to recent announcements like the open source Large World Model [1]. Absolutely no benchmarks against GPT4 present in the paper. Notably they used instruction response pairs generated from GPT4 for supervised fine tuning. Which has always felt like an experimental hack to me, but that’s how many folks are bootstrapp…

> Absolutely no benchmarks against GPT4 present in the paper. Table 4 on page 14 shows comparisons to GPT4V

[flagged]

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#47
post #4

The paper lists "first authors", "core authors", and "senior authors". My dream is to one day be listed on a seminal paper as "secondary forum reply author".

Similarly, I’d like the movie credit Second Assistant to the Second Second Assistant Director.

In that case, I highly recommend watching the movie Synecdoche New York (2008).

PS Can I be your hairdresser?

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#48

Earlier quoted context omitted.

Honest question: what do you (in the general sense, not specifically asking the parent) use Siri for? I think my main (only?) use case is setting a timer. Maybe I find conversational UIs awkward, or maybe I just got jaded REALLY quickly from Siri’s lacking capabilities early on, but I have hardly used it in the decade or whatever that it’s been around.

I use it almost daily for something that is simple but under appreciated I don’t know why it’s not in every marketing video: “Siri, remind me tomorrow at 10am to do X” I outsource so much of my memory to the phone via Siri ALL THE TIME. It’s so useful. Even for things in 20m. I’ll easily forget if I don’t do this, and it’s reliable so it gives me confidence. It also keeps the notification present until I actually do…

Yep. Reminders is #1 by far, followed by sending texts, turning lights on/off with HomeKit and timers which are similar.

I can’t imagine reminders w/o Siri because that’s how I add 90%+ of them. Grocery items, things to do at time X, or when I get to (or leave) work/home are the big ones.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#49

Earlier quoted context omitted.

I use it almost daily for something that is simple but under appreciated I don’t know why it’s not in every marketing video: “Siri, remind me tomorrow at 10am to do X” I outsource so much of my memory to the phone via Siri ALL THE TIME. It’s so useful. Even for things in 20m. I’ll easily forget if I don’t do this, and it’s reliable so it gives me confidence. It also keeps the notification present until I actually do…

Especially with Shortcuts, Siri can have some pretty useful functionality. My personal big improvement I'd like to see is being able to better able to tap into those actions without having to set things up in advance.

I’m really hoping for something like that.

A year or so ago I remember someone pointing out in a podcast how LLMs are great at taking something like general language and turning it into a series of predefined commands (the stuff available to shortcuts). It would instantly make Siri much more useful.

I think Federico Viticci rigged up something similar or at least a powerful demo using Siri + Shortcuts + ChatGPT to be able to answer all sort of questions better than native Siri.

Re: MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training

#50
post #23

If it’s going to take general artificial intelligent to get a voice assistant that can remember not one, but two entirely separate cooking timers, then so be it. Imagine the GPUs required! I’m still baffled at Siri and Google assistant. Virtually zero innovation in a decade. I just want to be able to turn on BBC radio while my hands are wet, is that really so hard?!

You mostly think there's no innovation because you speak English.
Post reply on HN