Live data from Hacker News

Ferret: A Multimodal Large Language Model

github.com

11–20 of 332 posts

Re: Ferret: A Multimodal Large Language Model

#11

Maybe the abstract of the paper is a better introduction to what this is: > We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates d…

[deleted]

Re: Ferret: A Multimodal Large Language Model

#12

> Usage and License Notices: The data, and code is intended and licensed for research use only. They are also restricted to uses that follow the license agreement of LLaMA, Vicuna and GPT-4. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes. Wait, how did "GPT-4" get in there?

[deleted]

Re: Ferret: A Multimodal Large Language Model

#17
post #5

We're watching Apple fill the moat in.

How so?

Running Multimodal LLMs on device and offline, i.e LLMKit for free equaling GPT-3.5 / 4 then Google will follow on Android.

Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps.

No need for web services like ChatGPT.

Re: Ferret: A Multimodal Large Language Model

#19
post #5

Earlier quoted context omitted.

How so?

Running Multimodal LLMs on device and offline, i.e LLMKit for free equaling GPT-3.5 / 4 then Google will follow on Android. Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps. No need for web services like ChatGPT.

So Apple is filling in ChatGPT's moat then, not their own? Pardon my confusion
Post reply on HN