Maybe the abstract of the paper is a better introduction to what this is: > We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates d…
Ferret: A Multimodal Large Language Model
11–20 of 332 posts
Re: Ferret: A Multimodal Large Language Model
#12> Usage and License Notices: The data, and code is intended and licensed for research use only. They are also restricted to uses that follow the license agreement of LLaMA, Vicuna and GPT-4. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes. Wait, how did "GPT-4" get in there?
Re: Ferret: A Multimodal Large Language Model
#13We're watching Apple fill the moat in.
Re: Ferret: A Multimodal Large Language Model
#14We're watching Apple fill the moat in.
Re: Ferret: A Multimodal Large Language Model
#15Re: Ferret: A Multimodal Large Language Model
#16Re: Ferret: A Multimodal Large Language Model
#17We're watching Apple fill the moat in.
How so?
Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps.
No need for web services like ChatGPT.
Re: Ferret: A Multimodal Large Language Model
#18We're watching Apple fill the moat in.
Re: Ferret: A Multimodal Large Language Model
#19Earlier quoted context omitted.
How so?
Running Multimodal LLMs on device and offline, i.e LLMKit for free equaling GPT-3.5 / 4 then Google will follow on Android. Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps. No need for web services like ChatGPT.