Ferret: A Multimodal Large Language Model
1–10 of 332 posts
Re: Ferret: A Multimodal Large Language Model
#2Re: Ferret: A Multimodal Large Language Model
#3Re: Ferret: A Multimodal Large Language Model
#4Can someone define the term “MLLM”?
Re: Ferret: A Multimodal Large Language Model
#5We're watching Apple fill the moat in.
Re: Ferret: A Multimodal Large Language Model
#6Re: Ferret: A Multimodal Large Language Model
#7Can someone define the term “MLLM”?
Re: Ferret: A Multimodal Large Language Model
#8> We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with 95K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination.
Re: Ferret: A Multimodal Large Language Model
#9Wait, how did "GPT-4" get in there?
Re: Ferret: A Multimodal Large Language Model
#10> Usage and License Notices: The data, and code is intended and licensed for research use only. They are also restricted to uses that follow the license agreement of LLaMA, Vicuna and GPT-4. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes. Wait, how did "GPT-4" get in there?