MiniGPT-4
321–330 of 337 posts
Re: MiniGPT-4
#322Re: MiniGPT-4
#323Re: MiniGPT-4
#324Re: MiniGPT-4
#325Re: MiniGPT-4
#326On a technical level, they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs. But the results are pretty amazing. It completely knocks Openflamingo && even the original blip2 models out of the park. And best of all, it arrived before OpenAI's GPT-4 Image Modality did. Real win for Open Sourc…
Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…
Re: MiniGPT-4
#327Earlier quoted context omitted.
Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…
BLIP2 is a contrastive Image-Language model. The embeddings from the BLIP2 image model are already both aligned with text, and linear. It should not be a surprise that only a projection is required to translate it to LLaMA's embedding space.
Re: MiniGPT-4
#328Earlier quoted context omitted.
Its poor form to be calling it 'Open' AI. So I guess its swings and roundabouts. Everyone is leeching where they can.
To be fair they were open when that name was picked and it looks like they may be trying to transition to just 'ai.com'.
Re: MiniGPT-4
#329On a technical level, they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs. But the results are pretty amazing. It completely knocks Openflamingo && even the original blip2 models out of the park. And best of all, it arrived before OpenAI's GPT-4 Image Modality did. Real win for Open Sourc…
>so hopefully by tomorrow it'll be runnable by 3090/4090 users. Taking a step back, this is just a wild statement. I know there's some doom and gloom out there, but in certain aspects, it's an awesome time to be alive.
Re: MiniGPT-4
#330Earlier quoted context omitted.
Hey, guys. Hey. Ready to talk plate processing and residue transport plate funneling? Why don't we start with joust jambs? Hey, why not? Plates and jousts. Can we couple them? Hell, yeah, we can. Want to know how? Get this. Proprietary to McMillan. Only us. Ready? We fit Donnely nut spacing grip grids and splay-flexed brace columns against beam-fastened derrick husk nuts and girdle plate Jerries, while plate flex tan…
Just tell me do we need a turbo encabulator or not?