Live data from Hacker News

MiniGPT-4

minigpt-4.github.io

321–330 of 337 posts

Re: MiniGPT-4

#326
post #115
post #6

On a technical level, they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs. But the results are pretty amazing. It completely knocks Openflamingo && even the original blip2 models out of the park. And best of all, it arrived before OpenAI's GPT-4 Image Modality did. Real win for Open Sourc…

Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…

Man you need to look at this - https://llava-vl.github.io/. They project with a linear layer from Clip directly. With blip-2, you could say it already converts RGB into token space.

Re: MiniGPT-4

#327
post #115

Earlier quoted context omitted.

Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…

BLIP2 is a contrastive Image-Language model. The embeddings from the BLIP2 image model are already both aligned with text, and linear. It should not be a surprise that only a projection is required to translate it to LLaMA's embedding space.

apparently you can project directly with CLIP. See here - https://llava-vl.github.io/. This seems pretty wild to me.

Re: MiniGPT-4

#328

Earlier quoted context omitted.

Its poor form to be calling it 'Open' AI. So I guess its swings and roundabouts. Everyone is leeching where they can.

To be fair they were open when that name was picked and it looks like they may be trying to transition to just 'ai.com'.

Save some money with closedai.com

Re: MiniGPT-4

#329
post #6

On a technical level, they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs. But the results are pretty amazing. It completely knocks Openflamingo && even the original blip2 models out of the park. And best of all, it arrived before OpenAI's GPT-4 Image Modality did. Real win for Open Sourc…

>so hopefully by tomorrow it'll be runnable by 3090/4090 users. Taking a step back, this is just a wild statement. I know there's some doom and gloom out there, but in certain aspects, it's an awesome time to be alive.

I've never seen anything quite like it.

Re: MiniGPT-4

#330

Earlier quoted context omitted.

Hey, guys. Hey. Ready to talk plate processing and residue transport plate funneling? Why don't we start with joust jambs? Hey, why not? Plates and jousts. Can we couple them? Hell, yeah, we can. Want to know how? Get this. Proprietary to McMillan. Only us. Ready? We fit Donnely nut spacing grip grids and splay-flexed brace columns against beam-fastened derrick husk nuts and girdle plate Jerries, while plate flex tan…

Just tell me do we need a turbo encabulator or not?

I'll take 2
Post reply on HN