Live data from Hacker News

MiniGPT-4

minigpt-4.github.io

221–230 of 337 posts

Re: MiniGPT-4

#221
post #115

Earlier quoted context omitted.

Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…

BLIP2 is a contrastive Image-Language model. The embeddings from the BLIP2 image model are already both aligned with text, and linear. It should not be a surprise that only a projection is required to translate it to LLaMA's embedding space.

This is the best answer. It makes sense to me. Thank you :-)

Re: MiniGPT-4

#222
post #183

Earlier quoted context omitted.

They would have full-rank because all the embedding space is used. There are no unused large pockets.

The weight matrix's rank would decrease for each feature in the target space that cannot be expressed as as a linear combination of features in the input space (plus a bias). For example, if the target space has a feature representing a non-visual quality like "smelliness," it would not be expressible as a linear combination of features representing visual attributes like "redness," "blueness," and "greenness," etc.…

A random nxn matrix is full rank... So it's kinda the default: any amount of noise in the embedding is going to result in full-rank transformations.

So it's really less-than-full rank which would require an explanation - ie, why does this image representation project into this perfectly isolated subspace of the language representation (or vice versa)?

If that happened I would start looking for things like a vocabulary of smell which is completely distinct and non-overlapping with any visual context. But we use cross-modal analogies in language /constantly/ (many smells are associated with things we can see - 'smells like a rose') so you wouldn't expect any clean separations for different modalities... Maybe there's some branch of analytic philosophy which has managed to completely divorce itself from the physical world...

Re: MiniGPT-4

#223
post #76

Earlier quoted context omitted.

It's more like gardening: 1. plant seed 2. ...wait a very long time... 3. observe completely unexpected but cool result The unexpected part of step 3 is what makes this very different from any kind of engineering, even webdev. Of course, there is a lot of engineering involved in good ML, but that is more comparable to agricultural engineering in the sense that it's just a lot of dumb plumbing that any engineer can do…

I mean, for me, the unexpected part of 3 is what got me into programming in general. The first time you type a mysterious incantation into an editor and a few more mysterious incantations into the console and the console prints "Hello, world" like it was supposed to, it's unexpected because it's hard to believe that any of this mysterious incantation stuff actually works at all. As you get better at programming you h…

I like how you've expressed this insight, and it is so true.

Becoming great at a particular technology stack means modelling it in great detail in your head, so you can move through it without external assistance. But that leaves an arena without discovery, where you just reinforce the same synapses, leading to rigidity and an absence of awe.

Re: MiniGPT-4

#224

Earlier quoted context omitted.

That.... That explains why I can't find it and makes a ton of sense..... I think that's such a silly name for it, but oh well Thanks for the correction!

Just to add to the confusion, there's another older RTX 6000 with 24GB ram. This is from an even older generation, same as the GeForce 20 series.

You're kidding? So they called it the RTX 6000, then called it the RTX A6000 for ampere, then back to RTX 6000 for Ada?

Why do they do this? Sometimes consumer products are versioned weirdly to mislead customers (like intel cpus) - but these wouldn't even make sense to do that with as they're enterprise cards?

Re: MiniGPT-4

#225

Earlier quoted context omitted.

Nvidia and the highest amount of vram you can get. Currently the 4090, the rumor is the 4090ti will have 48gb of vram, idk if its worth waiting or not. The more VRAM the higher paremeter count you can run all in memory (fastest by far). AMD is almost a joke in ML. The lack of CUDA support (which is nvidia proprietary) is straight lethal, and also even though ROCM does have much better support these days, from what I'…

I think the most recent rumors were amended to it having 24, unfortunately.

Darn.

I mean in all honestly there's no reason a gaming card would need 48gb at the moment when so few games even use 24gb.

48GB really only makes sense for workstation cards.

Re: MiniGPT-4

#226
post #102

Earlier quoted context omitted.

Windows generally works but there may be a somewhat small performance hit. IMO linux is much easier to get to work judging by all the github issue threads I see able SD/LLaMa stuff on windows - but I don't use windows so I dont have personal experience. 4090 24GB is 1800USD, The Ada A6000 48GB is like 8000USD and idk where you buy it? So if you want to run games and models locally the 4090 is honestly the best option…

If I was going to spend $8000 on a video card I’d hunt on eBay for an A100 80GB rather than settle for the A6000

Honestly yeah a used A100 80GB sounds like a better idea.

Re: MiniGPT-4

#227

Earlier quoted context omitted.

Hard disagree. Outside of the brand name ChatGPT, lay members of the general public are way more likely to call these chatbots (like Bard and Bing) “AIs” than “GPTs”. And although GPT could technically refer to any model that uses a Generative Pre-trained Transformer approach (although it probably wouldn’t be an open-and-shut case), the mark “GPT-4” definitely is associated with OpenAI and their product, and you can’…

So OpenAI ostensibly owns "GPT4" according to your argument. But does it own "MiniGPT4"? I hope you see the absurdity of this. Let's not discuss the amount of copyright licenses OpenAI has already infringed, too

I’ll put it this way:

At Brewer’s Art in Baltimore, MD they just released a beer called GPT (Green Peppercorn Tripel)[1]. They’re likely allowed to do that because a reasonable consumer would probably not actually think they had collaborated with OpenAI, because OpenAI does not make beer.

OP is releasing a model called “MiniGPT-4”. A reasonable consumer could look at that name and become confused about the origin of the product, thinking it was from OpenAI. This would be understandable, since OpenAI also makes large language models and has a well known one that they’ve been promoting whose brand name is “GPT-4”. If MiniGPT-4 does not meet that consumer’s expectation of quality (which has been built up through using and hearing about GPT-4) it may cause them to think something like “Wow, I guess OpenAI is going downhill”.

Trademark cases are generally decided on a “reasonable consumer” basis. So yeah, they can seem a little arbitrary. But it’s important for consumers to be able to distinguish the origin of the goods they are consuming and for creators to be able to benefit from their investment in advertising and product development.

[1] https://www.thebrewersart.com/bottles-cans

Re: MiniGPT-4

#228
post #3

Giving this a quick spin and I like what I'm seeing. I gave it a picture of Paolo Veronese's Lament over the Dead Christ [0] and asked what was going on in the background. > The painting depicts the scene of the crucifixion of Jesus Christ. In the foreground, there is a group of people standing around the cross, including Mary, the mother of Jesus, and the two thieves who were crucified with him. In the background, t…

How about asking it to describe a Robert Williams painting? Here's a relatively easy one: [1] [1] - http://www.illustrationartgallery.com/acatalog/WilliamsTimmy...

Used this prompt

Could you describe the image as thoroughly as possible? Please include the content, any noteworthy symbology, techniques or styles, what kind of media, guess at the artist or identify any attribution, then attempt to describe the theme or idea it represents?

Got this

The image is a painting of a child running down a dirt road with a kite in hand. The child is wearing a yellow shirt and blue shorts. The background is a mix of urban and rural elements, with buildings and trees in the background and a field of grass in the foreground. The style of the painting is a mix of realism and abstraction, with a focus on the child and the kite. The artist may be attempting to convey a sense of freedom and playfulness, with the child running with the kite representing the joy and wonder of childhood.

The medium used is oil paint on canvas. The artist may have used brushes and other tools to create the different textures and effects in the painting. The artist may have also used underpainting to establish the composition and values before adding the final layers of color and detail.

It is difficult to identify the artist without more information, but the style and technique of the painting suggest that it may be the work of an artist who is skilled in realism and abstraction, and who is interested in exploring themes of childhood and playfulness.

Re: MiniGPT-4

#229
post #109

Earlier quoted context omitted.

> This ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy. Just like any discussion between advanced web devs would make any humble woodworker feel? And just like any discussion between advanced woodworkers would make a humble web dev feel? "It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a…

Hey, guys. Hey. Ready to talk plate processing and residue transport plate funneling? Why don't we start with joust jambs? Hey, why not? Plates and jousts. Can we couple them? Hell, yeah, we can. Want to know how? Get this. Proprietary to McMillan. Only us. Ready? We fit Donnely nut spacing grip grids and splay-flexed brace columns against beam-fastened derrick husk nuts and girdle plate Jerries, while plate flex tan…

This post is double great and I will never forgive Amazon for canceling that show.

For those that don't know this is from a show called Patriot.

https://en.wikipedia.org/wiki/Patriot_(TV_series)

Scene: https://youtube.com/watch?v=-F-IHvF5OCA

Re: MiniGPT-4

#230
post #150

Earlier quoted context omitted.

FWIW I had no real issues getting StableDiffusion to run on a 6800 I have in one of my systems. I haven't tried with LLaMA at all.

6800 is RDNA2, not RDNA3. The latter is still waiting for ROCm support 4 months post-launch: https://github.com/RadeonOpenCompute/ROCm/issues/1813

I'm aware that a 6800 is not RDNA3. You stated broadly:

> Current AMD consumer cards have terrible software support and IMO isn't really an option. On Windows you might be able to use SHARK or DirectML ports, but nothing will run out of the box.

I was merely sharing that I did not have that same experience that current consumer cards have terrible support.

Post reply on HN