Live data from Hacker News

MiniGPT-4

minigpt-4.github.io

121–130 of 337 posts

Re: MiniGPT-4

#121
post #60
post #59

Earlier quoted context omitted.

Near term it’s a frustrating decision, but if these gpt4 vision LLMs are anything to go by it will prove to be the right decision in the long term.

Why wouldn’t LIDAR in addition to computer vision with cameras be a strictly better idea?

It's all trade offs. I'm just spitballing here, but if you have limited resources, you can either spend cash/time on lidar or invest in higher-quality mass-produced optics, or better computer vision software. If you get to a functional camera-only system sooner, might everyone be better off as you can deploy it more rapidly.

Manufacturing capacity of lidar components might be limited.

Another might be reliability/failure modes. If the system relies on lidar, that's another component that can break (or brownout and produce unreliable inputs).

So in a vaccum, yea a lidar+camera system is probably better, but who knows with real life trade offs.

(again, I just made these up, I do not work on this stuff, but these are a few scenarios I can imagine)

Re: MiniGPT-4

#122
Hate to be the person complaining about the name, but we already saw how this plays out with DALL-E mini: if you name your project directly after something else like this, no matter how much extra explanatory text you attach to it a large number of people will assume it's an "official" variant of the thing it was named after.

Eventually you'll have to rename it, either to resolve the confusion or because OpenAI pressure you to do so, or both.

So better to pick a less confusing name from the start.

(This one is even more confusing because it's about image inputs, but GPT4 with image inputs had not actually been released to anyone yet - similar in fact to how DALL-E mini got massive attention because DALL-E itself was still in closed preview)

Re: MiniGPT-4

#123
post #46

Earlier quoted context omitted.

It's pretty simple actually. Get a 3090 or 4090. Forget about AMD.

I have access to an Nvidia A100. But as a layman, what specs does the rest of the system need to use it for some real work? I would assume there needs to be at least as much ram as vram and maybe a few terabytes of disk space. Does anyone have experience with this?

If you have an A100, which in its 80GB variant costs $23,667 [1], you would not generally quibble over the price of a few terabytes of disk space.

[1] https://www.dell.com/en-us/shop/nvidia-ampere-a100-pcie-300w...

Re: MiniGPT-4

#124
post #109
post #14

Earlier quoted context omitted.

> they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.

> This ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy. Just like any discussion between advanced web devs would make any humble woodworker feel? And just like any discussion between advanced woodworkers would make a humble web dev feel? "It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a…

I love your analysis.

Re: MiniGPT-4

#125
Do I understand this correctly: they just took Blip2 and replaced the LLM with Vicuna, and to do that they just added a single linear layer to translate between frozen vision encoder and (frozen) Vicuna? Additionally, and importantly, they manually create a high quality dataset for finetuning their model.

If that is the case, then this is really a very, very simple paper. But I guess simple things can lead to great improvements, and indeed their results seem very impressive. Goes to show how much low hanging fruit there must be in deep learning these days by leveraging the amazing, and amazingly general, capabilities of LLMs.

Re: MiniGPT-4

#126

Earlier quoted context omitted.

Wow, he waits until halfway through the article to mention A New Kind of Science. Usually he works it into the first couple of paragraphs!

I known it’s hard to believe but I sense LLMs have slightly knocked his ego down and injected a small dose of humility. https://youtu.be/z5WZhCBRDpU I pick that up in above video and also in the post above. Definitely healthy for him which just to be clear I’m a huge Wolfram fan and the ego doesn’t really bother me, it’s just part of who he is, however I do find it nice that LLMs are having him self reflect more than…

Not a big Wolfram fan myself. I gave him the benefit of the doubt and bought "A New Kind of Science" (freakin' expensive when it first came out), and read the whole 1280 pages cover to cover ... Would have been better presented as a short blog post.

I find it funny how despite being completely uninvolved in ChatGPT he felt the need to inject himself into the conversation and write a book about it. I guess it's the sort of important stuff that he felt an important person like himself should be educating the plebes on.

Predictably he had no insight into it and will have left the plebes thinking it's something related to MNIST and cat-detection.

Re: MiniGPT-4

#127

Earlier quoted context omitted.

Windows generally works but there may be a somewhat small performance hit. IMO linux is much easier to get to work judging by all the github issue threads I see able SD/LLaMa stuff on windows - but I don't use windows so I dont have personal experience. 4090 24GB is 1800USD, The Ada A6000 48GB is like 8000USD and idk where you buy it? So if you want to run games and models locally the 4090 is honestly the best option…

The A6000 is actually the old generation, Ampere. The new Ada generation one is called 6000. Seems many places still sell A6000 (Ampere) for the same price as RTX 6000 (Ada) though, even though the new one is twice as fast. Seems you can get used RTX A6000s for around $3000 on ebay.

That.... That explains why I can't find it and makes a ton of sense.....

I think that's such a silly name for it, but oh well

Thanks for the correction!

Re: MiniGPT-4

#128
post #106

Earlier quoted context omitted.

That's because the decided they do not need lidar.

Sure, sure, but would it have killed them to drop in a few five dollar "don't hit this object" ultrasonic proximity sensors?

While ultrasonic sensors would be fine for parking, they don't have very good range so they aren't much help in avoiding, for example, crashing into stationary fire trucks or concrete lane dividers at freeway speeds.

Re: MiniGPT-4

#129
post #101
post #26

Earlier quoted context omitted.

Someone needs to write a buyer's guide for GPUs and LLMs. For example, what's the best course of action if don't need to train anything but do want to eventually run whatever model becomes the first local-capable equivalent to ChatGPT? Do you go with Nvidia for the CUDA cores or with AMD for more VRAM? Do you do neither and wait another generation?

For a general guide, I recommend: https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni... There's a subreddit r/LocalLLaMA that seems like the most active community focused on self-hosting LLMs. Here's a recent discussion on hardware: https://www.reddit.com/r/LocalLLaMA/comments/12lynw8/is_anyo... If you're looking just for local inference, you're best bet is probably to buy a consumer GPU w/ 24GB of RAM (309…

FWIW I had no real issues getting StableDiffusion to run on a 6800 I have in one of my systems.

I haven't tried with LLaMA at all.

Re: MiniGPT-4

#130
post #109
post #14

Earlier quoted context omitted.

> they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.

> This ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy. Just like any discussion between advanced web devs would make any humble woodworker feel? And just like any discussion between advanced woodworkers would make a humble web dev feel? "It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a…

Web devs have become blue collar!? =P

Great idea, actually. I do hope for a curriculum that enables kids on the trade school path to learn more about programming. Why not Master/Journeyman/Apprentice style learning for web dev??

Post reply on HN