Live data from Hacker News

MiniGPT-4

minigpt-4.github.io

141–150 of 337 posts

Re: MiniGPT-4

#141

Earlier quoted context omitted.

Remember that "cameras" aren't as good as human perception because human eyes interact with the environment instead of being passive sensors. (That is, if you can't see something you can move your head.) Plus we have ears, are under a roof so can't get rained on, are self cleaning, temperature regulating, have much better dynamic range, wear driving glasses…

And we still get into millions of accidents every year…

I keep hearing this argument over and over, but I find it uncompelling. As a relatively young person with good vision, who has never been in an accident after many years of driving, and who doesn't make the kind of simple mistakes I've seen the absurd mistakes self-driving cars make and I would not trust my life to a self-driving car.

Asking people to accept a driverless car based on over-arching statistics is papering over some very glaring issues. For example, are most accidents in cars being caused by "average" drivers or are they young / old / intoxicated / distracted / bad vision? Are the statistics randomly distributed (e.g. any driver is just as likely as the next to get in accidents)? Because the driverless cars seem to have accidents at random in unpredictable ways, but human drivers can be excellent (no accidents, no tickets ever), or terrible (drive fast, tickets, high insurance, accidents, etc). The distribution of accidents among humans is not close to uniform, and is usually explainable. I wouldn't trust a poor human driver on a regular basis, nor would I trust an AI because I'm actually a much better driver than both (no tickets, no accidents, can handle complex situations the AI can't). Are the comparisons of human accidents being treated as homogenous (e.g. the chance of ramming full speed into a parked car the same as a fender-bender?). I see 5.8M car crashes anually, but deaths remain fairly low (~40k, .68%), vs 400 driverless accidents with ~20 deaths (5%), I'm not sure we're talking about the same type of accidents.

tl;dr papering over the complexity of driving and how good a portion of drivers might be by mixing non-homogenous groups of drivers and taking global statistics of all accidents and drivers to justify unreliable and relatively dangerous technology would be a strict downgrade for most good drivers (who are most of the population).

Re: MiniGPT-4

#142
post #115

Earlier quoted context omitted.

Indeed, really simple. And yes, the results are shockingly good. But what I find most remarkable about this is that the ViT-L+Q-former's hidden states are related by only a linear projection (plus bias) to the Vicuna-13B's token embeddings: emb_in_vicuna_space = emb_in_qformer_space @ W + B These two models are trained independently of each other, on very different data (RGB images vs integer token ids representing s…

I think it’s just that affine transforms in high dimensions are surprisingly expressive. Since the functions are sparsely defined they’re much less constrained compared to the low dimensional affine transformations we usually think of.

Good point. Didn't think of that. It's a plausible explanation here, because the dimensionality of the spaces is so different, 5120 vs 768. Not surprisingly, the trained weight matrix has rank 768: it's using every feature in the lower-dimensional space.

Still, it's kind of shocking that it works so well!

I'd be curious to see if the learned weight matrix ends up being full-rank (or close to full-rank) if both spaces have the same dimensionality.

Re: MiniGPT-4

#143
post #109

Earlier quoted context omitted.

> This ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy. Just like any discussion between advanced web devs would make any humble woodworker feel? And just like any discussion between advanced woodworkers would make a humble web dev feel? "It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a…

Web devs have become blue collar!? =P Great idea, actually. I do hope for a curriculum that enables kids on the trade school path to learn more about programming. Why not Master/Journeyman/Apprentice style learning for web dev??

That's kind of how I think about bootcamps pumping out web devs. They're like trade schools, teaching you just enough fundamentals to know how to use existing tools.

Re: MiniGPT-4

#144

From a radiology world this is fascinating. I'm not worried about job security as I'm an interventionalist. What I'm wondering is about go-to-market strategies for diagnostics. I do some diagnostic reads and I would love to have something like this pre-draft reports (especially for X-Rays). There are tons of "AI in rads" companies right now, none of which have models that come anywhere close to GPT-4 or even this. Pe…

Your profession and... a few hundred others ?

Re: MiniGPT-4

#145

Earlier quoted context omitted.

Remember that "cameras" aren't as good as human perception because human eyes interact with the environment instead of being passive sensors. (That is, if you can't see something you can move your head.) Plus we have ears, are under a roof so can't get rained on, are self cleaning, temperature regulating, have much better dynamic range, wear driving glasses…

And we still get into millions of accidents every year…

Which sounds like a lot until you realize 1) we drive over three trillion miles a year in the US, and 2) the majority of those accidents are concentrated to a fraction of all drivers. The median human driver is quite good, and the state of the art AI isn't even in the same galaxy yet.

Re: MiniGPT-4

#146
post #44

It's hard to distinguish non-Google projects with Google Sans in their templates from actual Google Research papers, as the font is meant to be exclusively used by Google[1]. [1] https://developers.google.com/fonts/faq#how_can_i_get_a_lice...

Surely most people would read the authors list to determine provenance rather than the font?

I didn't think about it consciously but I think I did implicitly assume it was a Google project because of the font

Re: MiniGPT-4

#147
post #109
post #14

Earlier quoted context omitted.

> they're doing something really simple -- take BLIP2's ViT-L+Q-former, connect it to Vicuna-13B with a linear layer, and train just the tiny layer on some datasets of image-text pairs Oh yes. Simple! Jesus, this ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy.

> This ML stuff makes a humble web dev like myself feel like a dog trying to read Tolstoy. Just like any discussion between advanced web devs would make any humble woodworker feel? And just like any discussion between advanced woodworkers would make a humble web dev feel? "It's really simple, they're just using a No. 7 jointer plane with a high-angle frog and a PM-V11 blade to flatten those curly birch boards, then a…

Hey, guys. Hey. Ready to talk plate processing and residue transport plate funneling? Why don't we start with joust jambs? Hey, why not? Plates and jousts. Can we couple them? Hell, yeah, we can. Want to know how? Get this. Proprietary to McMillan. Only us. Ready? We fit Donnely nut spacing grip grids and splay-flexed brace columns against beam-fastened derrick husk nuts and girdle plate Jerries, while plate flex tandems press task apparati of ten vertipin-plated pan traps at every maiden clamp plate packet. Knuckle couplers plate alternating sprams from the t-nut to the SKN to the chim line. Yeah. That is the McMillan way. And it's just another day at the office.

Re: MiniGPT-4

#148
post #106

Earlier quoted context omitted.

Sure, sure, but would it have killed them to drop in a few five dollar "don't hit this object" ultrasonic proximity sensors?

While ultrasonic sensors would be fine for parking, they don't have very good range so they aren't much help in avoiding, for example, crashing into stationary fire trucks or concrete lane dividers at freeway speeds.

Just disable autopilot 0.00001 seconds before impact and it becomes the driver's fault.

Re: MiniGPT-4

#149
post #26

Earlier quoted context omitted.

Someone needs to write a buyer's guide for GPUs and LLMs. For example, what's the best course of action if don't need to train anything but do want to eventually run whatever model becomes the first local-capable equivalent to ChatGPT? Do you go with Nvidia for the CUDA cores or with AMD for more VRAM? Do you do neither and wait another generation?

Nvidia and the highest amount of vram you can get. Currently the 4090, the rumor is the 4090ti will have 48gb of vram, idk if its worth waiting or not. The more VRAM the higher paremeter count you can run all in memory (fastest by far). AMD is almost a joke in ML. The lack of CUDA support (which is nvidia proprietary) is straight lethal, and also even though ROCM does have much better support these days, from what I'…

I think the most recent rumors were amended to it having 24, unfortunately.

Re: MiniGPT-4

#150
post #101

Earlier quoted context omitted.

For a general guide, I recommend: https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni... There's a subreddit r/LocalLLaMA that seems like the most active community focused on self-hosting LLMs. Here's a recent discussion on hardware: https://www.reddit.com/r/LocalLLaMA/comments/12lynw8/is_anyo... If you're looking just for local inference, you're best bet is probably to buy a consumer GPU w/ 24GB of RAM (309…

FWIW I had no real issues getting StableDiffusion to run on a 6800 I have in one of my systems. I haven't tried with LLaMA at all.

6800 is RDNA2, not RDNA3. The latter is still waiting for ROCm support 4 months post-launch: https://github.com/RadeonOpenCompute/ROCm/issues/1813
Post reply on HN