PaLI-3 Vision Language Models
11–20 of 29 posts
Re: PaLI-3 Vision Language Models
#12Re: PaLI-3 Vision Language Models
#13Re: PaLI-3 Vision Language Models
#14Re: PaLI-3 Vision Language Models
#15Re: PaLI-3 Vision Language Models
#16Even undigitized materials aren't safe any more.
Re: PaLI-3 Vision Language Models
#17Something that stood out to me skimming the paper - that was somewhat buried - they finetune the model on each benchmark. "Finally, for each individual task (benchmark), we fine-tune the PaLI-3 model with frozen ViT image encoder on the task’s training data as described in the cor- responding section. For most tasks, we fine-tune the 812×812 resolution checkpoint, but for two document understanding tasks, we go up to…
Re: PaLI-3 Vision Language Models
#18It's getting really awkward seeing these papers from Google. "We're here too! We're totally not woefully behind everyone else in the field!". No model, no reasonable comparisons, just generic bragging.
I'm astounded as an ML researcher how Google can be doing so incredibly badly. They have access to unlimited compute, good people, and great infrastructure. Yet something about their internal culture means they are unable to compete with OpenAI, Facebook, and even the open source community. They constantly brag about how good their models are (even in private) and then every time they deploy anything its performance is pathetic (like Bard and Bard with vision).
You can tell why Google recently totally overhauled the leadership of Google Research/Deep Mind and shut down Google Brain.
Re: PaLI-3 Vision Language Models
#19Does the vision-language-model process raw image data, or does it process OCR character output?
Re: PaLI-3 Vision Language Models
#20No comparison against GPT-4V? How embarrassing! Where are they going to submit this? A conference where no one knows about GPT-4V? Ridiculous. It's getting really awkward seeing these papers from Google. "We're here too! We're totally not woefully behind everyone else in the field!". No model, no reasonable comparisons, just generic bragging. I'm astounded as an ML researcher how Google can be doing so incredibly bad…