Note that a video is just a sequence of images: OpenAI has a demo with GPT-4-Vision that sends a list of frames to the model with a similar effect: https://cookbook.openai.com/examples/gpt_with_vision_for_vid... If GPT-4-Vision supported function calling/structured data for guaranteed JSON output, that would be nice though. There's shenanigans you can do with ffmpeg to output every-other-frame to halve the costs too.…
The number of tokens used for videos - 1,841 for my 7s video, 6,049 for 22s - suggests to me that this is a much more efficient way of processing content than individual frames. For structured data extraction I also like not having to run pseudo-OCR on hundreds of frames and then combine the results myself.
The killer app of Gemini Pro 1.5 is using video as an input
471–480 of 507 posts
Re: The killer app of Gemini Pro 1.5 is using video as an input
#472At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…
The whole matter of tokens from video is one that has a lot of ambiguity, and is often presented as if these are some unique weird encoding of the contents of the video. But logically the only possible tokenization of videos (or images, or series of images ala video) is basically an image to text model that takes each frame and generates descriptive language -- in English in Gemini -- to describe the contents of the…
From the Gemini report https://arxiv.org/abs/2312.11805
>The visual encoding of Gemini models is inspired by our own foundational work on Flamingo (Alayrac et al., 2022), CoCa (Yu et al., 2022a), and PaLI (Chen et al.,2022), with the important distinction that the models are multimodal from the beginning and can natively output images using discrete image tokens (Ramesh et al., 2021; Yu et al., 2022b).
These are the papers Google say the multimodality in Gemini is based on.
Flamingo - https://arxiv.org/abs/2204.14198
Pali - https://arxiv.org/abs/2209.06794
The images are encoded. The encoding process tokenizes the images and the transformer is trained to predict text with both the text and image encodings.
There is no conversion to text for Gemini. That's not where the token number comes from.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#473Earlier quoted context omitted.
Your website and blog are very low on details on how this is working. Downloading and installing an mai directly feels unsafe imo. Especially when I don't know how this software is working. Is it recording a video, performing OCR continuously, taking just screenshots No mention of using any LLMs in there at all which is how you are presenting it in your comment here.
Feedback taken. I'll add more details on how this works for us technical people. LLM integration is in progress and coming soon. Any idea what would make you feel safe? 3rd party verification? I had it verified and published by the Microsoft Store. I feel eventually it all comes down to me being a decent person.
* In-depth technical explanation with architecture diagrams
* Open-source and self-hosted version
Also I didn't understand if it talks to a remote server or not. Because that's a big blocker for me.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#474Earlier quoted context omitted.
The whole matter of tokens from video is one that has a lot of ambiguity, and is often presented as if these are some unique weird encoding of the contents of the video. But logically the only possible tokenization of videos (or images, or series of images ala video) is basically an image to text model that takes each frame and generates descriptive language -- in English in Gemini -- to describe the contents of the…
This explanation is wrong as I've already said (256 is not the result of any conversion to text) but no one has to take my word for it. From the Gemini report https://arxiv.org/abs/2312.11805 >The visual encoding of Gemini models is inspired by our own foundational work on Flamingo (Alayrac et al., 2022), CoCa (Yu et al., 2022a), and PaLI (Chen et al.,2022), with the important distinction that the models are multimod…
As much as I would love to waste my time replying again to your magic thinking, instead I'll just politely chuckle and move on. Good luck.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#475Re: The killer app of Gemini Pro 1.5 is using video as an input
#476Re: The killer app of Gemini Pro 1.5 is using video as an input
#477Earlier quoted context omitted.
This explanation is wrong as I've already said (256 is not the result of any conversion to text) but no one has to take my word for it. From the Gemini report https://arxiv.org/abs/2312.11805 >The visual encoding of Gemini models is inspired by our own foundational work on Flamingo (Alayrac et al., 2022), CoCa (Yu et al., 2022a), and PaLI (Chen et al.,2022), with the important distinction that the models are multimod…
Stewing so much you had to double-dip reply? Ouch. As much as I would love to waste my time replying again to your magic thinking, instead I'll just politely chuckle and move on. Good luck.
You have your head so far up your ass even direct confirmation from the model builders themselves won't sway you. The comment wasn't for you. The comment is linked sources for the original poster and for the curious.
You see I don't have to hide behind a veneer of "Trust me bro. It works like this".
Re: The killer app of Gemini Pro 1.5 is using video as an input
#478Earlier quoted context omitted.
Nobody would complain on HN if Google Gemini was generating pictures of Lincoln existing as a... gasp ... white person. This absurd level of woke censorship is not doing them any good.
People would 1000% complain if a search for “skilled scientist” was 90% white men, even if that tracked completely true to statistical reality.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#479Earlier quoted context omitted.
https://en.wikipedia.org/wiki/Incitement_to_ethnic_or_racial...
There are 6 countries listed in that article, out of the nearly 200 countries in the world. Hardly "most." And there doesn't appear to be examples of those 6 countries imprisoning people for those laws.
[1]: https://www.reddit.com/r/MapPorn/comments/qh7ua1/hate_speech...
[2]: https://www.nytimes.com/2022/09/23/technology/germany-intern...
Re: The killer app of Gemini Pro 1.5 is using video as an input
#480Earlier quoted context omitted.
Stewing so much you had to double-dip reply? Ouch. As much as I would love to waste my time replying again to your magic thinking, instead I'll just politely chuckle and move on. Good luck.
>As much as I would love to waste my time replying again to your nonsense, instead I'll just politely chuckle and move on. Good luck. You have your head so far up your ass even direct confirmation from the model builders themselves won't sway you. The comment wasn't for you. The comment is linked sources for the original poster and for the curious. You see I don't have to hide behind a veneer of "Trust me bro. It wor…
Linking papers that you clearly haven't read and can't contextually apply -- as with the ViT or your misunderstanding of image tiling -- is not the sound strategy you hope it is. It doesn't confirm your claims.
I'm not asking anyone to "Trust me bro". So...have you called the Gemini Pro 1.5 API and tokenized an image or a video yet?
There is a certain element of this that is just spectacularly obvious to anyone who spent even a moment of critical thought -- if they're so capable -- on it. Your claim is that a high resolution image is tiled to a 16x16 array...and the magic model can at some later point magically on demand extract any and all details, such as OCR, from that 16x16. This betrays a fundamental ignorance of even the most basic of information theory.
Again, I would love to just block you and avoid the defensive insults you keep hurling, but this site lacks the ability. Stop replying to me, however many more contextually nonsensical citations you think will save face. Thanks.