Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

451–460 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#451

Earlier quoted context omitted.

The whole matter of tokens from video is one that has a lot of ambiguity, and is often presented as if these are some unique weird encoding of the contents of the video. But logically the only possible tokenization of videos (or images, or series of images ala video) is basically an image to text model that takes each frame and generates descriptive language -- in English in Gemini -- to describe the contents of the…

This is not at all how this works. There's no separate model. Yes there's unique tokenization, if not the video as a whole then for each image. The whole video is ~1800 tokens because Gemini gets video as a series of images in context at 1 frame/s. Each image is about 258 tokens because a token in image transformer terms is literally a patch of the image. https://arxiv.org/abs/2010.11929

>This is not at all how this works.

You can literally convert the tokens returned from a video to text. What do you even think tokens are?

Like seriously, before you write another word on this feel free to call the API and retrieve tokens for a video or image. Now go through the magical process of converting those tokens back to their text form. It isn't some magical hyper-dimensional, inside-out spatial encoding that yields impossible compression.

This process is obvious and logical if actually thought through.

>Each image is about 258 tokens

Because Google set that as the "budget" and truncates accordingly. Again, call the API with an image or video and then convert those tokens to text.

>https://arxiv.org/abs/2010.11929

This is super weird, and does not remotely prove your point. I literally spend most of my days in ViTs, but thanks for the link.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#452

Earlier quoted context omitted.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

Most likely that’s because cocktail rhymes with Molotov What definition of 'rhymes' are you using here?

This definition: https://www.collinsdictionary.com/dictionary/english/figurat...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#453
post #99

At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…

Image tokens =/ Text tokens. Image tokens are patches of the image. Each image is divided into ~256 parts. Those parts are the tokens. There's no separate run to another OCR.

Completely wrong.

Well, aside from the edited in bit about OCR. Of course there isn't a separate run to do OCR because that was literally the first step during image analysis. You know, before the conversion to simple tokens.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#454

Earlier quoted context omitted.

I'd be ok with it refusing to explain how to create explosives or illegal drugs, and refusing to generate underage nudes.

Would that include: - How to make a baking soda volcano - How to make legal drugs at home from scratch (this violates patents) - Explaining how a fictional character in a popular TV show created the drugs shown on screen - Giving you the titles of legally sold books that explain how illegal drugs are made It's an interesting thought experiment.

It's not even a thought experiment, it's a philosophical debate on morals and laws vs freedom and whatnot. It's not an easy one, and it goes back decades if not hundreds of years; remember things like the Anarchist's Cookbook?

(Sidenote, there's a conspiracy theory that the Anarchist's Cookbook is intentionally wrong with some formulations to foil would-be bombers)

Re: The killer app of Gemini Pro 1.5 is using video as an input

#455

Earlier quoted context omitted.

Image tokens =/ Text tokens. Image tokens are patches of the image. Each image is divided into ~256 parts. Those parts are the tokens. There's no separate run to another OCR.

Completely wrong. Well, aside from the edited in bit about OCR. Of course there isn't a separate run to do OCR because that was literally the first step during image analysis. You know, before the conversion to simple tokens.

There's no run to any OCR, first step or not.

And you have no idea what you're talking about.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#456
post #430

Earlier quoted context omitted.

So how crippled would you like them to be? Would you put any guard rails in place?

Assuming the person interacting with it is an adult, does it need any guard rails at all?

Yes it does, I don't want AI generating something that is illegal in my country. And it cannot make assumptions about where I live, due to VPNs and the like.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#457

Earlier quoted context omitted.

Completely wrong. Well, aside from the edited in bit about OCR. Of course there isn't a separate run to do OCR because that was literally the first step during image analysis. You know, before the conversion to simple tokens.

There's no run to any OCR, first step or not. And you have no idea what you're talking about.

[deleted]

Re: The killer app of Gemini Pro 1.5 is using video as an input

#458

Earlier quoted context omitted.

This is not at all how this works. There's no separate model. Yes there's unique tokenization, if not the video as a whole then for each image. The whole video is ~1800 tokens because Gemini gets video as a series of images in context at 1 frame/s. Each image is about 258 tokens because a token in image transformer terms is literally a patch of the image. https://arxiv.org/abs/2010.11929

>This is not at all how this works. You can literally convert the tokens returned from a video to text. What do you even think tokens are? Like seriously, before you write another word on this feel free to call the API and retrieve tokens for a video or image. Now go through the magical process of converting those tokens back to their text form. It isn't some magical hyper-dimensional, inside-out spatial encoding tha…

>You can literally convert the tokens returned from a video to text. What do you even think tokens are?

Tokens are patches of each image.

It's amazing to me how people will confidently spout utter nonsense. It only takes looking at the technical report for the Gemini models to see that you're completely wrong.

https://arxiv.org/abs/2312.11805

>The visual encoding of Gemini models is inspired by our own foundational work on Flamingo (Alayrac et al., 2022), CoCa (Yu et al., 2022a), and PaLI (Chen et al.,2022), with the important distinction that the models are multimodal from the beginning and can natively output images using discrete image tokens (Ramesh et al., 2021; Yu et al., 2022b).

Re: The killer app of Gemini Pro 1.5 is using video as an input

#459
post #430

Earlier quoted context omitted.

Assuming the person interacting with it is an adult, does it need any guard rails at all?

Yes it does, I don't want AI generating something that is illegal in my country. And it cannot make assumptions about where I live, due to VPNs and the like.

Do you want AI to follow the blasphemy laws of every country that has them?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#460

Earlier quoted context omitted.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

I think it's the COCK in cocktail.

Scunthorpe problem; I thought an AI should be smart enough to know the difference? https://en.wikipedia.org/wiki/Scunthorpe_problem
Post reply on HN