Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

111–120 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#112
post #86

Earlier quoted context omitted.

If 7 second video consumed 1k token, I'd assume the budget must be insane to process such prompt.

That's a 7 second video from an HD camera. When recording a screen, you only really need to consider whats changing on the screen.

That’s not true. What content is important context on the screen might change dependent on the new changes.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#113

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

I would hate that so much.

IKR, Who wouldn't want another Clippy constantly nagging you, but this time with a higher IQ and more intimate knowledge of you? /s

Re: The killer app of Gemini Pro 1.5 is using video as an input

#114

Earlier quoted context omitted.

Unless you live in the EU and have laws that should protect you from that.

Is it true or more of a myth? Based on my online read, Europe has "think of the children" narrative as common if not more than other parts of the world. They tried hard to ban encryption in apps many times.[1] [1]: https://proton.me/blog/eu-council-encryption-vote-delayed

> They tried hard to ban encryption in apps many times.

That's true of most places. We should applaud the EU's human rights court for leading the way by banning this behavior: https://www.eureporter.co/world/human-rights-category/europe...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#115

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

> Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. And then announcing "I can do your job now. You're fired."

That's why we would want it to run locally! Think about a fully personalized model that can work out some simple tasks / code while you're going out for groceries, or potentially more complex tasks while you're sleeping.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#116
> That 7 second video consumed just 1,841 tokens out of my 1,048,576 token limit.

is this simply an approximation done by Gemini in order to add some artificial limit on the amount of video?

Or do video frames actually equate directly to tokens somehow?

I guess my question is, is there a real relationship between videos and tokens as we understand them (i.e. "hello" is a token) or are they just using the term "tokens" because it's easy for a user to understand, and an image is not literally handled the same way a token is?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#117
post #99

At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…

Remember, if it's using a similar tokeniser to GPT-4 (cl100k_base iirc), each token has a dimension of ~100,000.

So 258x100,000 is a space of 25,800,000 floats, using f16 (a total guess) that's 51.6kB, probably enough to represent the image at ok quality with JPG.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#118

> That 7 second video consumed just 1,841 tokens out of my 1,048,576 token limit. is this simply an approximation done by Gemini in order to add some artificial limit on the amount of video? Or do video frames actually equate directly to tokens somehow? I guess my question is, is there a real relationship between videos and tokens as we understand them (i.e. "hello" is a token) or are they just using the term "tokens…

There's a new section at the bottom of the article about that.

It looks like an image is 258 tokens, and Gemini splits videos into one frame per second and processes those as images.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#119

Earlier quoted context omitted.

incentives cannot be fixed with just prohibitive laws, war on drags should've taught you something

War on drags? I thought that was just in Florida

please consider commenting more thoughtfully. I understand this is a joke but we don't want this site to devolve into Reddit.
Post reply on HN