Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

311–320 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#311
post #181
post #144

> 7 second video consumed just 1,841 tokens How? Video is a massive amount of data

It turns out it's 258 tokens per frame, and they only sample one frame every second.

So what sort of information is being left out?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#312

I wonder if the real killer app is Googles hardware scale verses OpenAi' s(or what Microsoft gives them). Seems like nothing Google's done has been particular surprising to OpenAi's team, it's just they have such huge scale maybe they can iterate faster.

The real moat is that Google has access to all the video content from YouTube to train the AI on, unlike anyone else.

[deleted]

Re: The killer app of Gemini Pro 1.5 is using video as an input

#313
post #308

In the same vein as "agents watching your screen" - what about "agents watching your posture"? Pages like [0] and [1] exist because people experience great benefits from becoming (even slightly) more aware of the way they are holding their bodies. Imagine this idea taken to the extreme, with a local agent intelligently reminding you to tighten your core, square your shoulders, relax your tongue, or warning of potenti…

> He sat as still as he could on the narrow bench, with his hands crossed on his knee. He had already learned to sit still. If you made unexpected movements they yelled at you from the telescreen.

agreed, anything FAANG/internet facing with this capability is Orwellian, which is why

> local

is explicitly included in the idea.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#314

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

If that much processing power is that cheap, this phase you’re describing is going to be fleeting because at that point I feel like it could just come up with ideas and code it itself.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#315
post #115

Earlier quoted context omitted.

That's why we would want it to run locally! Think about a fully personalized model that can work out some simple tasks / code while you're going out for groceries, or potentially more complex tasks while you're sleeping.

"AI Companion" is a bit like spouse. You are married to it in the long run, unless you decide to divorce it. Definitely TRUST is the basis of marrage, and it should be the same for AI models. As in human marriage, there should be a law that said your AI-companion cannot be compelled to testify against you :-)

But unlike a spouse you can reset it back to an earlier state you preferred.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#316

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

I'm working on this! https://www.perfectmemory.ai/ It's encrypted (on top of Bitlocker) and local. There's all this competition who makes the best, most articulate LLM. But the truth is that off-the-shelf 7B models can put sentences together with no problem. It's the context they're missing.

This looks cool, I hope you support macOS at some point in the future

Re: The killer app of Gemini Pro 1.5 is using video as an input

#317

Earlier quoted context omitted.

Your website and blog are very low on details on how this is working. Downloading and installing an mai directly feels unsafe imo. Especially when I don't know how this software is working. Is it recording a video, performing OCR continuously, taking just screenshots No mention of using any LLMs in there at all which is how you are presenting it in your comment here.

Feedback taken. I'll add more details on how this works for us technical people. LLM integration is in progress and coming soon. Any idea what would make you feel safe? 3rd party verification? I had it verified and published by the Microsoft Store. I feel eventually it all comes down to me being a decent person.

welp. this pretty much convinces me that its time I get out of tech. lean into the tradework I do in my spare time.

because I'm sure you and people like you will succeed in your endeavors, naively thinking you're doing good. and you or someone like you will sell out, the most ruthless investor will take what you've built and use it as one more cludgel of power to beat the rest of us with.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#318

Earlier quoted context omitted.

It's not that bad. With Perfect Memory AI I see ~9GB a month. That's 108 GB/year. HDD/SSDs are getting bigger than that every year. The storage also varies by what you do, your workflow and display resolution. Here's an article I wrote on my finding of storage requirements. https://www.perfectmemory.ai/support/storage-resources/stora... And if you want to use the data for LLM only, then you don't need to store the sc…

Does storage use scale linearly with the number of connected monitors (assuming each monitor uses the same resolution)?

Most screenshots are of the application window in the foreground, so unless your application spans all monitors, there is no significant overhead with multiple monitors. DPI on the other hand has a significant impact. The text is finer, taking more pixels...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#319

Guess the author didn't bother to check that those books actually are correct? The first one I checked, "Growing Up with Lucy by April Henry" doesn't exist. The actual book is by Steve Grand, and it's very obviously so in the video used as input. So a cool demo, but sadly useless for anything more.

Thanks for this comment. I am yet to see any “art” produced by AI that is not superficial or hollow (best case) or deeply unsettling (common case).

Re: The killer app of Gemini Pro 1.5 is using video as an input

#320
post #287

Earlier quoted context omitted.

Eventually someone will realise that it'd also be great for telling you where you left your keys, if it'd film everything you see instead of just your screen.

I simply am not going to have my entire life filmed by an form of technology, I don't care what the advantages are. There's a limit to the level of dystopian dependent uses of these technologies I'm going to put up with. I sincerely hope the majority of the human race feels the same way.

People already fill their homes with nanny cams. Very soon someone will hook those up to LLMs so you can ask it what happened at home while you were gone.
Post reply on HN