Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

261–270 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#261

The “cocktail” thing is real. A while back I tried to get DALLE to imagine characters from Moby Dick [1], but it completely refused. You’d think an AI company could come up with a better obscenity filter! [1] https://superb-owl.link/shapes-of-stories/#1513

I told Azure AI to summarize a chat thread and it gave me a paragraph. I said “use bullets” and got myself flagged for review.

Good gracious could I please just use an unfiltered model? Or maybe one which isn’t so sensitive?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#262

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

Finally, early-aughts 1337 a3s7h37ic can be cool again

Re: The killer app of Gemini Pro 1.5 is using video as an input

#263
post #131

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

> encrypted and local of course Only for people who'd pay for that. Free users would become the product.

I noticed you code this way, may i recommend a Lenovo Thinkpad with an Intel Xeon processor? You're sure to "wish everything was a Lenovo."

Re: The killer app of Gemini Pro 1.5 is using video as an input

#264

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

I'm working on this! https://www.perfectmemory.ai/ It's encrypted (on top of Bitlocker) and local. There's all this competition who makes the best, most articulate LLM. But the truth is that off-the-shelf 7B models can put sentences together with no problem. It's the context they're missing.

[dead]

Re: The killer app of Gemini Pro 1.5 is using video as an input

#266

Earlier quoted context omitted.

I feel like the storage requirements are really going to be these issue for these apps/services that run on "take screenshots and OCR them" functionality with LLMs. If you're using something like this a huge part of the value proposition is in the long term, but until something has a more efficient way to function, even a 1-year history is impractical for a lot of people. For example, consider the classic situation o…

This is where Microsoft (and Apple) has a leg up -- they can hook the UI at the draw level and parse the interface far more reliably + efficently than screenshot + OCR.

Google too, for all practical purposes, since presumably this is mostly just watching you use chrome 90% of the time.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#267

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

We're months into this technology being available so it's not a surprise that the various "safeties" have not been perfectly tuned. Perhaps Google knew they couldn't be perfect right now and they could err on the side of the model refusing to talk about cocktails, or err on the side of it gladly spouting about cocks. They may have made a perfectly valid choice for the moment.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#268
Cool and all, but are we going to pass the need for prompts already? I can see big usage for video access but the prompt mechanism is making it like a toy, is there an auto processing, where I predefine what to look for and feed the video and as long as the video is running it will process based on the criteria?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#269

Earlier quoted context omitted.

I’m not sure I would necessarily call YouTube a moat-creator for Google, since the content on YouTube is for all intents and purposes public data.

There is a difference between downloading a few videos and having access to ALL of them.

A good dataset to train on. Now if after a Zoom call collegue ask you to like their video and subscribe to them on YouTube it would look a little suspicious.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#270

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

I pre-ordered the rewind pendant. It will listen 24/7 and help you figure out what happened.

I bet meta is thinking of doing this with quest once the battery life improves.

https://rewind.ai/pendant

Post reply on HN