Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

421–430 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#421

Note that a video is just a sequence of images: OpenAI has a demo with GPT-4-Vision that sends a list of frames to the model with a similar effect: https://cookbook.openai.com/examples/gpt_with_vision_for_vid... If GPT-4-Vision supported function calling/structured data for guaranteed JSON output, that would be nice though. There's shenanigans you can do with ffmpeg to output every-other-frame to halve the costs too.…

On the other hand, a picture is a video with a single frame.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#422
Things are going to get strange as soon as we have AI wearables that monitor everything a person does/sees/hears in real time and privately offers them suggestions. It will seem great at first, vigilant life-coaching for people who need help, or knowledge/memory enhancement to make effective people even more effective. But what happens when people really start to trust the voice whispering in their ear and defer all their decision making to it? They'll probably become addicted to it, then enslaved to it. They will become meat puppets for the AI.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#423

It's Google. I'd rather avoid sharing my thoughts and interests with this Borg-like entity.

Yeah, I really hope open sources catches up quickly. Why on earth would I want to create a Google account just to use this, especially in work settings?

I think it is only a matter of time before open source vision LLMs have the ability to process videos. The tricky part might be getting to 1M token context length, which even proprietary LLMs (other than Gemini) are struggling with.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#424
post #416

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

I was fighting with ChatGPT yesterday because it wouldn't translate "fuck". I was quoting Office Space's "PC Load Letter? What the fuck does that mean?" Likewise it won't generate passive-aggressive answers meant for comedic reasons. I hate having to negotiate with AI like it's a difficult child.

I wonder, if you put asterisks like 'f***' it would translate that appropriately. Like, as a figleaf.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#425

Earlier quoted context omitted.

Paywall

If you are averse to seeing links to paywalled articles you probably shouldn't use HN

If you see a comment complaining about a paywall, it's usually a request for someone to archive it for everyone's benefit, and it's usually a request that gets fulfilled.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#427

Earlier quoted context omitted.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

Most likely that’s because cocktail rhymes with Molotov What definition of 'rhymes' are you using here?

It is like a joke saying. Saying something rhymes with something that doesn't actually rhyme is saying that the two things go together and when one hears the first they think the second also

Re: The killer app of Gemini Pro 1.5 is using video as an input

#428

Earlier quoted context omitted.

Most screenshots are of the application window in the foreground, so unless your application spans all monitors, there is no significant overhead with multiple monitors. DPI on the other hand has a significant impact. The text is finer, taking more pixels...

Why should DPI matter if the app is taking screenshots?

Because screenshots are in pixels, not inches.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#429
post #380

Earlier quoted context omitted.

I got 90% of this built on Linux (around KDE Wayland) before other interests/priorities took over: https://github.com/Zetaphor/screendiary/

This seems very very interesting. I'm still learning python so probably can't build on this. But like a cheap mans' version of this would be to take a screenshot every couple of minutes, OCR it and send to it gpt for some kind of processing (or not, just keep it as a log). Right? Or am I missing something?

Yes, that's exactly what's happening here, minus the sending it off to a third-party.

I didn't see the benefit when the OCR content is fully searchable, in addition to not wanting to pay OpenAI to spy on me.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#430
post #393

Earlier quoted context omitted.

Deeply agree with the sentiment. AIs are so throttled and crippled that it makes me sad every time gemini or chatgpt refuses to answer my questions. Also agree that it’s mostly policed by American companies who follow the American culture of “swearing is bad, nudity is horrible, some words shouldn’t even be said”

So how crippled would you like them to be? Would you put any guard rails in place?

Assuming the person interacting with it is an adult, does it need any guard rails at all?
Post reply on HN