Note that a video is just a sequence of images: OpenAI has a demo with GPT-4-Vision that sends a list of frames to the model with a similar effect: https://cookbook.openai.com/examples/gpt_with_vision_for_vid... If GPT-4-Vision supported function calling/structured data for guaranteed JSON output, that would be nice though. There's shenanigans you can do with ffmpeg to output every-other-frame to halve the costs too.…
The killer app of Gemini Pro 1.5 is using video as an input
401–410 of 507 posts
Re: The killer app of Gemini Pro 1.5 is using video as an input
#402Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…
This is already possible to implement today, so it's very likely that we'll all have our own personal AIs that know us better than we do.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#403> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.
Deeply agree with the sentiment. AIs are so throttled and crippled that it makes me sad every time gemini or chatgpt refuses to answer my questions. Also agree that it’s mostly policed by American companies who follow the American culture of “swearing is bad, nudity is horrible, some words shouldn’t even be said”
Re: The killer app of Gemini Pro 1.5 is using video as an input
#404Re: The killer app of Gemini Pro 1.5 is using video as an input
#405I'd rather avoid sharing my thoughts and interests with this Borg-like entity.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#406Re: The killer app of Gemini Pro 1.5 is using video as an input
#407Really. I am not that impressed. It is not something radically different from doing the same thing with a still photo which by now is trivial for those models. What is being tested here doesn't require a video. It is not showing to be able to derive any meaning from a short clip. It is fucking doing very fancy OCR, that's all. What would impress me is if shown a clip of an open chest surgery it was able to comment wh…
You mean like in this demo? https://www.youtube.com/watch?v=wa0MT8OwHuk
Re: The killer app of Gemini Pro 1.5 is using video as an input
#408It's Google. I'd rather avoid sharing my thoughts and interests with this Borg-like entity.
Re: The killer app of Gemini Pro 1.5 is using video as an input
#409> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.
I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.
What definition of 'rhymes' are you using here?
Re: The killer app of Gemini Pro 1.5 is using video as an input
#410Earlier quoted context omitted.
Unless you live in the EU and have laws that should protect you from that.
Is it true or more of a myth? Based on my online read, Europe has "think of the children" narrative as common if not more than other parts of the world. They tried hard to ban encryption in apps many times.[1] [1]: https://proton.me/blog/eu-council-encryption-vote-delayed