Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

481–490 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#481

Earlier quoted context omitted.

>As much as I would love to waste my time replying again to your nonsense, instead I'll just politely chuckle and move on. Good luck. You have your head so far up your ass even direct confirmation from the model builders themselves won't sway you. The comment wasn't for you. The comment is linked sources for the original poster and for the curious. You see I don't have to hide behind a veneer of "Trust me bro. It wor…

>even direct confirmation from the model builders themselves Linking papers that you clearly haven't read and can't contextually apply -- as with the ViT or your misunderstanding of image tiling -- is not the sound strategy you hope it is. It doesn't confirm your claims. I'm not asking anyone to "Trust me bro". So...have you called the Gemini Pro 1.5 API and tokenized an image or a video yet? There is a certain eleme…

>So...have you called the Gemini Pro 1.5 API and tokenized an image or a video yet?

You continue to blow my mind. Have you...have you even used the gemini pro api before ? You can't use the api to get the image tokens.

>This betrays a fundamental ignorance of even the most basic of information theory.

Wow, something else you don't understand. Go figure.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#482
post #287

Earlier quoted context omitted.

It doesn't even have to coach you at your job, simply a LLM-powered fuzzy retrieval would be great. Where did I put that file three weeks ago? What was that trick that I had to do to fix that annoying OS config issue? I recall seeing a tweet about a paper that did xyz about half a year ago, what was it called again? Of course taking notes and bookmarking things is possible, but you can't include everything and it tak…

Eventually someone will realise that it'd also be great for telling you where you left your keys, if it'd film everything you see instead of just your screen.

Also, just in case someone thinks this is an exaggeration, Meta is actively working to realize this with the Aria glasses. They just released another large dataset with such daily activities.

https://twitter.com/_akhaliq/status/1760502294016036986

Privacy concerns will not stop it, just like it didn't stop social media (and other) tracking. People have been taught the mantra that "if you have nothing to hide, ...", and everyone accepts it.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#483
post #287

Earlier quoted context omitted.

Eventually someone will realise that it'd also be great for telling you where you left your keys, if it'd film everything you see instead of just your screen.

I simply am not going to have my entire life filmed by an form of technology, I don't care what the advantages are. There's a limit to the level of dystopian dependent uses of these technologies I'm going to put up with. I sincerely hope the majority of the human race feels the same way.

This is not how most people think. If it's convenient and has useful features, it will spread. Soon enough it will be expected that you use it, just like it's expected today to have a smartphone and install apps to participate in events, or to use zoom etc.

By the way, Meta is already working to realize such a device. Like Alexa on steroids, but it also sees what you see and remembers it all. It's not speculation, it is being built.

https://twitter.com/_akhaliq/status/1760502294016036986

Re: The killer app of Gemini Pro 1.5 is using video as an input

#484

Earlier quoted context omitted.

But they are not a "single integer" either as in, like a byte... I don't have any good examples but I'm pretty sure the tokens are in the range of thousands of dimensions. It has to encode the properties of the patch of the image it derives from, and even a small 40x40 RGB pixel patch has plenty of information you have to retain.

A token is a single integer from a dictionary of a given model's vocabulary (e.g. GPT-4 has a vocab of ~100k different tokens, Gemma has ~256k). You are discussing embeddings which are a deeper, different element of models. https://platform.openai.com/tokenizer In the given example the video was condensed to a sequence of 258 tokens, and clearly it was a very minimalist, almost-entirely-ocr extraction from the video.

Yeah but we're not talking about LLMs here but vision transformers, which don't use the same type of token vocabulary to produce embeddings from the input as the LLMs do. The pixel data is much more dense than a few characters is, per token.

I looked it up - the original ViT models directly projected for example 16x16 pixel patches into 768-dimensional "tokens". So a 224x224 image ended up as 14*14=196 "tokens" each of which is a 768-dimensional vector. The positional encoding is just added to this vector.

This blog-post has the specific numbers, which makes it a bit less abstract than in the original paper: https://amaarora.github.io/posts/2021-01-18-ViT.html

Re: The killer app of Gemini Pro 1.5 is using video as an input

#485

Earlier quoted context omitted.

If you are averse to seeing links to paywalled articles you probably shouldn't use HN

If you see a comment complaining about a paywall, it's usually a request for someone to archive it for everyone's benefit, and it's usually a request that gets fulfilled.

Yes exactly its kind of implied, and not trying to be rude.. it would help if the person posting the paywalled link also posts an archive link of course!

Re: The killer app of Gemini Pro 1.5 is using video as an input

#487
post #364

Earlier quoted context omitted.

There are 6 countries listed in that article, out of the nearly 200 countries in the world. Hardly "most." And there doesn't appear to be examples of those 6 countries imprisoning people for those laws.

See this[1]. Most sampled countries have laws against hate speech. Certainly most of the ones western world care about. Also see [2] for examples of arrest. [1]: https://www.reddit.com/r/MapPorn/comments/qh7ua1/hate_speech... [2]: https://www.nytimes.com/2022/09/23/technology/germany-intern...

So basically, you have no real proof to back up your claim that "most" countries are "putting people in jail"

Re: The killer app of Gemini Pro 1.5 is using video as an input

#488
post #195

Earlier quoted context omitted.

To me this sounds like an opinion that would be common in the US, mostly because of where the trust and fears seem to be (private companies versus government). I think everybody (private companies, government, individuals) will try to influence and will affect your personal life. What I am worried about is who has the most efficient way to influence a lot the average person - because that entity can control on long t…

Can we please argue on the thing being discussed rather than where it is common? Are you saying influencing life through ads and putting me in jail have similar effect on me? If you combine all laws of my country I am pretty sure I would have broken few unintentionally. If government wants to just put me in jail they could retroactively find any of my past instance if they have the data. This is not some theoretical…

The "thing being discussed" is the efficacy of privacy laws. They work well, and the fact that you haven't been put on trial for your 'crimes' yet is tacit evidence.

In the real world, both corporations and governments are your enemy. You're mistakenly looking at it as a relativist comparison; the people influencing your life through advertising work with the people who put you in jail. They aggregate and sell data to Palantir which is used by dozens of well-meaning intelligence agencies to scrutinize their citizens. They threaten Apple and Google unless they turn over personally-identifying data and account details. Some of them even demand that corporate data is stored on state-owned servers.

So, what you actually want is to use the power of the "putting me in jail" people against your oppressors. If the law says that companies can't collect data unconditionally, then neither the corporation or the state can justly implicate you.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#489
post #364

Earlier quoted context omitted.

There are 6 countries listed in that article, out of the nearly 200 countries in the world. Hardly "most." And there doesn't appear to be examples of those 6 countries imprisoning people for those laws.

See this[1]. Most sampled countries have laws against hate speech. Certainly most of the ones western world care about. Also see [2] for examples of arrest. [1]: https://www.reddit.com/r/MapPorn/comments/qh7ua1/hate_speech... [2]: https://www.nytimes.com/2022/09/23/technology/germany-intern...

Reply to swigz: Apart from the link in previous comment, [1] has more examples

[1]: https://edition.cnn.com/2021/08/05/football/hate-crime-arres...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#490
post #465

Note that a video is just a sequence of images: OpenAI has a demo with GPT-4-Vision that sends a list of frames to the model with a similar effect: https://cookbook.openai.com/examples/gpt_with_vision_for_vid... If GPT-4-Vision supported function calling/structured data for guaranteed JSON output, that would be nice though. There's shenanigans you can do with ffmpeg to output every-other-frame to halve the costs too.…

How is sound handled? All I see in the Gemini docs is a terse sentence that says it isn’t included, which doesn’t sound like an optimal solution.

Models have to be trained to understand sound, it's not free.
Post reply on HN