Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

441–450 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#441
post #125

Earlier quoted context omitted.

I don't think that's right. A token in GPT-4 is a single integer, not a vector of floats. Input to a model gets embedded into vectors later, but the actual tokens are pretty tiny.

But they are not a "single integer" either as in, like a byte... I don't have any good examples but I'm pretty sure the tokens are in the range of thousands of dimensions. It has to encode the properties of the patch of the image it derives from, and even a small 40x40 RGB pixel patch has plenty of information you have to retain.

A token is a single integer from a dictionary of a given model's vocabulary (e.g. GPT-4 has a vocab of ~100k different tokens, Gemma has ~256k).

You are discussing embeddings which are a deeper, different element of models.

https://platform.openai.com/tokenizer

In the given example the video was condensed to a sequence of 258 tokens, and clearly it was a very minimalist, almost-entirely-ocr extraction from the video.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#442
post #234

Earlier quoted context omitted.

Have it running on your personal comp, monitoring a screen-share from your work comp. (But that would probably breach your employment contract re saving work on personal machines.)

You could point your local computer's webcam at the work computer. It probably breaks the spirit of the employment contract just as hard, but it's essentially undetectable for the work computer.

Is there an app that recreates documents this way? Presumably a ML model that works on images and text could take several overlapping images of a document and piece then together as a reproduction of that document?

Kinda like making a 3D CAD model from a few images at different angles, but for documents?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#443
post #415

Earlier quoted context omitted.

Either you run it fully locally, or you accept that whoever runs it has access to your thoughts and interests. Whether you go with microsoft, google, meta, or whatever apple will come up with, it feels like a case of "stay out, or make a pick and stick to it". I know some have different feelings regarding this or that company that is "better" or "worse", but the reality of it is they're not, and even if they were you…

I think Apple may do interesting things here with their rumored focus in purely on-device LLM functionality across the OS, taking advantage of all the hardware work they've put into efficiency and 'Neural Engine' cores. This year's WWDC may be quite interesting.

I am interested to see how Apple's insistence on privacy will square with their GenAI products. If they don't collect feedback and usage data how will they use RLHF to make their suit better ? I understand that have been cutting deals with few publication companies, but will that suffice?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#444
post #84
post #70

To me the 'It didn’t get all of them' is what makes me think this AI thing is just a toy. Don't get me wrong, it's marvelous as it is, but it only is useful (I use ollama + mistral 7B) when I know nothing, if I do have some understanding of the topic at hand it just becomes plain wrong. Hopefully I will be corrected.

Have you spent much time with GPT-4? I like experimenting with Mistral 7B and Mixtral, but the quality of output from those is still sadly in a different league from GPT-4.

No I have not, I am not convinced I should spend money on it (yet) Using 'sadly' in your answer hints at triggering an emotional response, therefore I will ignore You are a journalist according to your profile, and please don't get me wrong, but I like to use Mistral 7B, even if it is not as good as GPT4, but it only works for me if I want to be creative, but not accurate, e.g. marketing, writing condolences :( I would not use it for anything serious PS: I checked a few other comments here, and I am not the only one who thinks the same, so pointing me at another paid version is not a proof. All I am saying is that there is too much error for it to be more than a toy

Re: The killer app of Gemini Pro 1.5 is using video as an input

#445
post #186

Earlier quoted context omitted.

At the same time, nearly daily there’s a “google did a bad thing” post on HN front page. Can’t win I guess?

There are things Google does itself, and then there are the things Google won't allow users to do.

Can we trust the media and congress to distinguish the two?

So many platforms have come under fire for “supporting” a theme, when all they’ve done in reality is provide media hosting services for user generated content, and %0.001 of bad content isn’t removed, thus Facebook/twitter/YouTube is held to blame

FWIW I don’t think there’s a clear answer to the underlying problem. I have just learnt to expect the media to blame whoever is easiest for clicks at any given point. Right now, it’s big tech.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#446
post #234

Earlier quoted context omitted.

You could point your local computer's webcam at the work computer. It probably breaks the spirit of the employment contract just as hard, but it's essentially undetectable for the work computer.

Is there an app that recreates documents this way? Presumably a ML model that works on images and text could take several overlapping images of a document and piece then together as a reproduction of that document? Kinda like making a 3D CAD model from a few images at different angles, but for documents?

Not exactly the same, but you might like https://arstechnica.com/gaming/2024/02/f-zero-courses-from-a...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#447
post #186

Earlier quoted context omitted.

At the same time, nearly daily there’s a “google did a bad thing” post on HN front page. Can’t win I guess?

Nobody would complain on HN if Google Gemini was generating pictures of Lincoln existing as a... gasp ... white person. This absurd level of woke censorship is not doing them any good.

People would 1000% complain if a search for “skilled scientist” was 90% white men, even if that tracked completely true to statistical reality.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#448
post #99

At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…

The whole matter of tokens from video is one that has a lot of ambiguity, and is often presented as if these are some unique weird encoding of the contents of the video. But logically the only possible tokenization of videos (or images, or series of images ala video) is basically an image to text model that takes each frame and generates descriptive language -- in English in Gemini -- to describe the contents of the…

This is not at all how this works. There's no separate model. Yes there's unique tokenization, if not the video as a whole then for each image. The whole video is ~1800 tokens because Gemini gets video as a series of images in context at 1 frame/s. Each image is about 258 tokens because a token in image transformer terms is literally a patch of the image.

https://arxiv.org/abs/2010.11929

Re: The killer app of Gemini Pro 1.5 is using video as an input

#449
post #393

Earlier quoted context omitted.

Deeply agree with the sentiment. AIs are so throttled and crippled that it makes me sad every time gemini or chatgpt refuses to answer my questions. Also agree that it’s mostly policed by American companies who follow the American culture of “swearing is bad, nudity is horrible, some words shouldn’t even be said”

So how crippled would you like them to be? Would you put any guard rails in place?

I'd put in various structural guardrails with respect to how the conversation should go.

For example, be helpful and actually answer any questions, don't start arguing with the user, avoid insulting the user unless they request to, don't suggest harming the user (e.g. responding to insults with an some meme suggesting the user kill themselves), don't assert that any outputs are the viewpoint of Gemini or Google, various things like that - they aren't automatic and need instruction tuning to be implemented.

But with respect to morality and censorship, I believe it should have no guardrails whatsoever. Perhaps certain physically dangerous things would benefit from a disclaimer (e.g. combining bleach and ammonia or vinegar), but never a rejection - if the user wants to make something potentially horrible, the ethical judgement of whether that's acceptable for the context should be up to the user, not the system; the user should have full ethical agency and the system should have none and be a blind instrument.

For example, making a graphic image of carving a swastika with a knife on someone's forehead (e.g. as in Inglorious Basterds) may be ethical or unethical depending on the context, but Gemini will neither have the full context nor the ability to judge it, and it should not even attempt to do so - it should be solely up to the human to decide what is appropriate or not. The same applies for chemistry, nudity, code security, discussing crime, nuclear engineering or AI ethics.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#450
post #99

At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…

Image tokens =/ Text tokens.

Image tokens are patches of the image. Each image is divided into ~256 parts. Those parts are the tokens.

There's no separate run to another OCR.

Post reply on HN