Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

431–440 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#431
post #416

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

I was fighting with ChatGPT yesterday because it wouldn't translate "fuck". I was quoting Office Space's "PC Load Letter? What the fuck does that mean?" Likewise it won't generate passive-aggressive answers meant for comedic reasons. I hate having to negotiate with AI like it's a difficult child.

> I hate having to negotiate with AI like it's a difficult child.

Surely not in the list of things I expected to ever read in real life.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#432

Guess the author didn't bother to check that those books actually are correct? The first one I checked, "Growing Up with Lucy by April Henry" doesn't exist. The actual book is by Steve Grand, and it's very obviously so in the video used as input. So a cool demo, but sadly useless for anything more.

I think this post and others reactions and then your comment this far down really encapsulates where we’re at with this technology. Nearly 90 percent of comments on posts about LLMs are people talking about how the near future is about to boggle our minds and that general intelligence is near, but all my experiences with these LLMs show they’re capable of making the most basic of mistakes and doing so confidently and…

Humans are also perfectly capable of confidently doing mistakes.

The big difference here is that these models can scale the work beyond human capability.

Why pay 10000 mechanical turks to extract information from vids, if you can deploy N of these models, and get the work done at a fraction of the time?

Instead you can keep x% of the MTurks to check the vids where the model yields some high uncertainty score, and randomly audit other vids for quality assurance.

There's crazy amounts of potential in these things. Hell, the place I work at has already replaced certain human tasks with LLM-integrated solutions, with extremely good results.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#433

Guess the author didn't bother to check that those books actually are correct? The first one I checked, "Growing Up with Lucy by April Henry" doesn't exist. The actual book is by Steve Grand, and it's very obviously so in the video used as input. So a cool demo, but sadly useless for anything more.

Thanks for this comment. I am yet to see any “art” produced by AI that is not superficial or hollow (best case) or deeply unsettling (common case).

But then you have all the things you don't see. The CGI/fx artist that spent hours upon hours handcrafting realistic background CGI to some movie scene? Could very well be replaced in the not-so-distant future.

The first huge wave of ML/AI automation will involve all the things you don't notice straight away.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#434

Earlier quoted context omitted.

Open source isn't meant to give everyone control over a specific project. It's meant to make it so, if you don't like the project, you can fork it and chart your own direction for it.

It's meant to make it so, if you don't like the project, you can fork it and chart your own direction for it. ...accompanied by the wrath of countless others discouraging you from trying to fork if you even so much as give slight indications of wanting to do so, and then when you do, they continue to spread FUD about how your fork is inferior. I've seen plenty of discussions here and elsewhere where the one who sugge…

Is it up to the open source licenses to police the opinions people have?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#435

The “cocktail” thing is real. A while back I tried to get DALLE to imagine characters from Moby Dick [1], but it completely refused. You’d think an AI company could come up with a better obscenity filter! [1] https://superb-owl.link/shapes-of-stories/#1513

It's the Scunthorpe problem all over again

Re: The killer app of Gemini Pro 1.5 is using video as an input

#436

Earlier quoted context omitted.

Sounds like you'd choose the default woke option, and I'd choose the non-woke option. Choice is healthy. This world will leave you behind if you elect to substitute choice with monolithic wokism or any over-correcting ideology. Meanwhile: > Google is racing to fix its new AI-powered tool for creating pictures, after claims it was over-correcting against the risk of being racist. "It's missing the mark here," said Jac…

So, when you read here that they are fixing it, is that a good thing to you? Do you think that means they are turning down the censorship knob? Because in reality they are only replacing the feedback they already have in place with different feedback. Again, there is simply no such thing as an "uncensored" model if what you mean by that is something that performs as well as Gemini (or whatever) but has zero external…

I'm lost on most of your reply. No worries.

My armchair knowledge of AI tells me there's degrees of influence from the safety teams about what is permitted and what is not permitted.

My preference for "unconstrained" AI is a preference for less degrees of safety and more permissions. A preference for accuracy and objective truth over guardrails to words, facts, images, ideas.

The original definition of "woke" is morally sound, if provocative. Lately it is used as a smear due to the very incidents like this over-corrective safeguarded AI, which really is a hopeless blunder. Woke has become the descriptor for over-corrective social measures that in turn cause harm, offence, and misinformation.

Might the civil disagreement be reduced to "where should the moral baseline be". Perhaps we disagree only on that.

If I visited a sorcerer on the mountain top for advice, I'd expect unfiltered wisdom. Otherwise what's the point of walking all the way up the mountain.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#437
post #393

Earlier quoted context omitted.

Deeply agree with the sentiment. AIs are so throttled and crippled that it makes me sad every time gemini or chatgpt refuses to answer my questions. Also agree that it’s mostly policed by American companies who follow the American culture of “swearing is bad, nudity is horrible, some words shouldn’t even be said”

So how crippled would you like them to be? Would you put any guard rails in place?

I'd be ok with it refusing to explain how to create explosives or illegal drugs, and refusing to generate underage nudes.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#438
post #416

Earlier quoted context omitted.

I was fighting with ChatGPT yesterday because it wouldn't translate "fuck". I was quoting Office Space's "PC Load Letter? What the fuck does that mean?" Likewise it won't generate passive-aggressive answers meant for comedic reasons. I hate having to negotiate with AI like it's a difficult child.

> I hate having to negotiate with AI like it's a difficult child. Surely not in the list of things I expected to ever read in real life.

That's really how it feels. "ChatGPT, this is a quote from a movie. You don't need to be afraid of it. The man is angry at a printer, and it's funny. Let's just translate it to Pashto, it will take a few seconds and then we go back to simple questions, okay?"

Re: The killer app of Gemini Pro 1.5 is using video as an input

#439
post #99

At the end of the article, a single image of the bookshelf uploaded to Gemini is 258 tokens. Gemini then responds with a listing of book titles, coming to 152 tokens. Does anyone understand where the information for the response came from? That is, does Gemini hold onto the original uploaded non-tokenized image, then run an OCR on it to read those titles? Or are all those book titles somehow contained in those 258 to…

The whole matter of tokens from video is one that has a lot of ambiguity, and is often presented as if these are some unique weird encoding of the contents of the video.

But logically the only possible tokenization of videos (or images, or series of images ala video) is basically an image to text model that takes each frame and generates descriptive language -- in English in Gemini -- to describe the contents of the video.

e.g. A bookshelf with a number of books. The books seen are "...", "...", etc. A figurine of a squirrel. A stuffed owl.

And so on. So the tokenization by design would include the book titles as the primary information, as that's the easiest, most proven extraction from images.

From a video such tokenization would include time flow information. But ultimately a lot of the examples people view are far less comprehensive than they think.

It isn't surprising that many demonstrations of multimodal models always includes an image with text on it somewhere, utilizing OCR.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#440

Earlier quoted context omitted.

So how crippled would you like them to be? Would you put any guard rails in place?

I'd be ok with it refusing to explain how to create explosives or illegal drugs, and refusing to generate underage nudes.

Would that include:

- How to make a baking soda volcano

- How to make legal drugs at home from scratch (this violates patents)

- Explaining how a fictional character in a popular TV show created the drugs shown on screen

- Giving you the titles of legally sold books that explain how illegal drugs are made

It's an interesting thought experiment.

Post reply on HN