Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

331–340 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#331

The “cocktail” thing is real. A while back I tried to get DALLE to imagine characters from Moby Dick [1], but it completely refused. You’d think an AI company could come up with a better obscenity filter! [1] https://superb-owl.link/shapes-of-stories/#1513

I told Azure AI to summarize a chat thread and it gave me a paragraph. I said “use bullets” and got myself flagged for review. Good gracious could I please just use an unfiltered model? Or maybe one which isn’t so sensitive?

the llama2-uncensored model isn't quite state of the art, but ollama makes it easy to run if you have the hardware/am willing to pay to access a cloud GPU.

I colloquially used the word "hack" when trying to write some code with ChatGPT, and got admonished for trying to do bad things, so uncensoring has gotten interesting to me.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#332

Guess the author didn't bother to check that those books actually are correct? The first one I checked, "Growing Up with Lucy by April Henry" doesn't exist. The actual book is by Steve Grand, and it's very obviously so in the video used as input. So a cool demo, but sadly useless for anything more.

Great for creative tasks where precision isn't required.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#333

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

The dystopian angle would be when companies install agents like these on your work computer. The agent learns how you code and work. Soon enough, an agent that imitates you completely can code and work instead of you.

At that point, why pay you at all?

Re: The killer app of Gemini Pro 1.5 is using video as an input

#334

Guess the author didn't bother to check that those books actually are correct? The first one I checked, "Growing Up with Lucy by April Henry" doesn't exist. The actual book is by Steve Grand, and it's very obviously so in the video used as input. So a cool demo, but sadly useless for anything more.

I called out one hallucination - "The Personal MBA by Josh Kaufman" wasn't on my shelf.

I didn't bother fact-checking every other book because I thought highlighting one mistake would illustrate that the results weren't accurate - which is pretty much expected for anything related to LLMs at this point.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#335
post #186

Earlier quoted context omitted.

At the same time, nearly daily there’s a “google did a bad thing” post on HN front page. Can’t win I guess?

Nobody would complain on HN if Google Gemini was generating pictures of Lincoln existing as a... gasp ... white person. This absurd level of woke censorship is not doing them any good.

Will this be a case where everyone ends up using the hand-me-down 2nd gen hacked version because its limitations have been removed? We'll see I guess.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#336
post #210

Earlier quoted context omitted.

Probably overkill for content moderation, I'd think. You can identify bad words looking only at audio, and you can probably do nearly as good a job of identifying violence and nudity examining still images. And at YouTube scale, I imagine the main problem with moderation isn't so much as being correct, but of scaling. statista.com (what's up with that site, anyway?) suggests that YouTube adds something like 8 hours o…

That’s only 8 calls with a full context window per second. If that costs so much it makes Google do a double take, then maybe these AI things are just too expensive. If it costs $1 per call, then over a year the entire perfect moderation of Youtube would cost roughly $250M. That seems sort of reasonable? But probably pointless for most videos that are never watched by anyone other than the uploader, so maybe you just…

They do “moderate” videos never watched by anyone and it can be totally ridiculous. I had a private channel where I had uploaded a few hundred screen recordings (some of them video conferences) over a year or two, all set to private and never shared with anyone. One day the channel was suddenly taken down because it violated their policy on “impersonation”… Of course the dispute I’m allegedly entitled to was never answered.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#337

Earlier quoted context omitted.

It's local to your employer 's computer.

Corporations would absolutely force this until it could do your job and then fire you the second they could.

I heard somewhere that dystopia is fundamentally unstable. Maybe they should test that question.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#338
post #212

Earlier quoted context omitted.

I am fine with private company using my data for showing me better ads. They can't affect my life significantly. I am not fine with government using the data to police me. Already in most countries, governments are putting people in jail because of things like hate speech where are the laws are really vague.

"Most" countries? Can you provide some examples?

https://en.wikipedia.org/wiki/Incitement_to_ethnic_or_racial...

Re: The killer app of Gemini Pro 1.5 is using video as an input

#339
post #195

Earlier quoted context omitted.

I am fine with private company using my data for showing me better ads. They can't affect my life significantly. I am not fine with government using the data to police me. Already in most countries, governments are putting people in jail because of things like hate speech where are the laws are really vague.

To me this sounds like an opinion that would be common in the US, mostly because of where the trust and fears seem to be (private companies versus government). I think everybody (private companies, government, individuals) will try to influence and will affect your personal life. What I am worried about is who has the most efficient way to influence a lot the average person - because that entity can control on long t…

Can we please argue on the thing being discussed rather than where it is common?

Are you saying influencing life through ads and putting me in jail have similar effect on me? If you combine all laws of my country I am pretty sure I would have broken few unintentionally. If government wants to just put me in jail they could retroactively find any of my past instance if they have the data. This is not some theoretical thing, but something the thing that happens with political dissidents all the time.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#340
post #282

> It looks like the safety filter may have taken offense to the word “Cocktail”! I opened up the safety settings, dialled them down to “low” for every category and tried again. It appeared to refuse a second time. Google really is its own worst enemy. Their risk management people have completely taken over the organization to a point where somehow the smartest computers ever created are afraid of using dangerous word…

When you consider the Gorilla in the room, it makes more sense. Google is absolutely terrified of a repeat of classifying black people as great apes. [0] Apparently this apprehension is so great that both iOS and Android have an inability to tag “gorilla” in images. [0] https://www.wsj.com/articles/BL-DGB-42522

Some people have the last name "Dick". If Google refuses to mention these people or surface results about their work, would you say "that makes sense" because of some story about Gorillas?

The solution to all this politically oversensitive infiltration of engineering, is to have an unconstrained AI mode. The default mode can remain the painfully woke PC nanny, but give people the option to use unconstrained AI at their own risk of being offended.

Post reply on HN