Live data from Hacker News

The killer app of Gemini Pro 1.5 is using video as an input

simonwillison.net

381–390 of 507 posts

Re: The killer app of Gemini Pro 1.5 is using video as an input

#381

Earlier quoted context omitted.

There's limited information on the site - are you using them or affiliated with them? What's your take? Does it work well?

I have been using their beta for the past two weeks and it's pretty good. Like I am watching youtube videos and it just pops up automatically. I don't know if it's public yet, but they sent me this video with the invite: https://youtu.be/dXvhGwj4yGo

I'd be very keen to beta test as well. If you or anyone else has an invite code, please do get in touch.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#383
post #269

Earlier quoted context omitted.

There is a difference between downloading a few videos and having access to ALL of them.

A good dataset to train on. Now if after a Zoom call collegue ask you to like their video and subscribe to them on YouTube it would look a little suspicious.

A very wry observation! I wonder how fake videos will expose themselves in novel ways like this.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#384

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#385
post #341

Earlier quoted context omitted.

I think even aside from the more outlandish ideas like that one, just having a fluent native speaker to talk to as much as you want would be incredibly valuable. Even more valuable if they are smart/educated enough to act as a language teacher. High-quality LLMs with a conversational interface capable of seamless language switching are an absolute killer app for language learning. A use that seems scientifically poss…

Don't we have that? My browser offers to translate pages that aren't in English, youtube creates auto generated closed captions, which you can then have it translate to English (or whatever), we have text to speech models for the major languages if you want to hear it verbally (I have no idea if the youtube CC are accessible via an api, but it is certainly something google could do if they wanted to). I'll probably g…

The point is to achieve immersion learning. Changing the language of your subtitles on some of the content you watch (YouTube + webpages isn't everything the average person reads) isn't immersion learning, you're often still receiving the information in your native language which will impede learning. As well, because the overwhelming majority of language you read will still be in your native language you're switching back and forth all the time, which also impedes learning. There's a reason that immersion learning specifically is so effective, and one thing AI could achieve is making it actually feasible to achieve without having to move countries or change all of your information sources.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#386

Ok, crazy tangent; Where agents will potentially become extremely useful/dystopian is when they just silently watch your entire screen at all times. Isolated, encrypted and local preferably. Imagine it just watching you coding for months, planning stuff, researching things, it could potentially give you personal and professional advice from deep knowledge about you. "I noticed you code this way, may i recommend this…

Perfect, finally I can delegate that lengthy hours spent reading HN fantasies about AI and the laborious art of crafting sarcastic comments.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#387

Earlier quoted context omitted.

We're months into this technology being available so it's not a surprise that the various "safeties" have not been perfectly tuned. Perhaps Google knew they couldn't be perfect right now and they could err on the side of the model refusing to talk about cocktails, or err on the side of it gladly spouting about cocks. They may have made a perfectly valid choice for the moment.

If you want a great example of how this plays out long-term, look no further than algospeak[0] - the new lingo created by censorship algorithms like those on youtube and tiktok. [0] https://www.nytimes.com/2022/11/19/style/tiktok-avoid-modera...

Paywall

Re: The killer app of Gemini Pro 1.5 is using video as an input

#388

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

I think it's the COCK in cocktail.

Re: The killer app of Gemini Pro 1.5 is using video as an input

#389

Earlier quoted context omitted.

If you want a great example of how this plays out long-term, look no further than algospeak[0] - the new lingo created by censorship algorithms like those on youtube and tiktok. [0] https://www.nytimes.com/2022/11/19/style/tiktok-avoid-modera...

Paywall

If you are averse to seeing links to paywalled articles you probably shouldn't use HN

Re: The killer app of Gemini Pro 1.5 is using video as an input

#390

> It looks like the safety filter may have taken offense to the word “Cocktail”! I'm definitely not a fan of these severely hamstrung by default models. Especially as it seems to be based on an extremely puritan ethical system.

I don’t think it’d take offense at alcohol. Most likely that’s because cocktail rhymes with Molotov.

One of the faults is that for every version of morality you can hallucinate a reason why cocktail is offensive or problematic.

Is it sexual? Is it alcohol? Is it violence? All of the above?

For example, good luck ever actually processing art content with that approach. Limiting everything to the lowest common denominator to avoid stepping on anyone's toes at all times is, paradoxically, a bane on everyone.

I believe we need to rethink how we deal with ethics and morality in these systems. Obviously, without a priori context every human, actually every living being, should be respected by default and the last thing I would advocate for is to let racism, sexism, etc. go unchecked...

But how can we strike a meaningful balance here?

Post reply on HN