Live data from Hacker News

Lip Reading as a Service (Read Their Lips by Symphonic Labs)

readtheirlips.com

1–10 of 15 posts

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#2
Great way to build labeled training data.

User-submitted videos (with audio for STT), user-crafted bounding boxes (we might not need these soon), and user-guided RLHF.

The submitted videos are likely diverse, challenging (otherwise the human might just do it), and representative of solving actual customer problems.

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#3
Has anyone tried this with some video where they know what the person is saying?

I'd be interested to know how accurate it is, from what angles it will read lips at (front facing, side, etc).

Sounds promising if it works well. Imagine all the historical videos without sound you could try to finally know what was being said.

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#4
Thinking through some potentially interesting sources for videos where two people are talking but we don't know what was said and well, I think this is a decent starting point: https://www.youtube.com/watch?v=KLcfpU2cubo

Sadly, doesn't work too great in this situation:

> That they didnt go through but i would tell you theyre just a chill look at here lets do it chills with all of our great men and they look at every chance they go oh do you want to the black man well thats my gosh thats my gosh thats my gosh thats my gosh thats my gosh thats my gosh thats

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#7
post #2

Great way to build labeled training data. User-submitted videos (with audio for STT), user-crafted bounding boxes (we might not need these soon), and user-guided RLHF. The submitted videos are likely diverse, challenging (otherwise the human might just do it), and representative of solving actual customer problems.

Doesn't even need to be user guided. Use videos that have audio. You could have one AI that generates a transcript using the audio/video and another that watches the video on mute and tries to read the lips. Feedback would then be provided by the AI that had access to the audio.

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#8
post #2

Great way to build labeled training data. User-submitted videos (with audio for STT), user-crafted bounding boxes (we might not need these soon), and user-guided RLHF. The submitted videos are likely diverse, challenging (otherwise the human might just do it), and representative of solving actual customer problems.

Doesn't even need to be user guided. Use videos that have audio. You could have one AI that generates a transcript using the audio/video and another that watches the video on mute and tries to read the lips. Feedback would then be provided by the AI that had access to the audio.

I am thinking of the millions of hours of tv news. Presenters are almost always going to be the same position in frame and may already have high quality transcripts.

Re: Lip Reading as a Service (Read Their Lips by Symphonic Labs)

#10
post #3

Has anyone tried this with some video where they know what the person is saying? I'd be interested to know how accurate it is, from what angles it will read lips at (front facing, side, etc). Sounds promising if it works well. Imagine all the historical videos without sound you could try to finally know what was being said.

Experienced lip readers are lucky to get half of what is said. Better than nothing but not reliable enough for anything and so better to use something else if possible.

'i love you' and 'island view' have the same lip movements is the clasical example.

Post reply on HN