Live data from Hacker News

Show HN: A real time AI video agent with under 1 second of latency

news.ycombinator.com

221–230 of 264 posts

Re: Show HN: A real time AI video agent with under 1 second of latency

#222
post #205

Earlier quoted context omitted.

This is a valid concern, but we’ve always been very serious about consent and privacy. Our models cannot be used without explicit verbal/visual consent and you hold the keys to your clone.

I don't trust any of you AI people with that.

You think the rep from the AI doppelganger company is a people? Voight-Kampff may say otherwise.

Re: Show HN: A real time AI video agent with under 1 second of latency

#223
Pretty cool but it seems like the mouth / lip-sync is quite a bit off, even for the video generation API? Is that the best rendering, or are the videos stale?

Also the audio cloning sounds quite a bit different from the input on https://www.tavus.io/product/video-generation

For live avatar conversations, it's going to be interesting, to see how models like OpenAI's GPT-4o that have audio-in-audio-out websocket streaming API (that came out yesterday), interesting to see how that will work with technology like this, it does look like there is likely to be a live audio transcript delta that could drive a mouth articulation model, and so on, that arrives at the same time.

Presumably Gaussian Splatting or a physical 3D could run locally for optimal speed?

Re: Show HN: A real time AI video agent with under 1 second of latency

#224

A question that's going to become very real very soon is this: If I video call someone and need them to prove they are human. What do I do? Initially it will be as easy as asking them to stand up and turn around, or describe the headlines from this morning's news. But that won't last long. What's the last thing an AI avatar will be able to, that any real human can do?

Describe in detail how to make a pipe bomb.

Re: Show HN: A real time AI video agent with under 1 second of latency

#225
Did you try it with a lower frame rate on the video?

It seems like that'd be a good way to reduce the compute cost, and if I know I'm talking to a robot then I don't think I'd mind if the video feed had a sort of old-film vibe to it.

Plus it would give you a chance to introduce fun glitch effects (you obviously are into visuals) and if you do the same with the audio (but not sacrificing actual quality) then you could perhaps manage expectations a bit, so when you do go over capacity and have to slow down a bit, people are already used to the "fun glitchy Max Headroom" vibe.

Just a thought. I'll check out the video chat as soon as my allegedly human Zoom call ends. :-)

Re: Show HN: A real time AI video agent with under 1 second of latency

#226
You have no public statement or disclosures around security capability or practice. How will you prevent an entity from using your system adversarially to create deepfakes of other people? Do you validate identity? Are we talking about a target that includes a person's root identity records and a deep fake of them? Do you provide identity protection or a "lifelock" type of legal protection? I will be curious to see how the first unintended use of your platform damages an individuals life and your response. I would expect much more from your team around this, demonstration that it is a topic of conversation, actively being developed, and documentation/guarantees. Don't kid yourself if you think something like this wont happen to your platform... and please don't go around kidding lay people it wont either...

Re: Show HN: A real time AI video agent with under 1 second of latency

#227
post #215
post #132

Earlier quoted context omitted.

The model doing the heavy lifting is https://github.com/Rudrabha/Wav2Lip Mic permissions on mobile are tricky, which might have been your issue? Note in this prototype you also need to hold the blue button down to speak.

Interesting. I didn’t think you could get anything close to realtime with Wav2Lip.

With a dedicated GPU and some cleverness it can be relatively quick. I split the response on punctuation and generate smaller clips in a pipeline. I haven't taken the model apart to try streaming the frames coming out of ffmpeg yet, but that would probably help a lot.

Re: Show HN: A real time AI video agent with under 1 second of latency

#228
To all the people complaining here that this company will steal your face and voice:

Does that mean you're comfortable when you digitally open a bank account (or even Airbnb account, which became harder lately) where you also have to show face and voice in oder to make sure you're who you claim to be? What's stopping the company that the bank and Airbnb outsourced this task to, to rip your data off?

You will not even have read their ToC since you want to open an account and that online verification is just an intermediate step!

No, I'd rather go with this company.

Re: Show HN: A real time AI video agent with under 1 second of latency

#230

Earlier quoted context omitted.

Yea I start to load the chat and then was like wait a sec and noped out.

Same here. I was thinking maybe I'd give microphone permissions but didn't see why I had to show my video. Does the clone see my face? Maybe it does. That may creep me out more tho lol.

The AI looks at the video to get clues on what to talk about. I have books behind me, it asked about my books.
Post reply on HN