Live data from Hacker News

Unsupervised learning of probably symmetric deformable 3D objects from images

robots.ox.ac.uk

11–20 of 39 posts

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#12

Video here: https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?

Because reading from a script for five minutes is likely to require multiple takes for someone who isn't a practiced voice actor, while text to speech requires no extra effort on their part?

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#13
post #3

Stuff like this can obviously be used to make things like deepfakes 'better'. But i think it might be cool for creating virtual meeting rooms, where you can take a webcam shot of a persons face, normalize the skin tones for lighting, map to a 3d surface, then relight it for the virtual room. When you can rig the meshes to drive each other you could wear 'masks' of other peoples faces (or critters).

> But i think it might be cool for creating virtual meeting rooms, where you can take a webcam shot of a persons face, normalize the skin tones for lighting, map to a 3d surface, then relight it for the virtual room. You might be interested in project HeadOn at TU München: https://www.niessnerlab.org/projects/thies2018headon.html Justus Thies gave a presentation at our University about a year ago. IIRC they don't use…

Oh man, this plus the obs thread earlier plus zoom. The mind boggles haha.

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#14
post #11

Please fix the title of submission. It is " Unsupervised Learning of Probably Symmetric Deformable 3D Objects from Images in the Wild". No anime examples to be found in the paper :(

Figure 11, "Reconstruction on abstract face drawings", has a reconstruction of Naruto's face. The results are... unsettling.

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#16

Video here: https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?

None of the authors appear to be native English speakers, so perhaps they're self-conscious about their accents?

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#17

Video here: https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?

Because reading from a script for five minutes is likely to require multiple takes for someone who isn't a practiced voice actor, while text to speech requires no extra effort on their part?

> Because reading from a script for five minutes is likely to require multiple takes for someone who isn't a practiced voice actor

This depends on how much you can tolerate speech errors. Most listeners will gloss over them, preferring the human voice to the speech synthesizer while not even really noticing the errors.

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#18
post #16

Video here: https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?

None of the authors appear to be native English speakers, so perhaps they're self-conscious about their accents?

My thought as well. The TTS is good enough that it won't take much of an accent before the accent is harder to understand than the TTS as well. I know my own accent is strong enough that I'd have to put in very conscious effort to be easier to understand than this video.

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#19
post #16

Video here: https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?

None of the authors appear to be native English speakers, so perhaps they're self-conscious about their accents?

Could also have speech problems. Could be lazy. Could want to save time. Could be useful at producing consistent CC information across mediums. Could allow people to choose arbitrary voice synthesis in the future which super futurists may like the idea of. Could have used a translator to produce the text (I haven't listened) and not know English atall.

Personally, I'll take the human voice unless you literally cannot speak (e.g. disability) or feel uncomfortable.

Re: Unsupervised learning of probably symmetric deformable 3D objects from images

#20
post #19
post #16

Earlier quoted context omitted.

None of the authors appear to be native English speakers, so perhaps they're self-conscious about their accents?

Could also have speech problems. Could be lazy. Could want to save time. Could be useful at producing consistent CC information across mediums. Could allow people to choose arbitrary voice synthesis in the future which super futurists may like the idea of. Could have used a translator to produce the text (I haven't listened) and not know English atall. Personally, I'll take the human voice unless you literally cannot…

Many academics like the ability to "compile" latex, and probably want to "compile" their videos too, complete with autogenerated script. That way, when they make a small change to their source code, the new version will autogenerate a new video with an updated script.
Post reply on HN