Live data from Hacker News

SHARP, an approach to photorealistic view synthesis from a single image

apple.github.io

71–80 of 114 posts

Re: SHARP, an approach to photorealistic view synthesis from a single image

#71
I note the lack of human portraits in the example cases.

My experience with all these solutions to date (including whatever apple are currently using) is that when viewed stereoscopically the people end up looking like 2d cutouts against the background.

I haven't seen this particular model in use stereoscopically so I can't comment as to its effectiveness, but the lack of a human face in the example set is likely a bit of a tell.

Granted they do call it "Monocular View Synthesis", but i'm unclear as to what its accuracy or real-world use would be if you cant combine 2 views to form a convincing stereo pair.

Re: SHARP, an approach to photorealistic view synthesis from a single image

#73
post #57
post #51

Earlier quoted context omitted.

Until your comment I didn't realise I'd also read it wrong (despite getting the gist of it). Attempted rephrase of the original sentence: Imagine history documentaries where they take an old photo, free objects from the background, and then move them round to give the illusion of parallax.

I'd suggest a different verb like "detach" or "unlink".

isolate from the background?

Re: SHARP, an approach to photorealistic view synthesis from a single image

#74

I note the lack of human portraits in the example cases. My experience with all these solutions to date (including whatever apple are currently using) is that when viewed stereoscopically the people end up looking like 2d cutouts against the background. I haven't seen this particular model in use stereoscopically so I can't comment as to its effectiveness, but the lack of a human face in the example set is likely a b…

They're using their Depth Pro model for depth estimation, and that seems to do faces really well.

https://github.com/apple/ml-depth-pro

https://learnopencv.com/depth-pro-monocular-metric-depth/

Re: SHARP, an approach to photorealistic view synthesis from a single image

#75
post #61

Earlier quoted context omitted.

Early AI „everything turns into dog heads“ vibes. Beautiful.

I miss those. Anyone know if it's still possible to get the models etc. needed to generate them?

I also wanted to generate one of those this year, so I'll camp around here just in case anybody comments on it :)

Re: SHARP, an approach to photorealistic view synthesis from a single image

#76
post #51

Earlier quoted context omitted.

The "free" in this case is a verb. The objects are freed from the background.

Until your comment I didn't realise I'd also read it wrong (despite getting the gist of it). Attempted rephrase of the original sentence: Imagine history documentaries where they take an old photo, free objects from the background, and then move them round to give the illusion of parallax.

Free objects in the background.

Re: SHARP, an approach to photorealistic view synthesis from a single image

#78

I note the lack of human portraits in the example cases. My experience with all these solutions to date (including whatever apple are currently using) is that when viewed stereoscopically the people end up looking like 2d cutouts against the background. I haven't seen this particular model in use stereoscopically so I can't comment as to its effectiveness, but the lack of a human face in the example set is likely a b…

They're using their Depth Pro model for depth estimation, and that seems to do faces really well. https://github.com/apple/ml-depth-pro https://learnopencv.com/depth-pro-monocular-metric-depth/

Im not sure how the depth estimation alone translates into the view synthesis, but the current implementation on-device is definitely not convincing for literally any portrait photographs I have seen.

True stereoscopic captures are convincing statically, but don't provide the parallax.

Re: SHARP, an approach to photorealistic view synthesis from a single image

#79

Earlier quoted context omitted.

Simulation. It takes a lot of effort today to bring up simulations in various fields. 3 D programming is very nontrivial and asset development is extremely expensive. If I have a workspace I can take a photo of and just use it to generate a 3d scene I can then use it in simulations to test ideas out. This is particularly useful in robotics and industrial automation already.

I don't see any examples of a 3D scene information usable for simulation. If you want to simulate something hitting a table, you need the whole table (surface) in space, not just some spatial illusion effect extrapolated from an image of a table. I also think modelling the 3D objects for simulation is the least expensive part of an simulation... the simulation is the expensive thing. I doubt this will be useful for r…

With research like this you need to start somewhere. The fact we can get 3d information helps. There are people looking into making splats capture collision information [1].

I have worked on simulation and in my day job do a lot of simulation. While physics is oftem hard and expensive you only need to write the code once.

Assets? You need to comission 3d artists and then spend hours wrangling file formats. Its extremely tedious. If we could take a photo and extract meshes Im sure we'd have a much easier time.

[1] https://trianglesplatting.github.io/

Re: SHARP, an approach to photorealistic view synthesis from a single image

#80

I understand AI for reasoning, knowledge, etc. I haven't figured out how anyone wants to spend money for this visual and video stuff. It just seems like a bad idea.

Photo apps on phones (can you still call them cameras?) already have a lot of "AI" to enhance photos and videos taken. Some of it is technological necessity, since you're capturing something through a tiny hole, a lot of it is sexying it up to appeal to people, because hey, people would prefer a cinema-quality depiction of their memories rather than the reality...
Post reply on HN