Live data from Hacker News

SPAD: Spatially Aware Multiview Diffusers

yashkant.github.io

31–37 of 37 posts

Re: SPAD: Spatially Aware Multiview Diffusers

#31

I’m confused why there is so much focus on text to images and models. If you spent five minutes talking to anyone with artistic ability, they would tell you that this is not how they generate their work. Making images involves entirely different parts of reasoning than that for speech and language. We seem to be building an entirely faulty model of image generation (outside of things like ControlNet) on the premise t…

Not sure what these cars are all about. Everyone travels by horse and buggy…

We’re building a model optimized for the machine, not people.

Artists can go collect clay to sculpt and flowers to convert to paint. Computers are their own context and should not be romantically anthropomorphized

In the same way fewer and fewer people go to church, fewer and fewer will see the nostalgia in being a data entry worker all day. Society didn’t stop when we all got our first beige box.

Re: SPAD: Spatially Aware Multiview Diffusers

#32
post #31

I’m confused why there is so much focus on text to images and models. If you spent five minutes talking to anyone with artistic ability, they would tell you that this is not how they generate their work. Making images involves entirely different parts of reasoning than that for speech and language. We seem to be building an entirely faulty model of image generation (outside of things like ControlNet) on the premise t…

Not sure what these cars are all about. Everyone travels by horse and buggy… We’re building a model optimized for the machine, not people. Artists can go collect clay to sculpt and flowers to convert to paint. Computers are their own context and should not be romantically anthropomorphized In the same way fewer and fewer people go to church, fewer and fewer will see the nostalgia in being a data entry worker all day.…

This is an incredibly dull, unthinking regurgitation of the nonsense people say online to feel better about their own lack of creative ability. My point wasn’t that computers can’t do the same thing as artists (they already can), it’s that computers won’t achieve the same result by having people describe the images they want to see because that’s fundamentally not how images are made or even perceived.

Re: SPAD: Spatially Aware Multiview Diffusers

#34
post #31

Earlier quoted context omitted.

Not sure what these cars are all about. Everyone travels by horse and buggy… We’re building a model optimized for the machine, not people. Artists can go collect clay to sculpt and flowers to convert to paint. Computers are their own context and should not be romantically anthropomorphized In the same way fewer and fewer people go to church, fewer and fewer will see the nostalgia in being a data entry worker all day.…

This is an incredibly dull, unthinking regurgitation of the nonsense people say online to feel better about their own lack of creative ability. My point wasn’t that computers can’t do the same thing as artists (they already can), it’s that computers won’t achieve the same result by having people describe the images they want to see because that’s fundamentally not how images are made or even perceived.

I play three instruments, draw, sculpt, and used to build houses.

No one ever set a goal for AI to achieve the same result; just replace labor.

Your post is the same dull strawman non-engineers (I also have a BSc in engineering and MSc in math) repeat about AI.

Find me a formal proof of how “images are made” and I’ll show you one possible model of an infinite number of possible models to explain it with a few axiomatic correct twists to the math since all of our symbolic logic is a leaky abstraction that fails to capture how anything is “fundamentally made”.

Pretentious semantic wank is all you’re shipping

Re: SPAD: Spatially Aware Multiview Diffusers

#35

I’m confused why there is so much focus on text to images and models. If you spent five minutes talking to anyone with artistic ability, they would tell you that this is not how they generate their work. Making images involves entirely different parts of reasoning than that for speech and language. We seem to be building an entirely faulty model of image generation (outside of things like ControlNet) on the premise t…

Text prompts aren't an essential part of this technology. They're being used as the interface to generation APIs because it's easy to build, easy to moderate, and for the discord models like Midjourney it's easy for people to copy your work.

With a local model you can find latent space coordinates any way you want and patch the pixel generation model any way you want too. (the above are usually called textual inversion and LoRAs.)

I would personally like to see a system that can input and output layers instead of a single combined image.

Re: SPAD: Spatially Aware Multiview Diffusers

#36
post #17

Earlier quoted context omitted.

My understanding is that while they allow masked textures, the geometry is still fully emitted by Nanite. That means you still need to start with a mesh that has multiple leaves baked into a single polygon plane, as opposed to starting with individual leaf geometry and then that is somehow baked automatically. This illustration from the page you linked to shows that as well: https://cdn2.unrealengine.com/nanite-in-fo…

Yeah alpha masking is inefficient under Nanite, but as they explain further down its handling of dense geometry is good enough that they were able to get away with not using masked materials for the foliage in Fortnite. The individual leaves are modelled as actual geometry and rendered with an opaque, non-masked material. https://cdn2.unrealengine.com/nanite-in-fortnite-chapter-4-t... https://cdn2.unrealengine.com/na…

That 2nd picture looks to me like they switched from Nanite to hand-crafted billboard textures in the 2nd or 3rd LOD.

Re: SPAD: Spatially Aware Multiview Diffusers

#37
post #28

Earlier quoted context omitted.

So how is the desired content communicated to the artist? (Also, reference images can absolutely be used to communicate style to a diffusion model)

The artist is given the content to be illustrated, extrapolates themes and overarching rhetorical or narrative aspects of the work, creates visual representations or metaphors corresponding to these aspects, generates 3-5 interpretations, shows them to the AD, who provides feedback on what has been extrapolated, as well as various design considerations.

"The artist is given the content to be illustrated" and the content is in what form?
Post reply on HN