Live data from Hacker News

SceneScript, a novel approach for 3D scene reconstruction

ai.meta.com

21–30 of 32 posts

Re: SceneScript, a novel approach for 3D scene reconstruction

#21
post #18

Earlier quoted context omitted.

A huge problem is privacy > without upload to a server Doesn't matter. If the app has to render them then it can still send them to the server. If the app does not render directly it can still infer them by doing collision testing against them. If it can't even check for collisions then there isn't much point in having the data in the first place. It gets worse. The more you want AR/MR to understand the scene the mor…

Companies don't spy on you because they don't like you, they spy on you because they make money out of it. Outlaw making money out of personal data (by outlawing targeted ads, data purchase for insurance, and data brokers) and the incentives to spy on you disappear. You can't be probusiness and pro-privacy at the same time though, you need to limit businesses in order to protect the people.

[deleted]

Re: SceneScript, a novel approach for 3D scene reconstruction

#22
post #18

This was a very important missing piece for mixed/augmented and virtual reality. The semantic understanding of the objects around you (without upload to a server). A detailed set of actual meshes of the objects in your living and working spaces. So much can be built on this key building block, to actually help people with their daily lives, low vision people are only one group that will see massive benefits. I also t…

A huge problem is privacy > without upload to a server Doesn't matter. If the app has to render them then it can still send them to the server. If the app does not render directly it can still infer them by doing collision testing against them. If it can't even check for collisions then there isn't much point in having the data in the first place. It gets worse. The more you want AR/MR to understand the scene the mor…

Yes, the camera sees everything. But we could avoid teaching the model certain things. For instance, an undressed body and a dressed body could both be taught as a body. Likewise, medicine pills as well as regular mints could just be taught as mints.

While this approach doesn’t address privacy completely, it avoids certain elements of our life that are considered really private.

I think half of the point of using the simulated dataset to train the model was to safeguard the above mentioned kind of privacy (with the other half being lack of real-world dataset).

Re: SceneScript, a novel approach for 3D scene reconstruction

#24
post #9
post #3

I'm sure the tech is neat. But I get a real sinking feeling from how hard tech giants like "Meta" are trying openly to create dystopia. The "Metaverse" is not aspirational! > That understanding would let AR glasses tailor content to you and your individual context, like seamlessly blending a digital overlay with your physical space Horrifying, if it ever takes off in the way that "Meta" wants it to. Take the modern s…

They still have to solve for: * How uncomfortable the headsets are to wear for long periods if time * How useless they are (what are they solution for?) * How ridiculous they look * How expensive they are * How expensive it is to produce content for After they solve these issues then I will worry about the dystopia we are in. VR headsets may look cool and futuristic in the movies but in reality they are just screens…

Much has been written about the dystopia we are already in, but I am specifically concerned about the dystopia they are earnestly building. Meta is aware of these problems and is trying very hard to solve them.

And they won't be useless! That's the worst part. They will provide compelling value, like smartphones, and as such will be impossible to boycott. At first the value will be simple, obvious utility, like route overlays. Later it will be value within the system, much like how the main utility of smartphones now is to run the mandatory apps to participate in aspects of society that previously did not require a smartphone. And of course they will be addictive.

Here is a short sci-fi film depicting what I fear: https://www.youtube.com/watch?v=YJg02ivYzSs

Re: SceneScript, a novel approach for 3D scene reconstruction

#25
post #18

This was a very important missing piece for mixed/augmented and virtual reality. The semantic understanding of the objects around you (without upload to a server). A detailed set of actual meshes of the objects in your living and working spaces. So much can be built on this key building block, to actually help people with their daily lives, low vision people are only one group that will see massive benefits. I also t…

A huge problem is privacy > without upload to a server Doesn't matter. If the app has to render them then it can still send them to the server. If the app does not render directly it can still infer them by doing collision testing against them. If it can't even check for collisions then there isn't much point in having the data in the first place. It gets worse. The more you want AR/MR to understand the scene the mor…

I think there must be a way of doing this so you just have to trust the OS. The camera feed itself could be essentially private and only pass on locations and types of objects. In this way you could whitelist / blocklist items or categories.

But you would defo have to trust the OS, and in this case Meta…

Re: SceneScript, a novel approach for 3D scene reconstruction

#26

This was a very important missing piece for mixed/augmented and virtual reality. The semantic understanding of the objects around you (without upload to a server). A detailed set of actual meshes of the objects in your living and working spaces. So much can be built on this key building block, to actually help people with their daily lives, low vision people are only one group that will see massive benefits. I also t…

Isn't this more of A semantic understanding of how language maps onto geometry rather than understanding of actual space and geometry?

I sometimes feel like this may not work without language helping it along

Re: SceneScript, a novel approach for 3D scene reconstruction

#27

My issue with this approach is that the synthetic data, and the language used to describe it, necessarily encode a viewpoint - a bias - on the part of the creators. I am firmly behind the need to use language as an intermediary signifier, not just for utility purposes, but because it unlocks the ability to use language as an anchor for arbitrary AR data. My problem with this particular approach is that it requires us…

> My problem with this particular approach is that it requires us all to agree on what things are called

Walk around your home and notice [almost] everything already has a name for it as it comes from some store (unique handmade items are an outlier). Let's just label things we all agree they are called (or give them multiple labels).

Re: SceneScript, a novel approach for 3D scene reconstruction

#28
post #27

My issue with this approach is that the synthetic data, and the language used to describe it, necessarily encode a viewpoint - a bias - on the part of the creators. I am firmly behind the need to use language as an intermediary signifier, not just for utility purposes, but because it unlocks the ability to use language as an anchor for arbitrary AR data. My problem with this particular approach is that it requires us…

> My problem with this particular approach is that it requires us all to agree on what things are called Walk around your home and notice [almost] everything already has a name for it as it comes from some store (unique handmade items are an outlier). Let's just label things we all agree they are called (or give them multiple labels).

I think you overestimate the degree to which your choice of words overlaps with those of others, especially when you try to be more specific than a generic noun by applying modifiers or less-used nouns. What you call a daybed I might call a couch or a sofa or a lounge chair or a chaise or a hassock or a flat couch or a downstairs bed or a family bed … and all this presumes I speak the same flavor of English you do, or that whoever made the data set is familiar with the type of furniture I live around, or how it’s used. The examples are literally limitless, before you even get into the truly subjective things - not just names but opinions about things - that people might want to use to connect digital things to real things. Stuff like ‘larger than average’, ‘wasteful’, ‘beautiful,’ ‘sacred’, ‘durable’, etc etc etc

Believing there is one map from the real world to the world of symbols is solipsism.

Re: SceneScript, a novel approach for 3D scene reconstruction

#29

This was a very important missing piece for mixed/augmented and virtual reality. The semantic understanding of the objects around you (without upload to a server). A detailed set of actual meshes of the objects in your living and working spaces. So much can be built on this key building block, to actually help people with their daily lives, low vision people are only one group that will see massive benefits. I also t…

Isn't this more of A semantic understanding of how language maps onto geometry rather than understanding of actual space and geometry? I sometimes feel like this may not work without language helping it along

With a multimodal llm watching a stream of these objects, should be able answer the question -

"where are my keys?"

It should be able to respond -

"Here on the dining room table" with an astar path drawn to them in the headset.

It's a stream of tokens that you can almost just reverse search on to find the last instance of "keys" in it (thats a simplification but...).

Re: SceneScript, a novel approach for 3D scene reconstruction

#30

This was a very important missing piece for mixed/augmented and virtual reality. The semantic understanding of the objects around you (without upload to a server). A detailed set of actual meshes of the objects in your living and working spaces. So much can be built on this key building block, to actually help people with their daily lives, low vision people are only one group that will see massive benefits. I also t…

Isn't this more of A semantic understanding of how language maps onto geometry rather than understanding of actual space and geometry? I sometimes feel like this may not work without language helping it along

So, potentially how humans turn language into a rendered version of reality? :)
Post reply on HN