Live data from Hacker News

Project Aria 'Digital Twin' Dataset by Meta

projectaria.com

41–50 of 94 posts

Re: Project Aria 'Digital Twin' Dataset by Meta

#41

Earlier quoted context omitted.

+1 I’ll “yes and” here…beyond AR/VR a more powerful use case is multi-modal learning (with RL) which is what Meta is probably the leader in IMO. Example paper here: “ Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning” https://arxiv.org/abs/2301.10931 This IMO is the pathway to AGI, as it combines all sense-plan-do data into a time coordinated stream an…

Just to make sure I understand your excitement: we need guinea-pigs ahem people to wear 'head mounted all day ego centric AR' with who knows how many integrated sensors for long stretches on end, so we can finally get to our fabled A.G.i? That is some B.F. Skinner level future we're aiming for--only this time around, humans become the fully surveilled 'teaching machine'.

Well no...not guinea pigs. But correct conceptually - if it's opt-in only and perfectly transparent to everyone what is happening, which in this specific case of Aria it absolutely is.

If we want to make machines with equivalent or better capacity as humans we have to transfer the process for scientific discovery, including the sum of our cognitive capacity and knowledge to them.

If you quantize human adult-infant interactions, then it boils down to Human adults introducing learning trajectories, labeling input data and biasing weights with reinforcing behaviors for new reinforcement agents. If we can re-build the infrastructure to do precisely that, where the agent is in the place of the infant and society is in the place of the "Human Adult" then we will have re-built at scale the process for human development.

The best way we know how to do this today is implementing transfer learning approaches from the basic human developmental research. I started down this road back in 2010 trying to follow the work of Frank Guerin out of the University of Aberdeen [1] [2].

[1]https://www.surrey.ac.uk/people/frank-guerin

[2] https://scholar.google.co.uk/citations?view_op=view_citation...

Re: Project Aria 'Digital Twin' Dataset by Meta

#42

While acknowledging the utility of this kind of object/environment mapping in AR applications - the privacy obliteration is stark.

Yes, and it scares the shit out of me.

Now. Facebook, the company that nobody trusts to do this sort of thing, is going to have to really work hard to demonstrate that they are to be trusted with this data. Apple, and to a lesser extent google, don't.

That cool startup could get away with lots of things, so long as people like the product.

Fortunately for us, AR glasses are limited by power consumption, this means that they can't really do always on realtime streaming of data to the backend for mining. Sure you could have always on mm accurate location, but you can't have video recording at the same time. If you want facial recognition, you'll have to stop the music playing.

Now, what would help is a decent set of privacy laws, ie:

Any cameras smaller than x, must only allow recording of data from persons that expressly allow it, unless in the public domain. People attempting to re-create personally identifiable data from such sensors will be liable to 5 years in jail and or an unlimited fine. (insert carveouts for legitimate research and persons working towards providing evidence for court cases)

This isnt perfect, but its a lot better than what we have now.

Re: Project Aria 'Digital Twin' Dataset by Meta

#43

Earlier quoted context omitted.

+1 I’ll “yes and” here…beyond AR/VR a more powerful use case is multi-modal learning (with RL) which is what Meta is probably the leader in IMO. Example paper here: “ Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning” https://arxiv.org/abs/2301.10931 This IMO is the pathway to AGI, as it combines all sense-plan-do data into a time coordinated stream an…

> a behavior authoring loop A behaviour that authors loops?

I would read it as “a loop that authors behaviors”.

Re: Project Aria 'Digital Twin' Dataset by Meta

#44

Earlier quoted context omitted.

+1 I’ll “yes and” here…beyond AR/VR a more powerful use case is multi-modal learning (with RL) which is what Meta is probably the leader in IMO. Example paper here: “ Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning” https://arxiv.org/abs/2301.10931 This IMO is the pathway to AGI, as it combines all sense-plan-do data into a time coordinated stream an…

Just to make sure I understand your excitement: we need guinea-pigs ahem people to wear 'head mounted all day ego centric AR' with who knows how many integrated sensors for long stretches on end, so we can finally get to our fabled A.G.i? That is some B.F. Skinner level future we're aiming for--only this time around, humans become the fully surveilled 'teaching machine'.

As with most technology, there are plusses and minuses.

If used correctly (if is doing lots of heavy lifting here) this type of system, eye gaze, imu & microphones would provide much much better hearing aids than the current state of the art, at a much cheaper price (go look up the price of hearing aids, its _extortion_ )

Using gate analysis, it would be possible predict when someone is prone to falls, allowing much longer independence for older people.

Assuming that its possible to understand who you are talking to and what they said, you could mitigate and support dementia much more than we can now.

However.

You also have a vast network of headsets with highly accurate always on location, able to see what you are looking at, who you talk to, what you say, and in somecases what you feel about things.

Add in some basic object/facial recognition and you have an authoritarian's wet dream.

now is the time to regulate, but alas, that wont happen.

Re: Project Aria 'Digital Twin' Dataset by Meta

#45

Earlier quoted context omitted.

This is common on many, many sites like this because they do not have any tracking cookies or anything else that they would need consent for, but they're still required to display a cookie banner "notifying" you that cookies are "in use" as per the terms of the old 2009 ePrivacy Directive. In this case, it appears that projectaria.com sets 1) one cookie for the user's DPR (1 or 2) so that the backend can serve optimi…

> but they're still required to display a cookie banner "notifying" you that cookies are "in use" Common misconception but this is not true. If you use cookies only for functional purposes (not for tracking for example), you do not need to show any cookie banners. Like if you have a shopping cart and you have a cookie for keeping track of what's in it, it's for functional purposes for the user and hence needs no noti…

> Common misconception but this is not true. If you use cookies only for functional purposes (not for tracking for example), you do not need to show any cookie banners. Like if you have a shopping cart and you have a cookie for keeping track of what's in it, it's for functional purposes for the user and hence needs no notice to be used.

Personally, I would not put a cookie banner of any kind on my website. However, given this text:

    The term 'strictly necessary' means that such storage of or access to information should be essential, rather than reasonably necessary, for this exemption to apply. However, it will also be restricted to what is essential to provide the service requested by the user, rather than what might be essential for any other uses the service provider might wish to make of that data. It will also include what is required to comply with any other legislation the person using the cookie might be subject to, for example, the security requirements of the seventh data protection principle.

    Where the setting of a cookie is deemed 'important' rather than 'strictly necessary', those collecting the information are still obliged to provide information about the device to the potential service recipient and obtain consent.
I think it's clear why a more risk-conscious organization like Meta might take a more conservative reading of "Strictly necessary" that does not apply to e.g. bandwidth optimizations related to a device's DPI

Re: Project Aria 'Digital Twin' Dataset by Meta

#46
post #29

Earlier quoted context omitted.

Then why does the banner say? >We use cookies to personalise and improve content and services, deliver relevant advertisements and increase the safety of our users

It’s probably the default language for the company. Technical, t he at allows them to have tracking cookies even if they don’t have them now

Deceptive

Re: Project Aria 'Digital Twin' Dataset by Meta

#47
> All sequences within the Aria Digital Twin Dataset have been captured using fully consented researchers in controlled environments in Meta offices.

The requirement and boastful nature of this heading is a frightening tell against the company/industry's perceived practices.

Re: Project Aria 'Digital Twin' Dataset by Meta

#48
post #22

Non-commercial licence :[

I wonder, let's say one uses it on an open source project that accepts donations. Is that non-commercial enough? Or is it a no-no because money is involved? I'm just curious. I'm glad such free datasets exist even if they are restricted to educational/hobby use.

You can't use it in an open-source project. Open source licenses don't restrict commercial use.

Re: Project Aria 'Digital Twin' Dataset by Meta

#49

Earlier quoted context omitted.

Just to make sure I understand your excitement: we need guinea-pigs ahem people to wear 'head mounted all day ego centric AR' with who knows how many integrated sensors for long stretches on end, so we can finally get to our fabled A.G.i? That is some B.F. Skinner level future we're aiming for--only this time around, humans become the fully surveilled 'teaching machine'.

Well no...not guinea pigs. But correct conceptually - if it's opt-in only and perfectly transparent to everyone what is happening, which in this specific case of Aria it absolutely is. If we want to make machines with equivalent or better capacity as humans we have to transfer the process for scientific discovery, including the sum of our cognitive capacity and knowledge to them. If you quantize human adult-infant in…

But what about observer effects? People act differently when recorded, and rarely do we catch humans acting natural when knowingly observed (some of the early 24h/day Twitch streamers come to mind). And what happens once trials are done? How would people feel about their actions becoming part of a technology potentially able to replace them?

Even when this barrier can be overcome (i.e. people become accustomed to wearing these devices), I worry about the opt-in nature of it. We've yet to see a disruptive technology adhering to this principle through-and-through, and if current learning efforts are anything to go by, training data is not something companies want to willingly let go or lose out on.

Taken both, this path has the potential to be quite coercive if no strong guarantees or safeties can be upheld, especially if early exciting trials generate an interest-boom similar to the one we're seeing right now in the LM-space.

Re: Project Aria 'Digital Twin' Dataset by Meta

#50

Earlier quoted context omitted.

Just to make sure I understand your excitement: we need guinea-pigs ahem people to wear 'head mounted all day ego centric AR' with who knows how many integrated sensors for long stretches on end, so we can finally get to our fabled A.G.i? That is some B.F. Skinner level future we're aiming for--only this time around, humans become the fully surveilled 'teaching machine'.

Well no...not guinea pigs. But correct conceptually - if it's opt-in only and perfectly transparent to everyone what is happening, which in this specific case of Aria it absolutely is. If we want to make machines with equivalent or better capacity as humans we have to transfer the process for scientific discovery, including the sum of our cognitive capacity and knowledge to them. If you quantize human adult-infant in…

Is that model (parents giving labeled input and affecting some weights in the child’s head with reinforcement) really a good fit for the reality of how people learn to do things?

It’s my understanding (though I haven’t looked at the primary sources myself) that one of the facts that inspired Chomsky’s language theories and work for instance, was that when you quantify the information communicated by parents to language learning children, there’s actually not very much of it. Not nearly enough to support that what’s going on is anything like the kind of learning embodied by machine learning models.

If that’s true, and there is something of how to act intelligently / humanly already encoded in children (maybe genetically?) and not communicated by this sort of training, wouldn’t ignoring that and trying to get to it purely in this machine learning way be.. at least not at all informed by evidence / examples of it working in nature?

Post reply on HN