Live data from Hacker News

Megaface

exposing.ai

51–60 of 114 posts

Re: Megaface

#51
post #45
post #35

Earlier quoted context omitted.

"Pictures are clearly personally identifiable data, so storing them violates the GDPR if you don't have permission to do so." Wouldn't a Creative Commons license express this permission?

IANAL, but I believe no; the CC license handles the rights that a photographer can hand out, but doesn't come with any kind of model release guarantees.

Model release is a good point but in many situations where people are photographed it does not apply. When you make photos of yourself or your family or even of people in public spaces, you do not require a model release. And I imagine this to be the main input of this training set.

When you hire a model, photograph the person and then use these photos for promotion or commercial activities, you do require a model release. But in that case it would be absurdly weird to publish such commercial material as CC NC on Flickr, makes no sense.

Re: Megaface

#52

Earlier quoted context omitted.

You can do the scraping in a jurisdiction where it is legal.

> You can do the scraping in a jurisdiction where it is legal. No such thing with GDPR. Why do you think so many US websites take the lazy-ass approach and block EU visitors to their websites ? Simple, its because either you comply with GDPR or you don't process the information of citizens of GDPR covered countries. End of story.

Well, no, only if you’re under the jurisdiction of the EU courts. They can rule against you as much as they like, but it’s not enforceable outside of the EU or a jurisdiction that chooses to enforce EU judgements.

Re: Megaface

#53
post #43
post #39

Earlier quoted context omitted.

I agree, but why would one share wedding photos using an open license like Creative Commons?

Because regulations are generally unknown to folks who don't spend their time solving tech problems. People simply assumed they could share it easily with friends and family.

This doesn't match my experience, and I run a photo community myself.

You're absolutely right that people are generally fairly clueless about licenses, especially in the amateur domain. And the main implication of that is that they don't bother with it at all and leave it at whatever the default is, which typically is "copyrighted, all rights reserved".

Those explicitly tinkering with licenses, which is a purposeful action, tend to actually know (somewhat) what they are doing.

Further, if you leave a photo's license to its default, copyrighted, absolutely nothing stops you from sharing it with friends and family. What would happen? You share it with them and then sue yourself?

Similarly, somebody you don't even know could use your copyrighted image and post it on social media. Again, nothing happens, as this is widespread behavior and called "fair use", which it legally absolutely isn't. But nobody cares, as nobody will sue over it unless there is a case of vast commercial usage.

Re: Megaface

#55

One of the difficulties with these training datasets is the currently understood rules around web scraping. The current legal precedent [0] is that web scraping is perfectly legal, despite what is in the websites terms of service, "licence" or robots.txt. If a human can navigate to it freely, you can scrape it using automated means. What you can't do with scraped data is republish it verbatim. Doing a data analysis o…

A ML model would be considered transformational.

Re: Megaface

#56
Anyone know have any stats on the ethnicities and genera of the people in this dataset?

Is this still widely used to test face recognition?

Re: Megaface

#57

Anyone know have any stats on the ethnicities and genera of the people in this dataset? Is this still widely used to test face recognition?

There is the DiveFace dataset/metadata [1], which is a subset of the Megaface dataset with six equally sized groups: three ethnicities times two genders.

1: https://github.com/BiDAlab/DiveFace

Re: Megaface

#58
post #44

Earlier quoted context omitted.

>if I write a style parody of William Shakespeare and add a '--Willy Shakespeare' at the end to round it off, have I revealed that I have secretly copied his work? // I doubt you're suggesting SD, Dall-E, etc., are producing parodies so bringing in parody considerations muddies the water a lot. Also, Shakespeare's works are out of copyright. If you sell a painting signed with a [facsimile] signature of Dali then it's…

> a painting signed with a [facsimile] signature of Dali That's not what's happening here though. If you look at the original tweet ( https://twitter.com/LaurynIpsum/status/1599953586699767808 ) it seems that the complaint is about the "mangled remains of an artist’s signature". I don't see any examples where it's actually copying the signature of a specific artist. (Please do share an example of that if there is one…

I do respect artist's concerns. I have a hard time getting this one though. The ai learned that humans usually put squiggly lines in the corners, and it does, too. What is wrong with this?

Re: Megaface

#59

Earlier quoted context omitted.

>Or is it republishing of the original data? If it's publishing _data_ then you're fine under regular copyright as it only protects artistic works and not things like data. You might fall shy of other IP legislation but not copyright. YMMV, this is not legal advice and represents my personal opinion unrelated to my employment.

The "data" here is photographs, which all jurisdictions I'm aware of treat as coprightable.

Which makes this case even more interesting to me. Some percentage of those photos’ copywrites’ are owned by corporations rather than the pictured individuals.

If it was simply a large group of selfies, I don’t expect much legal challenge from the allegedly aggrieved. But when companies with legal counsel get involved…

Post reply on HN