Live data from Hacker News

Personal and social information of 1.2B people discovered in data leak

dataviper.io

241–250 of 440 posts

Re: Personal and social information of 1.2B people discovered in data leak

#241
post #216
post #165

Earlier quoted context omitted.

I've been using ES off and on since before 1.0 came out. It has always baffled me that ES doesn't require a username and password by default. ES is a database that has to exist on a network to be usable. Heck, it expects that you have multiple nodes, and will complain if you don't. So one of the first things you do is expose it to the network so you can use it. Yes, it takes some serious incompetence to not realize y…

If you set up elasticsearch on a cloud service like AWS, by default your firewall will prevent the outside world from interacting with it, and no authentication is really necessary. If you do use authentication, you probably wouldn't want username+password, you would probably want it to hook into your AWS role manager thing. So to me, username+password seems useful, but it isn't going to be one of the top two most co…

I'd argue that the "pre-cloud" era is still going strong. And that is a good thing. My workplace has it's own data center. There are some downsides, but I prefer it.

So username+password really is needed. And should be included by default.

Also, I'd expect the same of something like MongoDB. That it doesn't have that by default is just baffling.

Re: Personal and social information of 1.2B people discovered in data leak

#242
post #236

Earlier quoted context omitted.

Do you think it’s reasonable to believe your name / address / SSN / DOB / etc is already out there? I’m of the opinion it’s too late for prevention and we need, instead, mitigation.

Exactly. The very reason for existence of the two companies, pdl and oxy, is to tie n pieces of data with m pieces of data. So depending on how the "anonymous" phone number was used, it's plausible that the number can be connected with other PII. In fact I wonder if there is any such thing as non-PII, given the existence of such companies.

Companies need to stop treating knowledge of this information as proof that you are who you say you are. I would have no problem publicly posting my name, social security number, birthday, mother's maiden name, etc., if not for the fact that someone can actually use this information to open a bank account or take out a loan in my name. It's ridiculous that this is all it takes in most cases.

Re: Personal and social information of 1.2B people discovered in data leak

#243
post #234

> Analysis of the “Oxy” database revealed an almost complete scrape of LinkedIn data, including recruiter information. "Oxy" most likely stands for Oxylabs[1], a data mining service by Tesonet[2], which is a parent company of NordVPN. It is probably safe to assume, that LinkedIn was scraped using a residential proxy network, since Oxylabs offers "32M+ 100% anonymous proxies from all around the globe with zero IP bloc…

How is that possible? LinkedIn blocked mining the data this way several years ago. Is it still possible if you pay LinkedIn enough? Or is this old data?

It is strictly impossible to "block mining data" on the public web. Double that if the miner has free access to a pool of residential IPs.

[source: experience]

Re: Personal and social information of 1.2B people discovered in data leak

#244
post #234

> Analysis of the “Oxy” database revealed an almost complete scrape of LinkedIn data, including recruiter information. "Oxy" most likely stands for Oxylabs[1], a data mining service by Tesonet[2], which is a parent company of NordVPN. It is probably safe to assume, that LinkedIn was scraped using a residential proxy network, since Oxylabs offers "32M+ 100% anonymous proxies from all around the globe with zero IP bloc…

The article says it is "Company 2: OxyData.Io (OXY)"* (http://oxydata.io)

Re: Personal and social information of 1.2B people discovered in data leak

#245

Earlier quoted context omitted.

For a while, our Comcast billing account accessed some other person’s account. Comcast didn’t take it seriously, and just told us to create a new account and not use the old one. (!!!) We had full access. I could have signed this person up for the most expensive package, or even canceled their service.

Let's be realistic here. Everyone knows it's not possible to cancel Comcast service.

Nice one. However, I cancelled in person a couple years ago (because I had equipment to return).

The first thing I said at the counter was "I know it's really hard to cancel Comcast, and I'm not going to accept anything but a cancel."

The girl at the counter smiled and said "We know ..." and immediately cancelled my account.

Re: Personal and social information of 1.2B people discovered in data leak

#246
post #165

I was at an Elasticsearch meetup yesterday where we had a good laugh about several similar scandals in Germany recently involving completely unprotected Elasticsearch running on a public IP address without a firewall (e.g. https://www.golem.de/news/elasticsearch-datenleak-bei-conrad... , in German). This beats any of that. Out of the box it does not even bind to a public internet address. Somebody configured this to…

I've been using ES off and on since before 1.0 came out. It has always baffled me that ES doesn't require a username and password by default. ES is a database that has to exist on a network to be usable. Heck, it expects that you have multiple nodes, and will complain if you don't. So one of the first things you do is expose it to the network so you can use it. Yes, it takes some serious incompetence to not realize y…

It's a marketing ploy by ES.

They aggregated the data and published it so that the viral breach would spread their name around because all publicity is good publicity.

Just riffing of course.

Re: Personal and social information of 1.2B people discovered in data leak

#247
post #240

Earlier quoted context omitted.

How is that possible? LinkedIn blocked mining the data this way several years ago. Is it still possible if you pay LinkedIn enough? Or is this old data?

A large number residential proxies and fake LinkedIn accounts would look the same to LinkedIn as normal browsing.

There's information on the leak that wouldn't be widely available without accessing LinkedIn data using their APIs. Phone numbers and emails, for example.

Re: Personal and social information of 1.2B people discovered in data leak

#248
post #202

Earlier quoted context omitted.

LinkedIn doesn't protection doesn't seem to be that sophisticated at the moment. Someone I know maintains ~weekly up-to-date profiles of a few million users via a headless scraper that uses ~10 different premium accounts and a very low number of different IPs.

That is a violation of ToS (using registerd accounts for scrape) and could carry potential legal implications.

So is leaking PII? ToS isn't a legal contract: it's not signed by anyone and it's changed every other week without consent of users. ToS is just a formal excuse why someone's account may be suspended.

Re: Personal and social information of 1.2B people discovered in data leak

#249
post #136

Earlier quoted context omitted.

Distributed bot and scraper networks. Thousands of IPs geographically dispersed throughout the world. There is only so much you can do with rate limiting.

They asked about LinkedIn, where the content is gated behind a login. If it was a rate limiting problem, that would be trivial. Needing to be logged in as the same user defeats the purpose of proxying to hide your physical origin. Registering thousands of different users to use in a distributed way is hard now that they require a text message verification for new accounts.

Public LinkedIn profiles (which is many of them) are open to scrapers and they lost a court case about it.

https://www.eff.org/deeplinks/2019/09/victory-ruling-hiq-v-l...

Re: Personal and social information of 1.2B people discovered in data leak

#250
post #240

Earlier quoted context omitted.

A large number residential proxies and fake LinkedIn accounts would look the same to LinkedIn as normal browsing.

There's information on the leak that wouldn't be widely available without accessing LinkedIn data using their APIs. Phone numbers and emails, for example.

The article mentions it is a blend of data from http://oxydata.io/ and https://www.peopledatalabs.com/

Both are aggregators that get data from many sources, correlate them, and sell it. The phone numbers and emails could have come from anywhere.

See this screenshot from PeopleDataLabs: https://d1ennknj6q36vm.cloudfront.net/images/cblead.png

Post reply on HN