Live data from Hacker News

A raw dump of companies from all over the world by LinkedIn handle

blog.bigpicture.io

101–108 of 108 posts

Re: A raw dump of companies from all over the world by LinkedIn handle

#102
post #48
post #41

Earlier quoted context omitted.

> In any case, web scraping is a sort of gray area of the law. I don’t think it’s grey. It seems to be legal as the data are made freely available and the only grey part is that companies don’t want this to happen and would rather charge and not have people scrape.

It is a gray area (in the US) in the sense that there is no clear consensus about it in the courts. There have been court rulings in both directions.

http get requests are legal.

Re: A raw dump of companies from all over the world by LinkedIn handle

#103

It's funny how OP does not address where this data comes from even though it's obviously from LinkedIn. I see many people in the comments asking questions so I will add my two cents as someone who is currently employed by LinkedIn and has an interest in web scraping. This dataset was taken from scraping the company pages from LinkedIn. A company has to pay to have this page, so this certainly does not include all com…

[dead]

Re: A raw dump of companies from all over the world by LinkedIn handle

#104
post #92
post #66

I always get "Oops! We ran into an error. Contact us at support@bigpicture.io" when I try to sign-up

Apologies. We've had a huge problem with bots, so we have a number of security measures in place. The Google Captcha component is probably flagging you as a bot. Try disabling your VPN if you're on one, or use a different IP.

Oh the irony.

Re: A raw dump of companies from all over the world by LinkedIn handle

#106
post #47

You use "open source" multiple times in the post, HN title, HN comments, but: 1. The source code for the project isn't shared anywhere. 2. The data isn't shared under any standard open source license. 3. The terms of your site explicitly prohibit commercial use of this data. So what exactly makes this "open source in the broadest sense"?

It's open source in the sense of OSINT [0]. Clearly confusing on a site like Hacker News, but this has been standard usage of the term for that community for a long time now. 0. https://en.wikipedia.org/wiki/Open-source_intelligence

Thank-you for the clarification. "Open-source" is definitely different from "Open Source."

Meaning Open-source (sourced from open sources) but claiming Open Source is disingenuous.

Wikipedia isn't helpful either, because it refers to OSS as Open-source Software [1].

Open Source meaning may be more useful in comparison to Free Software. Stallman refers to "Open-source" (hyphenated) only once in this article, but only to refer to it as confusing versus free software [2].

It's possible "OSINT as Open-source" has been in use for longer than Stallman's use of "open source," but definitely they are different.

It's strange a site would sell up a feature on HN as "Open-source content, in the meaning of OSINT" without being up-front about it. The default assumption would be "open source as code that is free to modify, etc."

The mental gymnastics would be

  1. They claim it is "open source."
  2. They are talking about _content._
  3. It must be the OSINT kind of "open."
This could be a pattern, because they're always needing to add another comment, "Just kidding, we meant OSINT open; we're not sharing the code."

... documentation could be open source too, though--in the sense of "free to modify, etc" and not "sourced from freely available data."

Could it be both? Only if they accept contributions, I guess.

[1] https://en.m.wikipedia.org/wiki/Open-source_software

[2] https://www.gnu.org/philosophy/open-source-misses-the-point....

Re: A raw dump of companies from all over the world by LinkedIn handle

#108
post #41

It's funny how OP does not address where this data comes from even though it's obviously from LinkedIn. I see many people in the comments asking questions so I will add my two cents as someone who is currently employed by LinkedIn and has an interest in web scraping. This dataset was taken from scraping the company pages from LinkedIn. A company has to pay to have this page, so this certainly does not include all com…

> In any case, web scraping is a sort of gray area of the law. I don’t think it’s grey. It seems to be legal as the data are made freely available and the only grey part is that companies don’t want this to happen and would rather charge and not have people scrape.

It reminds me of a similar conflict in public records.

Many municipalities charge for copies of public records. One can go to viewing rooms and examine the records at no charge. We all own those records as members of the public that funded them, and the place where they are kept.

Many municipalities want to prohibit photography because people taking their own picture of a public record does not involve the copy fee.

The municipality confuses access to records as a part of their fee, and that is where attempts to limit photography come from.

The fee actually funds the work necessary for an official copy to be made, optionally stamped to be admissible in court or accepted as an "original copy."

A photocopy of a death certificate is no different than the one the city clerk made, except for the stamp and clerk being able to testify about making the record copy.

Some companies and government agencies will accept a death cert no matter what. Others want an official copy.

Getting back to scraping:

Clearly people have access and can make their own data copies.

Maybe the answer is for companies to make official data products available, or something along those lines.

Post reply on HN