Live data from Hacker News

1.3B Worldcat scrape and data science mini-competition

annas-blog.org

41–50 of 95 posts

Re: 1.3B Worldcat scrape and data science mini-competition

#41
post #37
post #16

I looked into using Dewey Decimal for a hobby project. OCLC has a de facto monopoly on it due to the Worldcat database. They're a non-profit, but they're supported by having libraries pay a subscription fee for Worldcat. Back when OCLC was founded, the idea that people would want to have a copy of a card catalog for personal use was laughable, so I'm sympathetic to the people that set up their funding model. It's far…

I have been told that my organization developed a system, HABS, that pre-dated OCLC [0]. That OCLC used this system as an inspiration. However, I cannot confirm this. Closest I can do is to find a footnote that Thanks Fred Kilgore, the founder of OCLC [1]. I should reach out to Koh, a friend of a friend, while she is still alive to confirm the story. Nevertheless, we have a collection of punch cards in a dusty room i…

Neat! I hope you can learn more from Koh.

I know a bit about Henriette Avram and her work at LoC developing MARC, but it of course makes sense that other libraries were thinking along similar lines at the time.

Re: 1.3B Worldcat scrape and data science mini-competition

#42
post #40
post #36

Earlier quoted context omitted.

It was true, for sure. But the question is, is it still true?

Schools often ask for a specific edition of a classic book and the only way to be reasonably sure you're buying the correct one is to search by ISBN.

Well, just copy and paste from the email you got from school.

I mean, if IT can do anything, it is to solve this problem.

Re: 1.3B Worldcat scrape and data science mini-competition

#43
post #8

Earlier quoted context omitted.

It's not the software per se, which is generally fit for purpose but not amazing, but the traditions and economics underpinning how libraries maintain their bibliographic metadata. Libraries sharing metadata for their catalogs has a long history, dating back to at least 1902 when the Library of Congress started selling catalog cards for use by other libraries. In the 1960s, the Library of Congress embarked on various…

If you liked the comment-length analysis OCLC & want more, there's a whole essay on the subject. [1] >But one of the ironies of the scraping is that it's not going to be immediately helpful to the libraries who are unable to afford to participate in Worldcat. This is because the scrape didn't (and quite possibly never could have) capture the data in MARC format, which is what most library catalog software uses. While…

> While it would have been ideal to get all the data in MARC & as many other formats as possible, I wonder how true this is worldwide - many libraries don't use MARC or have a digital catalog at all. Maybe there are some ways the data could be processed that make it easier to integrate into such places, but of course local needs/desires will vary widely.

Indeed, MARC is not universal (and for that matter, it wouldn't surprise me if at this point the majority of records in Worldcat were _not_ derived from MARC sources), and there are certainly non-MARC library catalog platforms out there. That said, as the growth of Koha shows, for better or worse MARC has become a close to a global baseline for a lot of libraries.

Re: 1.3B Worldcat scrape and data science mini-competition

#44
post #16

I looked into using Dewey Decimal for a hobby project. OCLC has a de facto monopoly on it due to the Worldcat database. They're a non-profit, but they're supported by having libraries pay a subscription fee for Worldcat. Back when OCLC was founded, the idea that people would want to have a copy of a card catalog for personal use was laughable, so I'm sympathetic to the people that set up their funding model. It's far…

OCLC is a nonprofit membership cooperative and would argue that it itself is that international coalition of national libraries and archives.

Re: 1.3B Worldcat scrape and data science mini-competition

#45
post #16

I looked into using Dewey Decimal for a hobby project. OCLC has a de facto monopoly on it due to the Worldcat database. They're a non-profit, but they're supported by having libraries pay a subscription fee for Worldcat. Back when OCLC was founded, the idea that people would want to have a copy of a card catalog for personal use was laughable, so I'm sympathetic to the people that set up their funding model. It's far…

OCLC is a nonprofit membership cooperative and would argue that it itself is that international coalition of national libraries and archives.

OCLC is a parasitic company masquerading as a "membership cooperative".

Libraries (often publicly funded) produce all the work, OCLC claims ownership of the results of that work, Libraries pay to get it back (but they do get a discount if they contribute).

The only reason OCLC continues to exist is because libraries don't have the support or resources to fight them. It's very similar to the Elsevier issue in academic publishing, but OCLC does a better job with PR.

Re: 1.3B Worldcat scrape and data science mini-competition

#46
post #39

How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…

For real. Openly bragging about exploiting security flaws to scrape out data en-masse, which undoubtedly put massive strain on back-end systems, is a far cry from what is considered legal (politely scraping public information).

Re: 1.3B Worldcat scrape and data science mini-competition

#47
post #39

How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…

For real. Openly bragging about exploiting security flaws to scrape out data en-masse, which undoubtedly put massive strain on back-end systems, is a far cry from what is considered legal (politely scraping public information).

I think this is the least of their concerns considering the rest of their activities. I guess they've got a sort of pirate's privilege in that they can openly brag about this stuff since they're already starting from the point of openly flaunting the law.

Also, I wouldn't be surprised if there simply are no lawyers working at, for or with Anna's Archive.

Re: 1.3B Worldcat scrape and data science mini-competition

#48

>Over the past year, we’ve meticulously scraped all Worldcat records. At first, we hit a lucky break. Worldcat was just rolling out their complete website redesign (in Aug 2022). This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (!) of records in mere days. >After that, security flaws were sl…

I don't think anyone is (legally) going to prop up a business or non-profit using data that was admittedly taken from them using their security holes.

Re: 1.3B Worldcat scrape and data science mini-competition

#49
It's infuriating seeing non-profits gatekeep datasets that were compiled with grant money. At least Elsevier doesn't present itself as a charity.

I was recently trying to get my hands on the Switchboard and Fisher conversational speech datasets. Both were funded by DARPA grants, and maintained by the non-profit LDC, which charges you thousands of dollars for access (and no discounts for individual researchers) - that is, if they'll even pay attention to you without a .edu email address. And both are standard corpora in the field of audio NLP, which makes replicating studies impossible.

Sadly, I couldn't find any way to pirate the datasets - they're too niche. So I applaud the authors for sticking it to Worldcat and scraping their data.

Re: 1.3B Worldcat scrape and data science mini-competition

#50

Earlier quoted context omitted.

OCLC is a nonprofit membership cooperative and would argue that it itself is that international coalition of national libraries and archives.

OCLC is a parasitic company masquerading as a "membership cooperative". Libraries (often publicly funded) produce all the work, OCLC claims ownership of the results of that work, Libraries pay to get it back (but they do get a discount if they contribute). The only reason OCLC continues to exist is because libraries don't have the support or resources to fight them. It's very similar to the Elsevier issue in academic…

I mean you’ve clearly read Aaron Swartz’s diatribes, but you also clearly have no clue about OCLC’s business model. The catalog data is intellectually interesting, but the value is in the holdings data and more importantly the interlibrary loan service it enables.

OCLC is exactly what happens when the libraries want to avoid another EBSCOhost or Proquest situation with ILL.

Post reply on HN