How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…
1.3B Worldcat scrape and data science mini-competition
51–60 of 95 posts
Re: 1.3B Worldcat scrape and data science mini-competition
#52Re: 1.3B Worldcat scrape and data science mini-competition
#53Re: 1.3B Worldcat scrape and data science mini-competition
#54It's infuriating seeing non-profits gatekeep datasets that were compiled with grant money. At least Elsevier doesn't present itself as a charity. I was recently trying to get my hands on the Switchboard and Fisher conversational speech datasets. Both were funded by DARPA grants, and maintained by the non-profit LDC, which charges you thousands of dollars for access (and no discounts for individual researchers) - that…
Re: 1.3B Worldcat scrape and data science mini-competition
#55Earlier quoted context omitted.
A second issue is that ISBNs identify a specific SKU (different formats will have different ISBNs, different printings may even get different ISBNs, etc), but book-related projects typically want some way to identify "the same book" across all these different formats, printings, sometimes even editions and translations and collections. OCLC IDs are identifying a different space than ISBNs are.
What you're referring to is sometimes referred to as a "work" vs "edition" https://openlibrary.org/help/faq/editing#work-edition
* Work is defined as the intellectual or artistic content of a distinct creation. It refers to a very abstract idea of a creation e.g. Shakespeare’s Romeo and Juliet and not a specific expression.
* Expression is the intellectual or artistic realization of a work. The realization may take the form of text, sound, image, object, movement, etc., or any combination of such forms.
* Manifestation is the embodiment of an expression of a work. For example a particular edition of a book or a specific music recording.
* Item is a single exemplar of a manifestation. Cataloguing is generally done, based on an item directly available to a cataloguer
Re: 1.3B Worldcat scrape and data science mini-competition
#56Earlier quoted context omitted.
If you liked the comment-length analysis OCLC & want more, there's a whole essay on the subject. [1] >But one of the ironies of the scraping is that it's not going to be immediately helpful to the libraries who are unable to afford to participate in Worldcat. This is because the scrape didn't (and quite possibly never could have) capture the data in MARC format, which is what most library catalog software uses. While…
> While it would have been ideal to get all the data in MARC & as many other formats as possible, I wonder how true this is worldwide - many libraries don't use MARC or have a digital catalog at all. Maybe there are some ways the data could be processed that make it easier to integrate into such places, but of course local needs/desires will vary widely. Indeed, MARC is not universal (and for that matter, it wouldn't…
Re: 1.3B Worldcat scrape and data science mini-competition
#57This is theft.
Re: 1.3B Worldcat scrape and data science mini-competition
#58> We scraped ISBNdb, and downloaded the Open Library dataset, but the results were unsatisfactory. The main problem was that there was not a ton of overlap of ISBNs. What prevents ISBN number collisions between authors? Is there a central authority assigning them, or is there say a national prefix, with each government assigning ISBN's for local publications (perhaps delegating this to another body in that nation). S…
Companies can get a range of ISBNs before deciding what to publish under each ISBN, or whether to publish something at all. So the authorities assigning ISBNs don't necessarily know what they're being used for.
Re: 1.3B Worldcat scrape and data science mini-competition
#59Re: 1.3B Worldcat scrape and data science mini-competition
#60How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…
You realize they're not a business, right?