Live data from Hacker News

1.3B Worldcat scrape and data science mini-competition

annas-blog.org

51–60 of 95 posts

Re: 1.3B Worldcat scrape and data science mini-competition

#51
post #39

How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…

You realize they're not a business, right?

Re: 1.3B Worldcat scrape and data science mini-competition

#52
post #30
post #26

Earlier quoted context omitted.

Google launched Bard earlier this year.

Yes, but why weren't they first?

It's like asking why wasn't X invented earlier.

Google and everyone else had no idea how successful LLMs could be until OpenAI did it.

Re: 1.3B Worldcat scrape and data science mini-competition

#54

It's infuriating seeing non-profits gatekeep datasets that were compiled with grant money. At least Elsevier doesn't present itself as a charity. I was recently trying to get my hands on the Switchboard and Fisher conversational speech datasets. Both were funded by DARPA grants, and maintained by the non-profit LDC, which charges you thousands of dollars for access (and no discounts for individual researchers) - that…

I hope one day there will a piratebay for datasets (the pile) and ai models (llama)

Re: 1.3B Worldcat scrape and data science mini-competition

#55
post #38
post #29

Earlier quoted context omitted.

A second issue is that ISBNs identify a specific SKU (different formats will have different ISBNs, different printings may even get different ISBNs, etc), but book-related projects typically want some way to identify "the same book" across all these different formats, printings, sometimes even editions and translations and collections. OCLC IDs are identifying a different space than ISBNs are.

What you're referring to is sometimes referred to as a "work" vs "edition" https://openlibrary.org/help/faq/editing#work-edition

It gets much much more complicated than that. There are never ending discussions about FRBR: Functional Requirements for Bibliographic Records. https://www.ifla.org/references/best-practice-for-national-b...

* Work is defined as the intellectual or artistic content of a distinct creation. It refers to a very abstract idea of a creation e.g. Shakespeare’s Romeo and Juliet and not a specific expression.

* Expression is the intellectual or artistic realization of a work. The realization may take the form of text, sound, image, object, movement, etc., or any combination of such forms.

* Manifestation is the embodiment of an expression of a work. For example a particular edition of a book or a specific music recording.

* Item is a single exemplar of a manifestation. Cataloguing is generally done, based on an item directly available to a cataloguer

Re: 1.3B Worldcat scrape and data science mini-competition

#56

Earlier quoted context omitted.

If you liked the comment-length analysis OCLC & want more, there's a whole essay on the subject. [1] >But one of the ironies of the scraping is that it's not going to be immediately helpful to the libraries who are unable to afford to participate in Worldcat. This is because the scrape didn't (and quite possibly never could have) capture the data in MARC format, which is what most library catalog software uses. While…

> While it would have been ideal to get all the data in MARC & as many other formats as possible, I wonder how true this is worldwide - many libraries don't use MARC or have a digital catalog at all. Maybe there are some ways the data could be processed that make it easier to integrate into such places, but of course local needs/desires will vary widely. Indeed, MARC is not universal (and for that matter, it wouldn't…

Worse, definitely worse.

Re: 1.3B Worldcat scrape and data science mini-competition

#58
post #11

> We scraped ISBNdb, and downloaded the Open Library dataset, but the results were unsatisfactory. The main problem was that there was not a ton of overlap of ISBNs. What prevents ISBN number collisions between authors? Is there a central authority assigning them, or is there say a national prefix, with each government assigning ISBN's for local publications (perhaps delegating this to another body in that nation). S…

Companies can get a range of ISBNs before deciding what to publish under each ISBN, or whether to publish something at all. So the authorities assigning ISBNs don't necessarily know what they're being used for.

[deleted]

Re: 1.3B Worldcat scrape and data science mini-competition

#60
post #51
post #39

How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…

You realize they're not a business, right?

Yes, why do you ask?
Post reply on HN