Live data from Hacker News

1.3B Worldcat scrape and data science mini-competition

annas-blog.org

21–30 of 95 posts

Re: 1.3B Worldcat scrape and data science mini-competition

#21
post #20

It's unclear what exactly the competition is about. Just to poke around the dataset?

As a note, I wish I had enough space to mirror their library. Looking at this brings out the collector in me...a tendency that I've successfully suppressed. You can only keep so many terabytes of archive around.

Re: 1.3B Worldcat scrape and data science mini-competition

#22
post #16

I looked into using Dewey Decimal for a hobby project. OCLC has a de facto monopoly on it due to the Worldcat database. They're a non-profit, but they're supported by having libraries pay a subscription fee for Worldcat. Back when OCLC was founded, the idea that people would want to have a copy of a card catalog for personal use was laughable, so I'm sympathetic to the people that set up their funding model. It's far…

[dead]

Re: 1.3B Worldcat scrape and data science mini-competition

#23
post #4

Noob Question: Isn't this going to be an great source for training Language models? Is it safe to assume that OpenAI/Google/Meta etc already have these? In any case great work!

If you could somehow download the entire archive you could feed it into your LLM for training. This is a huge corpus and is sort of ill-gotten. That said, it would be pretty awesome.

Google has this sort of thing already, since they have that whole "let's digitize the world's books" project. Interesting as to why google never developed a ChatGPT, given that they literally have a large amount of the world's books digitized.

Re: 1.3B Worldcat scrape and data science mini-competition

#24
post #16

I looked into using Dewey Decimal for a hobby project. OCLC has a de facto monopoly on it due to the Worldcat database. They're a non-profit, but they're supported by having libraries pay a subscription fee for Worldcat. Back when OCLC was founded, the idea that people would want to have a copy of a card catalog for personal use was laughable, so I'm sympathetic to the people that set up their funding model. It's far…

I suppose one way to do it would be to allow patrons of subscriber libraries to access the database dumps and API.

The downside is that would still make it harder than necessary to access and leave some people out. The upside is that it's not that much of change from their existing model. I'm sure there would also be concerns about database dumps being shared publicly, although Anna's Archive has already released their entire database, and I suspect most people who would pay for formal access wouldn't use an authorized copy. Ultimately, I suspect OCLC would still be resistant to this change, as it would feel like a huge shift, even if I'm not sure it would change much from their perspective.

Re: 1.3B Worldcat scrape and data science mini-competition

#25
post #8
post #2

From the end: > We do want to give a genuine shout-out to the Worldcat team. Even though it was a small tragedy that your data was locked up, you did an amazing job at getting 30,000 libraries on board to share their metadata with you. I wonder what the story is behind Worldcat getting so many libraries across the world on board? I don't know much about the software but it must be pretty compelling.

It's not the software per se, which is generally fit for purpose but not amazing, but the traditions and economics underpinning how libraries maintain their bibliographic metadata. Libraries sharing metadata for their catalogs has a long history, dating back to at least 1902 when the Library of Congress started selling catalog cards for use by other libraries. In the 1960s, the Library of Congress embarked on various…

If you liked the comment-length analysis OCLC & want more, there's a whole essay on the subject. [1]

>But one of the ironies of the scraping is that it's not going to be immediately helpful to the libraries who are unable to afford to participate in Worldcat. This is because the scrape didn't (and quite possibly never could have) capture the data in MARC format, which is what most library catalog software uses. While MARC records could be cross-walked from the JSON, they will undoubtedly omit some data elements found in the original MARC.

While it would have been ideal to get all the data in MARC & as many other formats as possible, I wonder how true this is worldwide - many libraries don't use MARC or have a digital catalog at all. Maybe there are some ways the data could be processed that make it easier to integrate into such places, but of course local needs/desires will vary widely.

[1] https://core.ac.uk/download/pdf/11883899.pdf - it was also published in this book: https://archive.org/details/radicalcatalogin0000unse

Re: 1.3B Worldcat scrape and data science mini-competition

#26
post #23
post #4

Noob Question: Isn't this going to be an great source for training Language models? Is it safe to assume that OpenAI/Google/Meta etc already have these? In any case great work!

If you could somehow download the entire archive you could feed it into your LLM for training. This is a huge corpus and is sort of ill-gotten. That said, it would be pretty awesome. Google has this sort of thing already, since they have that whole "let's digitize the world's books" project. Interesting as to why google never developed a ChatGPT, given that they literally have a large amount of the world's books digi…

Google launched Bard earlier this year.

Re: 1.3B Worldcat scrape and data science mini-competition

#27
post #19

Earlier quoted context omitted.

ISBNs are messy. The International ISBN Agency coordinates assigning ISBN ranges to national agencies, who in turn will assign subranges to publishers. The publishers in turn assign specific numbers to their own works. However, the international agency does not itself maintain a universal database of assigned ISBNs - the most it operates is a global database of publishers and their assigned ranges. And since it's the…

Yes. And while in most countries you can't properly publish a book without an ISBN (ie, have it sold in bookshops), you can publish a Kindle book without it (if you opt to only offer the ebook). That leaves a huge part of publications completely out of the system. Kindle-only books are on Amazon servers and nowhere else.

>while in most countries you can't properly publish a book without an ISBN (ie, have it sold in bookshops)

I'm quite skeptical of this, given the amount of books I've personally seen published in recent decades without ISBNs, along with the limited & haphazard attempts to regulate what it means to 'publish' something or even to be a 'proper' bookseller. But if you have some experience I don't with this, I'm interested in hearing about it.

Re: 1.3B Worldcat scrape and data science mini-competition

#28
ISBN is the default ID when it comes to book related projects, yes it is convenient but not without its caveats. The often overlooked fact is ISBN was introduced in late 1960s, so books published prior to that obviously does not have that number; and not all countries adopted ISBN from day one, some like China was on its own catalog systems until 1980s; and bc ISBN are usually centralized managed by govt or commercial agencies, censorship with political or commercial reasons are not uncommon, some books were not able to get published, or may only see the world without an ISBN.

For obvious reasons, older / non-English / suppressed books may be those need more care when it comes to preserving.

Re: 1.3B Worldcat scrape and data science mini-competition

#29

ISBN is the default ID when it comes to book related projects, yes it is convenient but not without its caveats. The often overlooked fact is ISBN was introduced in late 1960s, so books published prior to that obviously does not have that number; and not all countries adopted ISBN from day one, some like China was on its own catalog systems until 1980s; and bc ISBN are usually centralized managed by govt or commercia…

A second issue is that ISBNs identify a specific SKU (different formats will have different ISBNs, different printings may even get different ISBNs, etc), but book-related projects typically want some way to identify "the same book" across all these different formats, printings, sometimes even editions and translations and collections. OCLC IDs are identifying a different space than ISBNs are.

Re: 1.3B Worldcat scrape and data science mini-competition

#30
post #26
post #23

Earlier quoted context omitted.

If you could somehow download the entire archive you could feed it into your LLM for training. This is a huge corpus and is sort of ill-gotten. That said, it would be pretty awesome. Google has this sort of thing already, since they have that whole "let's digitize the world's books" project. Interesting as to why google never developed a ChatGPT, given that they literally have a large amount of the world's books digi…

Google launched Bard earlier this year.

Yes, but why weren't they first?
Post reply on HN