Live data from Hacker News

1.3B Worldcat scrape and data science mini-competition

annas-blog.org

81–90 of 95 posts

Re: 1.3B Worldcat scrape and data science mini-competition

#81
post #19

Earlier quoted context omitted.

Yes. And while in most countries you can't properly publish a book without an ISBN (ie, have it sold in bookshops), you can publish a Kindle book without it (if you opt to only offer the ebook). That leaves a huge part of publications completely out of the system. Kindle-only books are on Amazon servers and nowhere else.

>while in most countries you can't properly publish a book without an ISBN (ie, have it sold in bookshops) I'm quite skeptical of this, given the amount of books I've personally seen published in recent decades without ISBNs, along with the limited & haphazard attempts to regulate what it means to 'publish' something or even to be a 'proper' bookseller. But if you have some experience I don't with this, I'm intereste…

bookstores don't want to carry a book without an isbn because no isbn means it's not available to purchase from their distributor and it's easiest for a store to order through real distribution channels (like ingram in the u.s.).

but most stores carry a small amount of self-published books and sometimes those books have no isbn. those books are typically by local writers. but in my experience as a bookseller, self published books are a pain to work with. some self-published books aren't returnable, but returns are an important part of the bookstore business since a lot of books don't sell. working with a lot writers individually about ordering etc is more involved than going through a single distributor, this takes a lot of time for whoever has to do this.

> given the amount of books I've personally seen published in recent decades without ISBNs

i'd bet this on amazon? iirc, you can't always return self-published amazon books. i think the author decides this, bc they get charged a processing fee for returns.

> or even to be a 'proper' bookseller

you can totally sell your collection online without any isbns and you'd be considered a bookseller. you'll just need an sku system. there's a difference between a used/collectible seller and a bookseller who carries new books as well as used/collectibles. the new books require proper distribution channels.

Re: 1.3B Worldcat scrape and data science mini-competition

#82
post #39

How does Anna's Archive keep their all their lawyers from quitting? > Even though OCLC is a non-profit, their business model requires protecting their database. Well, we’re sorry to say, friends at OCLC, we’re giving it all away. :-) [...] > This included a substantial overhaul of their backend systems, introducing many security flaws. We immediately seized the opportunity, and were able scrape hundreds of millions (…

What lawyers?

Re: 1.3B Worldcat scrape and data science mini-competition

#84

Question for anyone from Anna's Archive or elsewhere: are catalogue metadata available from national library collections such as the British Library or US Library of Congress? (I've ... worked a bit with LoC classification and subject headings data, of which publicly-available data are only available in PDF or wordprocessing (MS Word or Wordperfect, if memory serves) formats. Which is ... somewhat unfortunate.)

Depending on what you're looking for, a _lot_ more is being published by the Library of Congress as Linked Data nowadays, including LoC classification and subject headings. Check out https://id.loc.gov/.

Re: 1.3B Worldcat scrape and data science mini-competition

#85

Question for anyone from Anna's Archive or elsewhere: are catalogue metadata available from national library collections such as the British Library or US Library of Congress? (I've ... worked a bit with LoC classification and subject headings data, of which publicly-available data are only available in PDF or wordprocessing (MS Word or Wordperfect, if memory serves) formats. Which is ... somewhat unfortunate.)

Depending on what you're looking for, a _lot_ more is being published by the Library of Congress as Linked Data nowadays, including LoC classification and subject headings. Check out https://id.loc.gov/ .

That seems to be just metadata about metadata.

Where do you go to actually find the Classifications or Subject Headings themselves, in whatever format?

Re: 1.3B Worldcat scrape and data science mini-competition

#86

Earlier quoted context omitted.

For real. Openly bragging about exploiting security flaws to scrape out data en-masse, which undoubtedly put massive strain on back-end systems, is a far cry from what is considered legal (politely scraping public information).

I think this is the least of their concerns considering the rest of their activities. I guess they've got a sort of pirate's privilege in that they can openly brag about this stuff since they're already starting from the point of openly flaunting the law. Also, I wouldn't be surprised if there simply are no lawyers working at, for or with Anna's Archive.

flaunting-> flouting

Re: 1.3B Worldcat scrape and data science mini-competition

#87

Earlier quoted context omitted.

I think this is the least of their concerns considering the rest of their activities. I guess they've got a sort of pirate's privilege in that they can openly brag about this stuff since they're already starting from the point of openly flaunting the law. Also, I wouldn't be surprised if there simply are no lawyers working at, for or with Anna's Archive.

Scraping and giving away the content is a better look than scraping and selling it as well. This is being done for public benefit, not private profit.

I'm not commenting on the morality, only the legality.

Re: 1.3B Worldcat scrape and data science mini-competition

#88

Earlier quoted context omitted.

I think this is the least of their concerns considering the rest of their activities. I guess they've got a sort of pirate's privilege in that they can openly brag about this stuff since they're already starting from the point of openly flaunting the law. Also, I wouldn't be surprised if there simply are no lawyers working at, for or with Anna's Archive.

flaunting-> flouting

Nice catch, thanks. I suppose they're flaunting their flouting of the law ;)

Re: 1.3B Worldcat scrape and data science mini-competition

#89

Earlier quoted context omitted.

Scraping and giving away the content is a better look than scraping and selling it as well. This is being done for public benefit, not private profit.

I'm not commenting on the morality, only the legality.

The legality is itself a function of commercial and political power amongst publishers.

In front of a jury of peers, the moral arguments might well be persuasive.

Re: 1.3B Worldcat scrape and data science mini-competition

#90

Earlier quoted context omitted.

Depending on what you're looking for, a _lot_ more is being published by the Library of Congress as Linked Data nowadays, including LoC classification and subject headings. Check out https://id.loc.gov/ .

That seems to be just metadata about metadata. Where do you go to actually find the Classifications or Subject Headings themselves, in whatever format?

https://id.loc.gov/download/ offers bulk downloads of many of the vocabularies, including the subject headings. In addition, individual records of various types can be search for.
Post reply on HN