Live data from Hacker News

Unpaywall: An open database of 31,903,705 free scholarly articles

unpaywall.org

51–57 of 57 posts

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#52

I'm gonna stick with scihub to be honest

Yeah, it seems a lot more well-known and I'd rather give my attention to something that isn't a copycat of hard work.

Unpaywall has very little to do with scihub: This isn't about giving access to works that aren't publicly available, but instead about giving links to existing free uploads of the material by the author (i.e. on arxiv or his/her personal website).

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#53
post #46

See also: https://core.ac.uk/

> See also: https://core.ac.uk/ Came here to say that! Also see OpenAlex, soon to be launching: https://docs.openalex.org/

Good news, OpenAlex has launched already in API and snapshot form! Includes most Unpaywall data.

OpenAlex featured in an HN thread a few days ago: . https://news.ycombinator.com/item?id=31271477

Web UI in the works.

Heather COI: cofounder and dev for Unpaywall and OpenAlex

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#54

- "To complete the download of the full dataset, please fill out this form, which helps us report usage back to our funders, and we'll immediately provide you with a download link." https://unpaywall.org/products/snapshot Is that dataset different from this un-gated one? They're both indexes of Crossref DOI's, and they're both 120 million records long. https://www.crossref.org/blog/new-public-data-file-120-milli...

Would not be surprised if the Crossref DB was their starting point.

It is indeed :) (COI: Unpaywall cofounder and dev)

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#55
post #2

Isn’t this what google scholar was supposed to be about? Hope they figure out how to monetize this before the cash runs out because a properly curated, non fire hose kind of source would be great. Particularly if scientific publishing continues to shift to open source vs. Paywalled.

Yes, Unpaywall is self-sustainable as a result of service-level agreements with companies who use the data in their products (specifically Web of Science, Scopus, many others). We also recently got a grant to help with additional development and integration into OpenAlex: https://blog.ourresearch.org/arcadia-2021-grant/

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#56

For some reason, I expected a big search box for articles on the front page. It exists at http://unpaywall.org/articles but the link's hidden away at the very bottom of the page.

the main way most people use it the Chrome or Firefox extension

Re: Unpaywall: An open database of 31,903,705 free scholarly articles

#57
post #50

Earlier quoted context omitted.

Hi, i'm the maintainer of scholar.archive.org. It looks like we are missing a bunch of your public papers, such as those published here: http://park.itc.u-tokyo.ac.jp/eigo/publication_en.html Both unpaywall and scholar.archive.org work best with papers that have persistent identifiers like DOIs, PMIDs, DOAJ article ids, or dblb records. Unpaywall currently works with Crossref DOIs exclusively. With scholar.archive.or…

Many thanks for the reply. Internet Archive Scholar—like everything else the Internet Archive does—is fantastic, and I am very grateful for all the efforts you and your colleagues make. Just for reference for anyone else reading this, here is an excerpt from an e-mail I sent you in March 2021, after IA Scholar was first mentioned on HN: “I contacted the people at [a large Japanese academic library]. ... I showed them…

It is true that some might need to be done manually, but Google Scholar shows that it can be done, with some level of accuracy, via HTML and PDF scraping. PIDs and more formalized metadata make things much easier. But Google Scholar did result in pressure on platforms/publishers/repositories to put at least minimal metadata in HTML meta tags, and this can be machine-extracted. And there is a ton of content and metadata available via OAI-PMH. Neither of these technologies cost anything to publishers on the margin, once they get them implemented, and many have to reap the discovery benefits of large search indices.
Post reply on HN