- "To complete the download of the full dataset, please fill out this form, which helps us report usage back to our funders, and we'll immediately provide you with a download link." https://unpaywall.org/products/snapshot Is that dataset different from this un-gated one? They're both indexes of Crossref DOI's, and they're both 120 million records long. https://www.crossref.org/blog/new-public-data-file-120-milli...
Unpaywall: An open database of 31,903,705 free scholarly articles
41–50 of 57 posts
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#42If a university pay walls a single paper or article they produce from the public they should immediately be stripped of all of their public funding. I can't think of a single reason why the public is forced to pay taxes to them if the only way they are able to see a benefit from it is by proxy of someone who is paying money to attend or someone spending their private money to access the information.
In the US, federally funded research is already required to made available without paywalls within 1 year of publication. Open-access advocates online generally seem to be unaware of this, and that current goalposts are therefore (a) immediate open access, and (b) open-access publication of all scientific research regardless of funding source. https://obamawhitehouse.archives.gov/blog/2013/02/22/expandi... https://ww…
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#43Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#44https://oa.mg/ is another initiative that indexes over 200 million papers and lets you know if a paper is open or not. They tend to have download links to a rather large amount of papers.
It would be great to see Zotero integrate these databases into a PDF search feature, or a plug-in
https://www.zotero.org/blog/improved-pdf-retrieval-with-unpa...
(Disclosure: Zotero dev)
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#45Hot take: Google Scholar is content collected means of donated human and robot indexers. It is not a database or service until they have an API that allows us to harvest. Yes I am salty af about it. It is taking energy away from curating ORCID records and the like.
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#46See also: https://core.ac.uk/
Came here to say that! Also see OpenAlex, soon to be launching:
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#47The Internet Archive has a similar project in beta: https://scholar.archive.org I have about ten academic papers that are accessible online for free, so I tried a vanity search. The Internet Archive has indexed one of those papers; Unpaywall has none.
It looks like we are missing a bunch of your public papers, such as those published here: http://park.itc.u-tokyo.ac.jp/eigo/publication_en.html
Both unpaywall and scholar.archive.org work best with papers that have persistent identifiers like DOIs, PMIDs, DOAJ article ids, or dblb records. Unpaywall currently works with Crossref DOIs exclusively.
With scholar.archive.org (and fatcat.wiki, the backing catalog), it is possible to submit metadata records directly, but it can be laborious and would be better for everybody if this process was automated.
Processing OAI-PMH feeds or extracting bibliographic metadata from HTML metadata would probably improve our coverage, and we are hoping to roll out that kind of scraping eventually. But it has been a challenge to clean and de-duplicate metadata at that scale.
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#48- "To complete the download of the full dataset, please fill out this form, which helps us report usage back to our funders, and we'll immediately provide you with a download link." https://unpaywall.org/products/snapshot Is that dataset different from this un-gated one? They're both indexes of Crossref DOI's, and they're both 120 million records long. https://www.crossref.org/blog/new-public-data-file-120-milli...
Unpaywall will both check if articles are actually available from the publisher (by following the DOI and parsing the landing page), and by looking for other versions elsewhere on the web (eg, a pre-print). It is simple in theory, but doing this reliably for millions of DOIs from thousands of publishers is a lot of work!
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#49Earlier quoted context omitted.
It would be great to see Zotero integrate these databases into a PDF search feature, or a plug-in
Zotero has had Unpaywall integration since 2018: https://www.zotero.org/blog/improved-pdf-retrieval-with-unpa... (Disclosure: Zotero dev)
Also: thank you for building a tool that actually works. I used it first in 2010, then switched to EndNote for better web integration, but I’m back to Zotero now that there’s a Safari add on. The PDF functionality in 5.0 has made apps like Notability and Goodnotes for the seemingly trivial task of managing and marking up highlights completely redundant.
It’s one of those OSS applications that genuinely works well and can be approachable to everyday people without feeling like a half-baked product.
Re: Unpaywall: An open database of 31,903,705 free scholarly articles
#50The Internet Archive has a similar project in beta: https://scholar.archive.org I have about ten academic papers that are accessible online for free, so I tried a vanity search. The Internet Archive has indexed one of those papers; Unpaywall has none.
Hi, i'm the maintainer of scholar.archive.org. It looks like we are missing a bunch of your public papers, such as those published here: http://park.itc.u-tokyo.ac.jp/eigo/publication_en.html Both unpaywall and scholar.archive.org work best with papers that have persistent identifiers like DOIs, PMIDs, DOAJ article ids, or dblb records. Unpaywall currently works with Crossref DOIs exclusively. With scholar.archive.or…
Just for reference for anyone else reading this, here is an excerpt from an e-mail I sent you in March 2021, after IA Scholar was first mentioned on HN:
“I contacted the people at [a large Japanese academic library]. ... I showed them your HN post [1] about the data you've already collected through J-STAGE, and, contrary to my own impression, they said you have probably already captured most of the metadata for Japanese academic journals that would be easily available. They also pointed out that J-STAGE includes a fair amount of publications from the humanities side of things, also contrary to my own impression.
“The main sticking point, they said, is journals that are published by universities or academic societies and have not been listed on J-STAGE. Many of those journals have never been digitized, they said, and those that are available in digital form are likely to be available only on those universities’ or societies’ individual websites. The library people didn’t know of any aggregators or indexes for such sites. The only way to find them, they suggested, would be for someone to hunt for the sites by hand.
“Over the years, I myself have been involved with the publication of several such journals and have set up websites for a couple, too. The ones published by departments at [a particular Japanese university] are included in [the university’s online repository] but not yet, it seems, on J-STAGE. A couple published by small academic societies are available only on those societies' websites. [Addendum: The Japanese academic societies I have been involved with—mostly in the humanities—would have difficulty getting DOIs or other persistent identifies for the papers they publish; it would take some effort even to convince them of the necessity. They are volunteer-run organizations, and just maintaining their websites is often a challenge for them.]
“Yet another impression of mine (also perhaps wrong) is that a higher percentage of academic research in Japan is published through such journals than in the U.S. It would be very valuable to have that research findable through IA Scholar, but the barriers to collecting it seem high.”