Live data from Hacker News

Dozens of scientific journals have vanished from the internet

sciencemag.org

121–130 of 136 posts

Re: Dozens of scientific journals have vanished from the internet

#121
post #98

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> If folks want to help, it would be great to have a "youtube-dl for open access papers". I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.

Unpaywall is very helpful! However, even for direct PDF links, publishing platforms will often do things like check for a session cookie; if you don't have the correct cookie you get bounced back to the landing page, where you need find and follow another link. This isn't super complicated to work around (persist a cookie jar, use a headless browser, etc), but it doesn't work out-of-the-box with our crawlers, the same way crawling youtube doesn't work out-of-the-box so we rely on the youtube-dl community.

Re: Dozens of scientific journals have vanished from the internet

#122

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring. Most of orgs/journals/conferences in this group just don't have the resources to be able do this, nor maintain something like this. It was funny last year when…

From what I have seen, the least technically resourced journals often use hosted platforms or free software like OJS (basically wordpress for journals), which comes with features like HTML meta tags and OAI-PMH by default.

The trickier cases are when folks write their own platforms, or even write their own raw HTML with no templating, in which case adding tags to all landing pages or supporting an API would be a relatively large amount of work.

Re: Dozens of scientific journals have vanished from the internet

#123
post #98

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> If folks want to help, it would be great to have a "youtube-dl for open access papers". I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.

Yes, the Internet Archive already used those URLs, or tried to. Bryan was so kind as to share some statistics about it: https://groups.google.com/g/unpaywall/c/AbNwXdyWZfE/m/OVoVmx...

Having an URL to something that is supposed to serve a PDF doesn't mean that you actually get it. Publishers like Elsevier or Wiley require a cookie/JavaScript dance and/or place heavy rate limits; whatever PDF hosted by them is never really accessible, even if it's ostensibly open access.

The "youtube-dl for papers" would do what youtube-dl does, i.e. take the URL and execute whatever JavaScript or other trickery the target page/website requires before it serves the content.

Re: Dozens of scientific journals have vanished from the internet

#124
post #98

Earlier quoted context omitted.

> If folks want to help, it would be great to have a "youtube-dl for open access papers". I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.

Unpaywall is very helpful! However, even for direct PDF links, publishing platforms will often do things like check for a session cookie; if you don't have the correct cookie you get bounced back to the landing page, where you need find and follow another link. This isn't super complicated to work around (persist a cookie jar, use a headless browser, etc), but it doesn't work out-of-the-box with our crawlers, the sam…

LOL, I waited since yesterday and ended up writing 5 min after you. :)

Re: Dozens of scientific journals have vanished from the internet

#125
post #3

In France we have HAL: https://en.wikipedia.org/wiki/Hyper_Articles_en_Ligne "Hyper Articles en Ligne, generally shortened to HAL, is an open archive where authors can deposit scholarly documents from all academic fields." I work at a university and I know library people and management check carefully that every paper we produce is deposited in HAL. Also: https://fr.wikipedia.org/wiki/Hyper_articles_en_ligne "Depuis…

In the UK we have to deposit with the institution you were working/studying at. It's a bit annoying as there's no central place to deposit.

You can use Zenodo (hosted at CERN with some EU funding). https://zenodo.org/

Re: Dozens of scientific journals have vanished from the internet

#126
post #7

This isn’t just a problem with scientific journals, but also niche research/enthusiast journals. When a magazine goes under, and the copyright holder is of a murky/unknown status, still too dangerous to digitize and make available, which is a shame.

Indeed. One part of the problem might be addressed in EU in the coming years if art. 8 of the 2019 copyright directive gets implemented well. https://www.communia-association.org/2019/12/10/implementing...

Re: Dozens of scientific journals have vanished from the internet

#127
post #23
post #4

This sounds like apologia from a big closed publisher (AAAS) explaining why open-access is supposedly bad. See, sometimes open-access journals fold, and when that happens, nobody knows what happens to the articles. But they'd like you to ignore two inconvenient facts: First, that traditional closed publishers effectively lose all articles by default by this metric! And second, that the Internet Archive, itself open-a…

I have been involved in the same projects furthering open access within Finnish universities as the corresponding author has, and I think the aim of the study is not make OA look bad, but to make it better by finding its shortcomings and then fixing them.

I think the criticism is not about the paper, only about the article describing it.

Re: Dozens of scientific journals have vanished from the internet

#128
post #31
post #4

This sounds like apologia from a big closed publisher (AAAS) explaining why open-access is supposedly bad. See, sometimes open-access journals fold, and when that happens, nobody knows what happens to the articles. But they'd like you to ignore two inconvenient facts: First, that traditional closed publishers effectively lose all articles by default by this metric! And second, that the Internet Archive, itself open-a…

Counter-point: we all know that "free" internet services are flaky. Eg many people prefer to pay for email, for reliability. This article introduces the perspective that OA journals are a bit like ad-supported email. I don't really agree that traditional publishers lose all articles by default by this metric. I think there is some value in a reliable record of the journal itself, as more than the sum of its parts. St…

Commercial publishing giants with subscription journals are ad-supported too, nowadays. You find ads in the weirdest places. https://mamot.fr/@nemobis/104755721369049700
Post reply on HN