Live data from Hacker News

Dozens of scientific journals have vanished from the internet

sciencemag.org

91–100 of 136 posts

Re: Dozens of scientific journals have vanished from the internet

#91
post #83
post #44

Earlier quoted context omitted.

Ironically, the French time format is best in my opinion, because some operating systems and services can’t handle colons in file names. ISO 8601 allows for Thhmmss.sss which can be represented in filenames, but I’d rather use something like 20h33m02.345 because it’s much easier to read at a glance than T203302.345, which looks like one decimal number.

It appeals to me as an EE (by training) too. I've occasionally slipped a '£2k3' or similar and had to explain...

[deleted]

Re: Dozens of scientific journals have vanished from the internet

#92

Earlier quoted context omitted.

What would be really useful is to know the average citation count of these journals. As someone who hates to see this stuff disappear, there is still a cynical person inside me that knows there are a glut of journals that are often used to bump publishing count for professors trying to get tenure. That cynic inside of me also realizes that a subset of the journal business is a bit of a scam anyway since frequently au…

It's worth noting that citation counts are an increasingly poor metric of paper quality (and always has been). There are multiple works to show that the rise of search engine like Google Scholar have meant that researchers are increasingly citing the same papers, because their searches are all returning the same thing. Meanwhile, there are some "sleeper" papers that are super relevant to a lot of works, offer great i…

Thank you, now I know the term to associate with one of our papers written by one of the students, a "sleeper paper". The twist is that Google ranks it consistently high in the very first page in Google search (but not Google Scholar) for well over a year now for two common keywords. Somehow it never get a good citation count mainly a few handful self-citations even it has over 99% classification in all the important metrics namely accuracy, sensitivity and specificity.

Curiously I have another co-author paper by another student that was published around the same time (same topic different application) and as of today it has over 60 citations count. It even cited by a Stanford University researcher but only with classification accuracy of below 90%!

I think this is a classic case of a sleeper paper where you have excellent results but then nobody want to cite you because it makes their paper look bad or negate their (fake) claim of novelty. FYI, most of the papers in the field use the same venerable open online database that make it easy to compare your results against others.

Re: Dozens of scientific journals have vanished from the internet

#93
post #77

Earlier quoted context omitted.

> Universities at worst don’t care. I wonder how aaronsw would feel about this statement.

Didn’t the university and publisher both request that the case be dropped? Wasn’t the DA the only one pushing for a conviction?

http://swartz-report.mit.edu/docs/report-to-the-president.pd...

"Very early in this post-arrest period, MIT decided to “remain neutral,” as between the government and Aaron Swartz, in the investigation and eventual prosecution. Initially this meant simply that MIT would not take a public position on the prosecution.Throughout the following (almost) two years, MIT’s decisions were mostly guided by this posture of neutrality."

"With regard to substance, MIT would make no statements, whether in support or in opposition, about the government’s decision to prosecute Aaron Swartz, the government’s decisions about charges in an indictment, or any possible plea bargain stances of the prosecution or the defense. [15]"

"[15]: This position of neutrality would not have necessarily extended to the sentencing phase of the prosecution, where MIT might have been prepared to advocate on behalf of Aaron Swartz had he been convicted."

Re: Dozens of scientific journals have vanished from the internet

#94

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> If folks want to help, it would be great to have a "youtube-dl for open access papers".

Zotero has an existing set of "translators".

And somewhat related to this request: I scratched out some notes last year about how to get more out of "zero-obligation communities" (like the pool of prospective contributors in open source) " rel="nofollow">https://www.colbyrussell.com/2019/02/15/what-happened-in-jan.... The long and short of it is that instead of saying something like "if folks want to help, it would be great[...]", you should provide a place for people to sign up, take them at their word that they're willing to help, and then lay out a concrete set of tasks/deliverables. People get weird about trying to avoid being seen as not gentle enough with volunteers, but the end result is a lot of unharnessed human potential. You've got a pool of mechanical turks at your disposal. Take a break from polishing the arrangement of instructions you give to the computer and focus some energy on writing the "programs" that you want to be executed by meatbags.

Re: Dozens of scientific journals have vanished from the internet

#95

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

How does one get to work at the Internet Archive?

Re: Dozens of scientific journals have vanished from the internet

#96
post #95

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

How does one get to work at the Internet Archive?

You could try https://archive.org/about/jobs.php

Re: Dozens of scientific journals have vanished from the internet

#97

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring.

Most of orgs/journals/conferences in this group just don't have the resources to be able do this, nor maintain something like this.

It was funny last year when a rep from google was doing a teleconference talk from Mountain View in Jakarta, covering things like this in front of a room full of like the top 200 journal managers in Indonesia (there are like tens of thousands of journals) and their eyes were glossing over when the rep was trying to address fixing some of the issues they encounter when crawling than hamper just indexing.

Might as well be coming from a different planet…

Re: Dozens of scientific journals have vanished from the internet

#98

At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example: https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3... and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ,…

> If folks want to help, it would be great to have a "youtube-dl for open access papers".

I think Unpaywall is already trying to do this? Or at least, for every DOI they index, they try to include a link to the direct article if known.

Re: Dozens of scientific journals have vanished from the internet

#99
post #3

In France we have HAL: https://en.wikipedia.org/wiki/Hyper_Articles_en_Ligne "Hyper Articles en Ligne, generally shortened to HAL, is an open archive where authors can deposit scholarly documents from all academic fields." I work at a university and I know library people and management check carefully that every paper we produce is deposited in HAL. Also: https://fr.wikipedia.org/wiki/Hyper_articles_en_ligne "Depuis…

It would be even better if the papers that are referenced were also archived in the same store... not for republication, which I guess could have copyright issues, but for the sake of preserving the prior art that may be necessary to understand context and the specific advances described in a paper

Re: Dozens of scientific journals have vanished from the internet

#100
post #53

I think there is a case to be made for a kind of "public utility" infrastructure for the distribution and storage of scholarly work given how 1. cheap it is, considering the size of the institutions that produce and benefit from them. 2. absurdly broken the private publishing industry has become.

I'm surprised that this doesn't exist already. There's many hundreds of open access journals already, yet they're not standardized on one internetworked system? They're all implementing the basics of a PDF repository independently, and poorly? Why?? Take arXiv and expand its scope and upload everything there. Boom, problem effectively solved using mostly existing tools.

https://en.wikipedia.org/wiki/Overlay_journal
Post reply on HN