Live data from Hacker News

The DOI System

doi.org

41–50 of 67 posts

Re: The DOI System

#41
post #25
post #20

Earlier quoted context omitted.

> non-hierarchy-revealing URLs that must be resolved using a database they control What do you mean? There is a journal prefix (Nature has 10.1038) and the link usually translates fairly directly to a URL for new articles. EDIT: Just looked it up in the DOI rules. The registrant chooses the suffix. So the journals are free to maintain 1:1 mapping to URLs if they wish. I assume DOI is popular because people trust that…

> I assume DOI is popular because people trust that the central entity might do a bit better job First, I don’t know how popular it is. Second, look at what an absolutely poor job the doi.org site themselves have done keeping their showcased example links from their demonstration document from breaking over the years. See my comment elsewhere here about what I found on archive.org. What am I missing here? It looks li…

> First, I don’t know how popular it is.

Very. The doi is now the unique key to identify a scientific article.

Researchers like it because it is an easy way to get canonical metadata without having to go through Google Scholar or (dog forbid) something clunky like Science Direct. You just put the doi and bang, you’ve got the article. No issue with a typo in the volume number, or authors who can go by different names (or the same name written differently), or anything like that. It has make managing bibliography databases massively easier.

Editors like it because it’s a very efficient way of checking the information in the bibliographies of submitted manuscripts.

I have no clue about their examples, but in the real life I have never had a link rot issue in ~10 years handling thousands of references.

Re: The DOI System

#42
post #29

Earlier quoted context omitted.

That’s why the display guidelines stipulate that DOIs are always expressed as URLs.

Sure, but 10.1083 is still a shitty link compared to nature.com, independent of which url prefix is used.

Dois are not supposed to replace urls. They serve very different purposes. A doi is a unique identifier for a scientific article, that’s all it is.

Re: The DOI System

#43
post #37
post #8

Earlier quoted context omitted.

Yes, it is a URI schema. FAQ #11, at https://www.doi.org/faq.html > DOI & URI: how does the DOI system work with web URI technologies? > DOI names may be expressed as URLs (URIs) through a HTTP proxy server. In addition, DOI is a registered URI within the info-URI namespace (IETF RFC 4452, the "info" URI Scheme for Information Assets with Identifiers in Public Namespaces). See the DOI Handbook, 2 Numbering and 3 Reso…

> The original paper is no longer accessible. (An alternative is to add the equivalent of a big red stamp on it saying "RETRACTED" while leaving the content accessible.) I think you are slightly confused - the "alternative" that you describe is actually what happened in this case (and in all the cases of retracted papers that I've encountered). A Retraction Notice has been published at a separate publication [1] with…

Indeed. I misread the retraction notice quite severely.

Thank you for the correction.

As an alternative example, I searched Retraction Watch.

https://retractionwatch.com/2021/09/23/alzheimers-diagnosis-... describes the paper:

> “Intracranial pressure waveform changes in Alzheimer’s disease and mild cognitive impairment” — which now seems to have disappeared entirely from the journal’s site — appeared in Surgical Neurology International in July.

It points out the paper is still in PubMed Central, at https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8168790/ . However, its DOI, 10.25259/SNI_48_2021 , redirects to a 404 page from the journal.

Re: The DOI System

#44
post #24
post #20

Earlier quoted context omitted.

> non-hierarchy-revealing URLs that must be resolved using a database they control What do you mean? There is a journal prefix (Nature has 10.1038) and the link usually translates fairly directly to a URL for new articles. EDIT: Just looked it up in the DOI rules. The registrant chooses the suffix. So the journals are free to maintain 1:1 mapping to URLs if they wish. I assume DOI is popular because people trust that…

Compare the readability and editability of: nature.com/2021/11/17/news/ versus: 10.1038.123.366.345643 And you ask what I mean. What I mean is the first one is transparently better in several ways. If I got the syntax wrong, just focus on the first part then, nature.com versus 10.1083. One is more readable than the other.

Neither your URL nor your DOI resolves. I meant rather the correspondence of

    https://www.nature.com/articles/1771046b0

    to

    10.1038/1771046b0
or

    https://www.nature.com/articles/s41586-021-04114-w
    
    to
    
    10.1038/s41586-021-04114-w

Re: The DOI System

#45
post #20

Earlier quoted context omitted.

> non-hierarchy-revealing URLs that must be resolved using a database they control What do you mean? There is a journal prefix (Nature has 10.1038) and the link usually translates fairly directly to a URL for new articles. EDIT: Just looked it up in the DOI rules. The registrant chooses the suffix. So the journals are free to maintain 1:1 mapping to URLs if they wish. I assume DOI is popular because people trust that…

To take your point to the next step, how do two otherwise unrelated journals (of which there are tens of thousands) link between each other and solve link rot? Answer: shared open non profit metadata infrastructure. Disclaimer - I’m at Crossref AMA.

Theoretically, the journals could publish a canonical URL for each article. But apparently organizations are not good at publishing and maintaining such URLs. Especially in case of mergers and rebrandings.

Re: The DOI System

#46
DOI is one implementation of handle system defined in RFCs 3650, 3651 and 3652.

"RFC 3650: Handle System Overview" https://www.rfc-editor.org/rfc/rfc3650.txt

"RFC 3651: Handle System Namespace and Service Definition" https://www.rfc-editor.org/rfc/rfc3651.txt

"RFC 3652: Handle System Protocol (ver 2.1) Specification" https://www.rfc-editor.org/rfc/rfc3652.txt

Digital Object Architecture and the Handle System - icann https://www.icann.org/en/system/files/files/octo-002-14oct19...

>The Digital Object Architecture (DOA) is an overall architecture for managing digital objects with an associated unique persistent identifier. In this context, digital objects are defined as a sequence or set of sequences of bits. The DOA resulted from the work by Dr. Robert Kahn and Dr. Vinton Cerf at the Corporation for National Research Initiatives (CNRI) in the late 1980s.

>The DOA has three core components: the identifier/resolution system, the Digital Object Repository system, and the Digital Object Registry system containing metadata about the repository objects. The Handle System is the original name of the identifier/resolution system of the DOA. Its governance is coordinated by the DONA Foundation, a Geneva-incorporated nonprofit organization founded by CNRI.

>The Handle System and the DONA Foundation are the elements of the DOA that are closest to ICANN organization’s role helping coordinate the Internet's system of unique identifiers. This report will focus on those two elements to better understand the technology and its usage, innovation, and limitations. This report is not intended to endorse the technology nor offer any recommendations regarding its operation. It is based on the set of publicly available technical documents, an analysis of the code published on CNRI website, and a number of interviews with Dr. Robert Kahn and his team at CNRI.

https://www.handle.net/

Re: The DOI System

#47
post #33

Earlier quoted context omitted.

I dunno, looks similar to the ISBN system for books.

ISBN = International Standard Book Number DOI does resemble ISBN, but I can find book info via ISBN without relying on the publisher to maintain a database. (via Google or ISBN.nu or maybe even the Library of Congress) ISBN is indeed assigned by the publisher, who has been granted an ISBN prefix by the ISBN Fairy. That's same as DOI. But DOI feels more like Amazon's book numbering system, their ASIN. That relies on a…

It would be awesome if there actually was a reliable, API-accessible (and reasonably cheap/free for small numbers of automated lookups) database of ISBNs. I want to build my own book inventory system and be able to look up ISBNs in my own tool.

Re: The DOI System

#48
post #9

Can someone tell me where to get a definitive list of DOI prefixes? I tried to find this some time ago and came away very frustrated. The system seems in reality very opaque.

Worth reading the Wikipedia article on the Handle System, of which the DOI system is a subset: https://en.wikipedia.org/wiki/Handle_System All handles that start with 10 are DOIs. Handles that start with 20 are issued by CNRI. I don't know of a published list of prefixes, but it would be easy to investigate since entering the prefix into the appropriate resolver will bring up information on the issuing authority: - h…

> All handles that start with 10 are DOIs. Handles that start with 20 are issued by CNRI.

Interesting! I manage a system that is responsible for two Handle prefixes: 10568 and 10947. Guess we got those from CNRI around 2009 and it was different then.

Re: The DOI System

#49
post #45

Earlier quoted context omitted.

To take your point to the next step, how do two otherwise unrelated journals (of which there are tens of thousands) link between each other and solve link rot? Answer: shared open non profit metadata infrastructure. Disclaimer - I’m at Crossref AMA.

Theoretically, the journals could publish a canonical URL for each article. But apparently organizations are not good at publishing and maintaining such URLs. Especially in case of mergers and rebrandings.

That canonical URL is literally what the DOI system is. Publishers getting together to agree on a shared identifier system, regardless of each publisher's business model.

Re: The DOI System

#50
post #44
post #24

Earlier quoted context omitted.

Compare the readability and editability of: nature.com/2021/11/17/news/ versus: 10.1038.123.366.345643 And you ask what I mean. What I mean is the first one is transparently better in several ways. If I got the syntax wrong, just focus on the first part then, nature.com versus 10.1083. One is more readable than the other.

Neither your URL nor your DOI resolves. I meant rather the correspondence of https://www.nature.com/articles/1771046b0 to 10.1038/1771046b0 or https://www.nature.com/articles/s41586-021-04114-w to 10.1038/s41586-021-04114-w

Look at the metadata of those two articles.

    https://www.nature.com/articles/1771046b0
    has
    
and

    https://www.nature.com/articles/s41586-021-04114-w
    has
    
Post reply on HN