Live data from Hacker News

The arXiv of the future will not look like the arXiv

ar5iv.labs.arxiv.org

41–50 of 52 posts

Re: The arXiv of the future will not look like the arXiv

#41

> the PDF is not a format fit for sharing, discussing, and reading on the web. PDFs are (mostly) static, 2-dimensional and non-actionable objects. It is not a stretch to say that a PDF is merely a digital photograph of a piece of paper. It is too far a stretch, murdering the poor subject: PDFs are the best format available for long-term information, such as research papers. They have the advantages of digital data: S…

Of all the properties you mention, that PDFs "preserve presentation across platforms" is the only one that isn't shared with responsibly wielded HTML (e.g. the sort of thing that Zotero produces when stashing a local copy—which uses SingleFile under the hood). It's also the one property that is net undesirable—being more liability than benefit.

Being sent a PDF of an academic paper to read (or do anything with other than send it to a printer) is about ten times lower on the user preference scale than having someone send a link to a blog post on the same subject. (The other reason for that being that when people are in the mode that involves writing an academic paper, they forget how to write anything that anyone would actually want to read. Most academic writing sucks.)

Of the properties you listed that PDF does share with self-contained HTML, on the other hand, there isn't one that PDF isn't worse at—not even "transmittable". (Initially I would have put them on the same level there, but of course that's wrong. When you're in an environment where for whatever reason a file copy is not an option, PDF's binary format makes it harder to transmit the bytestream than HTML.)

Who cares if a PDF looks the same everywhere if that means everyone who encounters it bounces away rather than having to slog through any attempt to actually read it?

Re: The arXiv of the future will not look like the arXiv

#42
post #9

Readers may find the Octopus project interesting: > Designed to replace journals and papers as the place to establish priority and record your work in full detail, Octopus is free to use and publishes all kinds of scientific work, whether it is a hypothesis, a method, data, an analysis or a peer review. > Publication is instant. Peer review happens openly. All work can be reviewed and rated. > Your personal page reco…

Similarly, the MIT-affiliated PubPub https://www.pubpub.org/>.

Re: The arXiv of the future will not look like the arXiv

#43
post #41

> the PDF is not a format fit for sharing, discussing, and reading on the web. PDFs are (mostly) static, 2-dimensional and non-actionable objects. It is not a stretch to say that a PDF is merely a digital photograph of a piece of paper. It is too far a stretch, murdering the poor subject: PDFs are the best format available for long-term information, such as research papers. They have the advantages of digital data: S…

Of all the properties you mention, that PDFs "preserve presentation across platforms" is the only one that isn't shared with responsibly wielded HTML (e.g. the sort of thing that Zotero produces when stashing a local copy—which uses SingleFile under the hood). It's also the one property that is net undesirable—being more liability than benefit. Being sent a PDF of an academic paper to read (or do anything with other…

> responsibly wielded HTML

That would be fantastic, but there are no available solutions that meet the specs I listed, including long-term preservation and annotation (what annotation subsystems are there for HTML?). ePub is 'responsibly wielded HTML', but it lacks annotation and long-term preservation is iffy.

I much prefer PDFs to blog posts, personally - they are mine, I can annotate them, etc. Also, I find much more thought is put into a PDF than a blog post (which both beat Twitter!).

Re: The arXiv of the future will not look like the arXiv

#44
post #39

Earlier quoted context omitted.

If I understand correctly, your comment addresses PDFs created from scanning paper. PDFs at arXiv are converted from LaTeX inputs, per the OP, and not via scanning and OCR; therefore they contain perfect renditions of the text. >> Searchable, copy-able, transmittable, and data is extractable. They are also an open format, don't rely on a central service to be available, and they preserve presentation across platforms…

I'm not talking about scanned data. I'm talking about digitally born PDF, which have to be OCRed nevertheless because their text layer is unusable. LaTex is one of the worst offenders when it comes to producing mangled text layers. Multi column text is often stored with both columns interleaving, or not at all. Verbatim is mangled, formulas are a hot mess. The text order between paragraphs is not preserved. LaTex is…

> LaTex is one of the worst offenders when it comes to producing mangled text layers. Multi column text is often stored with both columns interleaving, or not at all. Verbatim is mangled, formulas are a hot mess. The text order between paragraphs is not preserved.

That's interesting. I must not deal with many LaTeX-based PDFs. The text in electronically-born PDFs I use is usually nearly flawless, with the exceptions of the bizarre extra space inserted between some words, and the challenge of hyphenated words on lines that no longer wrap in that spot.

> All those extra features you mention make your "still readable in 50 years" requirement go out the window pretty quickly.

I don't have your expertise, but I've heard a different story from librarians regarding PDF and particularly PDF/A.

Re: The arXiv of the future will not look like the arXiv

#45
post #36

> the PDF is not a format fit for sharing, discussing, and reading on the web. PDFs are (mostly) static, 2-dimensional and non-actionable objects. It is not a stretch to say that a PDF is merely a digital photograph of a piece of paper. It is too far a stretch, murdering the poor subject: PDFs are the best format available for long-term information, such as research papers. They have the advantages of digital data: S…

> data is extractable That's true for plain text (in the best case), but try extracting an equation, table or a diagram. Stepping away from best case, PDFs in theory look the same everywhere, but turn into a mess on buggy implementations or differing rendering engines – due to the insistence on having a stable presentation, they assume positining and sizing always works, so when that fails, it fails worse than a bugg…

> That's true for plain text (in the best case), but try extracting an equation, table or a diagram.

Good point. From what format are tables, diagrams, and formulas extractable (while retaining format)? I've had good luck moving tables between my web browser and email applications, though it always surprises me that the html is implemented similarly enough.

> PDFs in theory look the same everywhere, but turn into a mess on buggy implementations or differing rendering engines

I don't deal with PDFs programatically, and it sounds like you might, but from the user end, and from running networks of thousands of users, I've hardly ever seen problems in practice except for the browsers' JavaScript renderers.

Re: The arXiv of the future will not look like the arXiv

#46
post #41

Earlier quoted context omitted.

Of all the properties you mention, that PDFs "preserve presentation across platforms" is the only one that isn't shared with responsibly wielded HTML (e.g. the sort of thing that Zotero produces when stashing a local copy—which uses SingleFile under the hood). It's also the one property that is net undesirable—being more liability than benefit. Being sent a PDF of an academic paper to read (or do anything with other…

> responsibly wielded HTML That would be fantastic, but there are no available solutions that meet the specs I listed, including long-term preservation and annotation (what annotation subsystems are there for HTML?). ePub is 'responsibly wielded HTML', but it lacks annotation and long-term preservation is iffy. I much prefer PDFs to blog posts, personally - they are mine, I can annotate them, etc. Also, I find much m…

> there are no available solutions that meet the specs I listed

We're going in circles. "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for.

If you're preparing material for print and it's important to be able to represent the exact printed layout (e.g. to print again), then PDF makes sense. If printing doesn't appear in the pipeline twice or even once, then PDF is very, very bad.

> I much prefer PDFs to blog posts, personally - they are mine, I can annotate them, etc.

You can do that with blog posts.

> I find much more thought is put into a PDF than a blog post

I don't. I find, as I alluded to before, that there's much less thought put into trying to express things clearly and economically. Instead, that concern is replaced with a concern for writing in a way that sounds "academic" but is painful to read.

PS:

> ePub is 'responsibly wielded HTML'

Not at all. EPUB is very irresponsibly designed. "The format works in my browser today" should have been the #1 sanity check on that workgroup's output. They failed.

Re: The arXiv of the future will not look like the arXiv

#47
post #46

Earlier quoted context omitted.

> responsibly wielded HTML That would be fantastic, but there are no available solutions that meet the specs I listed, including long-term preservation and annotation (what annotation subsystems are there for HTML?). ePub is 'responsibly wielded HTML', but it lacks annotation and long-term preservation is iffy. I much prefer PDFs to blog posts, personally - they are mine, I can annotate them, etc. Also, I find much m…

> there are no available solutions that meet the specs I listed We're going in circles. "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for. If you're preparing material for print and it's important to be able to represent the e…

> "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for.

I think you are missing the experiences of a large part of the user population. They put together their report, or book, or brochure or datasheet or whatever, and they want it to look a certain way, regardless of the platform, and they almost all use PDF. People care very much about how their work product looks. PDF is a solution that satisfies that constraint - I have seen it do that very consistently for a long time.

How do you annotate blog posts, and in way that is preserved for decades.

I almost suspect we are somehow talking different things, because even the most non-technical users know that about PDFs. But also it reads like you are finding a way to disagree with everything.

Re: The arXiv of the future will not look like the arXiv

#48
post #46

Earlier quoted context omitted.

> there are no available solutions that meet the specs I listed We're going in circles. "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for. If you're preparing material for print and it's important to be able to represent the e…

> "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for. I think you are missing the experiences of a large part of the user population. They put together their report, or book, or brochure or datasheet or whatever, and they want…

> I think you are missing the experiences of a large part of the user population.[...] People care very much about how their work product looks.

I'm aware such people exist, I just don't confuse that fact with a belief that a majority of readers don't have problems with the experience of e.g. trying to read a PDF on a phone or anything else that isn't at least A4-/letter-sized (heck, PDFs have fewer people read them than would otherwise even when the person doing the reading is using a desktop or laptop)—precisely because PDFs preserve the presentation for print.

> How do you annotate blog posts, and in way that is preserved for decades.

Is that a statement or a question? In any case, it's hard to begin to conceptualize what sort of misunderstandings about the relevant media could lead to either. Blog posts (published originally as HTML, that is) are not inherently less susceptible to being saved than an academic article published as PDF. I did, however, already mention Zotero (and SingleFile).

> I almost suspect we are somehow talking different things, because even the most non-technical users know that about PDFs. But also it reads like you are finding a way to disagree with everything.

Am I? I'm pretty sure that I understand what you're saying, at least, and that we're talking about the same things. I don't know what you expect, though, when the fundamental premises are in dispute. There's no way to just "yes-and" through disagreements like that.

Re: The arXiv of the future will not look like the arXiv

#49
post #48

Earlier quoted context omitted.

> "Preserve presentation across platforms" is an anti-feature. No one has created a solution that satisfies that constraint because it's (a) a lot of work for (b) something that is the opposite of what the people involved are actually aiming for. I think you are missing the experiences of a large part of the user population. They put together their report, or book, or brochure or datasheet or whatever, and they want…

> I think you are missing the experiences of a large part of the user population.[...] People care very much about how their work product looks. I'm aware such people exist, I just don't confuse that fact with a belief that a majority of readers don't have problems with the experience of e.g. trying to read a PDF on a phone or anything else that isn't at least A4-/letter-sized (heck, PDFs have fewer people read them…

> I'm pretty sure that I understand what you're saying, at least, and that we're talking about the same things. I don't know what you expect, though, when the fundamental premises are in dispute. There's no way to just "yes-and" through disagreements like that.

If you start from the premise that you know everything, there's not much to talk about. The basis of such interactions is to try to learn from the other person, be intellectually curious.

> Is that a statement or a question? In any case, it's hard to begin to conceptualize what sort of misunderstandings about the relevant media could lead to either.

That the is the language of someone trying to fight - about document formats!

Re: The arXiv of the future will not look like the arXiv

#50
post #48

Earlier quoted context omitted.

> I think you are missing the experiences of a large part of the user population.[...] People care very much about how their work product looks. I'm aware such people exist, I just don't confuse that fact with a belief that a majority of readers don't have problems with the experience of e.g. trying to read a PDF on a phone or anything else that isn't at least A4-/letter-sized (heck, PDFs have fewer people read them…

> I'm pretty sure that I understand what you're saying, at least, and that we're talking about the same things. I don't know what you expect, though, when the fundamental premises are in dispute. There's no way to just "yes-and" through disagreements like that. If you start from the premise that you know everything, there's not much to talk about. The basis of such interactions is to try to learn from the other perso…

You've resorted to personal attacks—and purely personal attacks—after imagining attacks from the other end. Contribute to the topic or don't, but don't change the subject as a substitute for engaging with it.
Post reply on HN