Live data from Hacker News

ArXiv now offers papers in HTML format

blog.arxiv.org

61–70 of 325 posts

Re: ArXiv now offers papers in HTML format

#61
post #4

Earlier quoted context omitted.

And here's the PDF of the same paper for comparison: https://arxiv.org/pdf/2312.12451.pdf

The contrast is massive. I'm much more likely to read the html version; that PDF is deeply off-putting in some hard to define way. Maybe it's the two columns, or the font, or the fact that the format doesn't adjust to fit different screen sizes.

For what it's worth, two column layouts are very common in the physical sciences, or at least in physics which I'm more familliar with. I have a feeling that the reason is at least partly to save page space when using displayed math (e.g. equations that are formatted in a break between blocks of text), which use the full text width (i.e. the width of one column) to display what may be much less than half a page wide.

Re: ArXiv now offers papers in HTML format

#63
post #23

doesn't work great with long author lists... https://browse.arxiv.org/html/2312.12907v1

The PDF is worse, so there is no simple answer to this: https://arxiv.org/pdf/2312.12907v1.pdf

At least the HTML version pairs each author with their affiliations, instead of the PDF which has all the names on page 1, and all the affiliations on page 2. That's completely unreadable.

Re: ArXiv now offers papers in HTML format

#64
post #61

Earlier quoted context omitted.

The contrast is massive. I'm much more likely to read the html version; that PDF is deeply off-putting in some hard to define way. Maybe it's the two columns, or the font, or the fact that the format doesn't adjust to fit different screen sizes.

For what it's worth, two column layouts are very common in the physical sciences, or at least in physics which I'm more familliar with. I have a feeling that the reason is at least partly to save page space when using displayed math (e.g. equations that are formatted in a break between blocks of text), which use the full text width (i.e. the width of one column) to display what may be much less than half a page wide.

It makes sense - for paper. But pixels are infinite - HTML is far better for screen display, which is how people read things nowadays.

The extra column next to the one I'm reading introduces a lot of visual noise, and the content is hard enough as it is. I'm sure physicists have all gotten used to it, but it certainly trips me up.

Re: ArXiv now offers papers in HTML format

#65

Earlier quoted context omitted.

The contrast is massive. I'm much more likely to read the html version; that PDF is deeply off-putting in some hard to define way. Maybe it's the two columns, or the font, or the fact that the format doesn't adjust to fit different screen sizes.

If you read a lot of papers in your line of work you will quickly appreciate the two columns and justification.

Admittedly, I don't read research papers. But with HTML, surely the choice between one or two columns is a checkbox away.

Re: ArXiv now offers papers in HTML format

#66
post #45

A lot of AI/ML papers these days have an accompanying interactive page like [0], will we see anything like these now directly in arXive? [0] https://voyager.minedojo.org/

I think then arXiv would have to deal with mantaining the tech stack and providing the presumably much higher server capacity to serve the more varied web pages that would result, so it seems like a tall order. arXiv already has an experimental integration with Papers with Code [0], which I guess provides similar results for the reader, though the authors have to figure out their own web hosting.

[0] https://info.arxiv.org/labs/showcase.html#arxiv-links-to-cod...

Re: ArXiv now offers papers in HTML format

#67
post #41

For anyone interested in staying informed about important new AI/ML papers on arXiv, check out https://www.emergentmind.com , a site I'm building that should help. Emergent Mind works by checking social media for arXiv paper mentions (HackerNews, Reddit, X, YouTube, and GitHub), then ranks the papers based on how much social media activity there has been and how long since the paper was published (similar to how HN a…

That looks great. No real feedback yet, but it's the kind of thing I've always been looking for as a better alternative to Twitter.

Thanks! I've got a lot more planned for it too. If anyone has any feedback that doesn't make sense to share here, or if you're a researcher who is open to some questions about how you currently follow arXiv papers, drop me a note at matt@emergentmind.com.

Re: ArXiv now offers papers in HTML format

#69

Earlier quoted context omitted.

If you read a lot of papers in your line of work you will quickly appreciate the two columns and justification.

Admittedly, I don't read research papers. But with HTML, surely the choice between one or two columns is a checkbox away.

Which checkbox?

I cannot find anything relevant in any of the 3 browsers I use (Vivialdi, Firefox, Chrome). Would really appreciate this option.

A quick search gave some apparently unmaintained browser extensions, and it's it.

Re: ArXiv now offers papers in HTML format

#70
post #6

This is a great UX addition. Why did it take them so long?

The conversion is still very error-prone. It can't convert a lot of packages, and the last paper I read, StarVector, half the HTML version is just missing. (I think it hit an error at a figure of some sort.) I reported an error, but I've been reporting errors against the ar5iv and abstracts for years now and the long tail of problems just seems like an incredible slog.

Can confirm. From an ar5iv standpoint, 2.56% articles currently fail to convert entirely, and 22.9% have known errors to the converter. That leaves 74.5% of nominally usable articles. This success rate is noticeably lower for the newest batches of arXiv submissions, as the converter hasn't caught up with the most recent package innovations.

We have a plan in place to meaningfully fall back for unknown packages, but that will take at least another year to put in place, and likely another couple of years to stabilize.

Meanwhile, there is some hope that with arXiv launching the HTML Beta we will get more contributions for package support (LaTeXML is an open source project, with public domain licensing, everybody benefits).

But again the original point is spot on. Coverage will be hit-or-miss for a while longer yet, for an arbitrary arXiv submission. The good news is that authors could work towards better support for their articles, if they wanted to.

Post reply on HN