Live data from Hacker News

alphaXiv: Open research discussion on top of arXiv

alphaxiv.org

171–180 of 190 posts

Re: alphaXiv: Open research discussion on top of arXiv

#171
post #102

Tenured prof here. Every paper of mine goes on Arxiv with no exceptions, published under CC BY-NC-ND licenses. Some of us are working hard to overcome the system (e.g. look at the IACR's efforts). Unfortunately, academics are still hindered by institutional inertia; in fact, many prefer the status quo, usually those who rely on prestige over actual quality to advance their careers.

> (...) in fact, many prefer the status quo, usually those who rely on prestige over actual quality to advance their careers. Your comment doesn't read like one from anyone with any relationship with academia. If you had, you'd know that the issue is not a vacuous "prestige" but funding being dependent on hard metrics such as impact factor, and in some cases with metrics being collected exclusively from a set of esta…

  > the issue is not a vacuous "prestige" but funding being dependent on hard metrics such as impact factor
These things are not in contention.

There is no singular problem to be solved, which is why it is so difficult. No smoking gun.

  > And ArXiv is not one of them.
ArXiv has a large impact on metrics and so called impact factor. But let's also not be delusional to the fact that a paper from a prestigious institution will always receive more citations (or any other metric) than an equal quality paper from a less prestigious institution. All our metrics can be hacked through publicity.

  > Reading your post
Reading yours, it sounds like you stand in the way of resolving issues in academia. Not because you don't have issues with it that you want to solve, but because you have already found the answer.

Your comment reads like someone who has a relationship with academia.

Re: alphaXiv: Open research discussion on top of arXiv

#172

Earlier quoted context omitted.

> (...) in fact, many prefer the status quo, usually those who rely on prestige over actual quality to advance their careers. Your comment doesn't read like one from anyone with any relationship with academia. If you had, you'd know that the issue is not a vacuous "prestige" but funding being dependent on hard metrics such as impact factor, and in some cases with metrics being collected exclusively from a set of esta…

> Your comment doesn't read like one from anyone with any relationship with academia. Your comment reads likewise. He didn't say he publishes them exclusively on Arxiv. It's quite common for professors to post it there as well as submit to journals. Many (most?) journals allow for it - they don't insist the ones in arxiv be taken down - as long as they're posting preprints and not the final (copyrighted) version. As…

  > Your comment reads likewise.
FWIW I thought it read like reviewer 2, which makes me actually think they have a relationship with academia.

The problem I have with their comment is that it rejects the critique by pointing to a different issue. As if there is a singular issue with academia that leads to the mess. So it comes off as a typical "reviewer 2" comment to me where the is more complaint and disagreement than critique.

FWIW, I think we in academia need to embrace the fuzziness and noise of evaluation. I think the issue is that by trying to make sense out of an extremely noisy process we placed too much value in the (still noisy) signals we can use. It is a problem to think that these are objective answers and deny the existence of Goodhart's Law (this is especially ironic in ML where any intro class discusses reward hacking). And in this, I think there's a strong coupling between cgshep's and chipdart's complaints.

As for publishing, I think we also lost sight of the main reason we publish: to communicate with our peers. Publishers played an important role since not even a century ago we could not make our works trivially available to our peers. But now the problem is information overload, not lack of information. And I think in part the review process is getting worse and worse each year as we place so little value on the act of reviewing, do not hold anyone accountable for a low quality review[0,1], do not hold ACs or Metas accountable for, and we have so many papers to review that I don't think we can expect high quality reviewing even if we actually incentivized it. I mean in ML we expect a few hundred ACs to oversee over ten thousand submissions?

My question is if we'll learn that the noisiness of the process and the randomness of rejection creates a negative feedback of papers. Where you "roll the dice" on the next conference. You resubmit without changes, as well as your new work (publish or perish baby). If we had quality reviewing at least this would push pressure for papers to get incrementally better instead of just being recycled. But recycling is a very efficient strategy right now and we've seen plenty of data to suggest it is.

[0] I understand the reasons for this. It is a tricky problem

[1] I'd actually argue we incentivize bad reviews. No one questions you when you reject a work and finding reasons to reject a work are far easier to accept one. There's always legitimate reasons to reject any work. Not to mention that the whole process is zero sum, since venue prestige is for some reason based on the percentage of papers rejected. As if there isn't even variance in year to year quality.

Re: alphaXiv: Open research discussion on top of arXiv

#173

One of the co-creators of this site. A lot of great suggestions I'm reading so far, a lot of them are currently in the works (zooming in/out, infra issues for slow loading times on some papers, google scholar claiming papers). For some more context, we are a group of 3 students with a background in AI research, and this site was initially built as an internal tool to discuss ai papers at Stanford. We've been dealing…

Moderation is typically the thing that doesn't scale. I am not sure it's a solvable problem (see reddit, stackoverflow, youtube, quora, etc. for negative examples and anti-patterns.) Often sites start out great and then degrade when they become popular.

My main recommendation was going to be organizational: to cooperate and work with arXiv itself, rather than risking a potentially adversarial or competitive relationship.

Now that I think about it however, I am convinced by a peer comment that was basically "leave arxiv the way it is and don't mess it up." So carry on then.

Re: alphaXiv: Open research discussion on top of arXiv

#174

Great idea. - The frontpage should directly show the list of papers, like with HN. You shouldn't have to click on "trending" first. (When you are logged in, you see a list of featured papers on the homepage, which isn't as engaging as the "trending" page. Again, compare HN: Same homepage whether you're logged in or not.) - Ranking shouldn't be based on comment activity, which ranks controversial papers, rather papers…

Counterpoint: please don't do any of the above and keep arxiv as it is. It is too valuable to mess it up, it is the few things on the internet that have not been ruined yet, and the "comment activity" can happen in the articles themselves at the scale of years, decades, and centuries.

> it is the few things on the internet that have not been ruined yet

Hmm, this is an interesting point.

Re: alphaXiv: Open research discussion on top of arXiv

#175

Earlier quoted context omitted.

> In general academics prefer PDF to HTML. In part, this is just because our tooling produces PDFs, so this is easiest. The tooling producing PDF by default absolutely makes the preference for PDF justifiable. However, tooling is driven by usage - if more papers come with rendered HTML (e.g. through Pandoc if necessary), and people start preferring to consume HTML, then tooling support for HTML will improve. > But al…

HTML still lacks one key feature: a way of storing the entire document as a single file that remains fully functional offline and can be reasonably expected to be widely supported for decades. Research papers are used both for communicating new results and for archiving them. The long-term stability needed for the latter has never been a strong point of web technology.

I agree that PDF is better than web technologies in terms of stability. I'm not objecting to PDFs being available (like you said, for archive purposes you want them provided by the authors), but to PDFs being the default, and oftentimes only, format available.

Re: alphaXiv: Open research discussion on top of arXiv

#176

Earlier quoted context omitted.

HTML still lacks one key feature: a way of storing the entire document as a single file that remains fully functional offline and can be reasonably expected to be widely supported for decades. Research papers are used both for communicating new results and for archiving them. The long-term stability needed for the latter has never been a strong point of web technology.

I agree that PDF is better than web technologies in terms of stability. I'm not objecting to PDFs being available (like you said, for archive purposes you want them provided by the authors), but to PDFs being the default , and oftentimes only , format available.

Note that ePub is basically just a zipped HTML file, and has become quite common for ebooks. I don’t know how that might be for archiving purposes?

I generally stick to PDF myself, but I do sometimes wish it would be more ergonomic to reflow a 2-column paper for reading on mobile on the go, for example. Also, ePub is easier to read in night mode than PDF recoloring, and seems easier to search through (try searching for a Greek letter in a PDF…).

EDIT: How is the math support in ePub though? Are people embedding KaTeX/MathJax or just relying on MathML, and how is the quality compared to TeX?

Re: alphaXiv: Open research discussion on top of arXiv

#177

Earlier quoted context omitted.

That sounds like a good thing if your goal is to advance mankind’s knowledge and not just your career. No one is forcing anyone to respond to the comments. Also, it’s not clear whether your “demands” example would even pass the moderation guidelines there.

If people can’t make a career doing science, science doesn’t get done. It’s one thing to optimize an abstract pursuit of knowledge, but you also gotta remember that you need this to be a job people are willing/able to do.

The problem with that line of thinking is that it comes from an average mindset. There’s nothing wrong with that unless you’re the one making complaints that a new system will make your work not look as good. Seems like a race to the bottom to me.

To put it another way, if it’s just a job to you, dealing with criticism is part of what you’re getting paid for anyway, same as everyone else. And if that’s still too much to ask, you can point the finger at the lack of research that isn’t profit-driven.

Re: alphaXiv: Open research discussion on top of arXiv

#178
It a shame that arXiv now it is not what it used to be, a very useful pre-print before actual publication. It looks like it is now a pseudo journal masquerading as a pre-print server since apparently arXiv has editorial and review teams that reject papers based on their 'expertises' [1].

Perhaps they think they are reputable now just because Perelman's proof papers were published there and they want to maintain their 'reputation' [2]. The irony is that Perelman would most probably not publish in arXiv if it is in their current pseudo journal status.

[1] Editorial Advisory Council:

https://info.arxiv.org/about/people/editorial_advisory_counc...

[2] Reclusive mathematician rejected honors for solving 100-year-old math problem, but he relied on Cornell's arXiv to publish:

https://news.cornell.edu/stories/2006/09/proof-100-year-old-...

Re: alphaXiv: Open research discussion on top of arXiv

#179

Earlier quoted context omitted.

I think development on the TeX-to-HTML compiler has slowed down at some point, and it's far from perfect yet. Some of the issues are probably HTML5 limitations, unlikely to be fixed any time soon (unless one wants formulas to become graphics). But there is another problem: It takes too long to load on mobile and doesn't reflow. I thought mobile was one of the reasons people wanted HTML in the first place!

In PDFs on arXiv, syntax highlighted codeblocks are graphics.

I think that's essentially only true if they are that in the original source. You can check for yourself, most papers have the TeX source available on arxiv.

Re: alphaXiv: Open research discussion on top of arXiv

#180

Great idea. - The frontpage should directly show the list of papers, like with HN. You shouldn't have to click on "trending" first. (When you are logged in, you see a list of featured papers on the homepage, which isn't as engaging as the "trending" page. Again, compare HN: Same homepage whether you're logged in or not.) - Ranking shouldn't be based on comment activity, which ranks controversial papers, rather papers…

> Use HTML rather then PDF. The PDF is the original paper, as it appears on arXiv, so using PDF is natural. In general academics prefer PDF to HTML. In part, this is just because our tooling produces PDFs, so this is easiest. But also, we tend to prefer that the formatting be semi-canonical, so that "the bottom of page 7" or "three lines after Theorem 1.2" are meaningful things to say and ask questions about. That sa…

I’m ok with the PDF but the title should be in HTML. The pdf failed to load for me due to tracker blockers (also why?!) so I was confused because there was no title but had comments
Post reply on HN