Live data from Hacker News

ArXiv declares independence from Cornell

science.org

291–300 of 300 posts

Re: ArXiv declares independence from Cornell

#291
I go back to xxx.lanl.gov days - that is, the beginning. Back then it was all physics, some math and a little quantitative finance (not bitcoin). And the quality was pretty good because it was a preprint archive. In fact, a headline from 2000:

APS and BNL Host XXX e-Print Archive Mirror Feb. 1, 2000

The APS is establishing, in cooperation with Brookhaven National Laboratory, the first electronic mirror in the United States for the Los Alamos e-Print Archive.

Today, from the landing page, it describes itself as "arXiv is a free distribution service and an open-access archive for nearly 2.4 million scholarly articles in the fields of [long list]. Materials on this site are not peer-reviewed by arXiv.

Well, that's a large part of the problem. A lot of the stuff there now will never see a journal (even of dubious quality) and there is limited filtering of what new submissions will be stored. GIGO.

Best thing ArXiv could do is go back to their roots - limit the fields and return to preprint only. Spin off the comp sci stuff for sure to someone else along with all its headaches.

fixed: url

Re: ArXiv declares independence from Cornell

#292

Earlier quoted context omitted.

You have to keep in mind that an increasing portion of their time and labor is going towards moderation and filtering due to a mass influx of nonsensical AI generated papers, non-academic numerology-tier hackery, and other useless drivel. Spinning the service off forces other the labor out onto other universities rather than leaving them to solely Cornell

Is the problem the storage cost for hosting them, the HDDs? I'm sure they can be offloaded to cold storage because most of that slop won't be opened by anyone. Arxiv doesn't need moderation. Nobody is asking for Arxiv moderation. It needs minimal checks to remove overtly illegal content.

When you stop moderating input, that's when someone builds a fuse filesystem on top of it. We had those for discord (dsfs), twitterfs, redditfs, yt-media-storage, etc. It's also when someone starts using it to distribute malware, like websites built on a combination of GitHub and a cdn.

Re: ArXiv declares independence from Cornell

#293

Earlier quoted context omitted.

I came here to say something similar. As someone who works in a field that applies machine learning but is not purely focused on it, I interact with people who think that arXiv is the only relevant platform and that they don't need to submit their work to any journal, as well as people who still think that preprints don't count at all and that data isn't published until it's printed in an academic journal. It can fee…

Simply anticipating basic push backs from reviewers makes sure that you do a somewhat thorough job. Not 100% thorough and the reviews are sometimes frivolous and lazy and stupid. But just knowing that what you put out there has to pass the admittedly noisily gatekept gate of peer review overall improves papers in my estimation. There is also a negative side because people try to hide limitations and honest assessment…

[dead]

Re: ArXiv declares independence from Cornell

#295
post #292

Earlier quoted context omitted.

Is the problem the storage cost for hosting them, the HDDs? I'm sure they can be offloaded to cold storage because most of that slop won't be opened by anyone. Arxiv doesn't need moderation. Nobody is asking for Arxiv moderation. It needs minimal checks to remove overtly illegal content.

When you stop moderating input, that's when someone builds a fuse filesystem on top of it. We had those for discord (dsfs), twitterfs, redditfs, yt-media-storage, etc. It's also when someone starts using it to distribute malware, like websites built on a combination of GitHub and a cdn.

We are talking about a different kind of moderation. People want to filter out incorrect information that in their opinion damages the reputation of Arxiv, eg covid stuff. It's not about dumping binary data.

This is a motte and bailey fallacy. The real question is about moderation with the goal of checking truth and the scientific content. Obviously illegal content and ddos type overloading attacks need to be blocked.

Very different philosophies are clashing here. Arxiv came about in an age of different zeitgeist. We may never get back to that moment.

Re: ArXiv declares independence from Cornell

#296
post #60

Earlier quoted context omitted.

For everyone else, The reason is because arxiv is growing significantly leading to 297,000 deficit in operating costs for 2025 alone. Corenell has helped with donation a long with other organizations that pay membership fees. As a result, donors + leaders of arxiv think it's best to spin off to increase funding.

What is unclear why they need stuff of 27 and 6.7 million to operate essentially static hosting website in 2026.

https://info.arxiv.org/about/reports/2024_arXiv_annual_repor...

A critical component of the arXiv-CE project is moving our services entirely off of Cornell University’s infrastructure — this goal is also known as Milestone 1. Milestone 1 completion is projected for the end of fiscal year 2026.

Assume if you are a library, and every day, half baked so-called books brought to the librarians where they have to make sure it is meaningful, readable and printable, 3000 of them, they accept and put them in the right bookshelf, and entire internet reads every one of them on the shelf multiple times by the AI bots, search engines and researchers.

They are not only making a new library, they are also maintaining both and syncing two libraries because Cornell cannot handle the volume of access by bots.

It is not static. It is essentially running two ships side-by-side, and two ships need to appear as one from the outside. And, the new ship is still only half built. The new ship is being designed, and being built. 27 seems small to me.

Re: ArXiv declares independence from Cornell

#297

Earlier quoted context omitted.

Silicon Valley is the only place in the United States where $300K is even close to the "middle" of anything. I just moved to SV a few months ago from the Midwest (and not a particularly cheap part of it). Telling my coworkers who aren't from the US what a house costs in Wisconsin, you'd have thought I was the one who moved from a foreign country.

As a datapoint, I get paid just under 250k/yr and I'm an above average developer in his very late career, at a midwest company. 300k avg for SV is about right. The local college and medical administrators are the ones that own the mansions in my city. I have a family, house and mortgage plus my large medical expenses (cardiac) I can handle...until I cant.

Holy moly, $250 in the midwest? Where do I get your job?

For reference, I just left a position in the Midwest for a job in SV that pays a little more than you're getting paid. $250 but with Midwestern rent would be life-changing. Sounds like we're in very different stages of our careers, though.

Re: ArXiv declares independence from Cornell

#298

Earlier quoted context omitted.

Silicon Valley is the only place in the United States where $300K is even close to the "middle" of anything. I just moved to SV a few months ago from the Midwest (and not a particularly cheap part of it). Telling my coworkers who aren't from the US what a house costs in Wisconsin, you'd have thought I was the one who moved from a foreign country.

> Silicon Valley is the only place in the United States where $300K is even close to the "middle" of anything. It does heavily cluster around SV, for sure, but Seattle/NewYork/Boston/Arlington will all get you there, and Chicago/Austin/etc aren't all that far behind at this point

I just left a position in Chicago because SV pays me about double.

Re: ArXiv declares independence from Cornell

#299

Earlier quoted context omitted.

Volunteer moderators are a valid option. And I think may work out better than paid employees.

volunteer moderators are a valid option however this is also the way peer review works and the system is unfortunately very problematic and exploitative. First pass sanity checks are also a lot less fun than proper peer review so paying moderators to do it is probably safer in the long run or else you end up with cliques of moderators who only keep moderating out of spite/personal vendettas against certain groups or…

How is wikipedia successful?

Re: ArXiv declares independence from Cornell

#300
post #196
post #184

Earlier quoted context omitted.

Page rank was inspired by bibliometrics and evaluation of science publications. It's messed up now because of the rankings. Further fiddling with ranking will not fix the problem.

+1, PageRank was taken from academia. They even cited it in their original work. Funny how the origins of these things get forgotten.

Inspired by, not taken. It's a clever solution to a hard problem!

Do you remember Ask Jeeves? Dogpile? Google was an incredible improvement!

Post reply on HN