Live data from Hacker News

ArXiv declares independence from Cornell

science.org

121–130 of 300 posts

Re: ArXiv declares independence from Cornell

#121

Earlier quoted context omitted.

Then drop it and make people upload a pdf and a zip of the latex sources. Most people I talk to hate that pipeline and spend a lot of debug hours on it when Arxiv can't compile what overleaf and your local latex install can.

Arxiv can recompile latex to support accessibility and html. Going to pdf submissions would be a major step backward.

Make it an external service then, and leave the thing that's already working great to just be.

The reason authors like and use arxiv is that it gives 1) a timestamp, 2) a standardized citable ID, and 3) stable hosting of the pdf. And readers like the no-nonsense single click download of the pdf and a barebones consistent website look.

All else is a side show.

Re: ArXiv declares independence from Cornell

#122
post #105

Earlier quoted context omitted.

Universities (outside a few) just have much weaker PR machines so you never hear what they do. Also their work is not user facing products so regular people, even tech power users won't see them.

Not sure about that. How would a university test scaling hypotheses in AI, for example? The level of funding required is just not there, as far as I know.

Universities are also not suited to test which race car is the fastest, but that does not obviate the need for academic research in mechanical engineering.

Re: ArXiv declares independence from Cornell

#123

Earlier quoted context omitted.

I came here to say something similar. As someone who works in a field that applies machine learning but is not purely focused on it, I interact with people who think that arXiv is the only relevant platform and that they don't need to submit their work to any journal, as well as people who still think that preprints don't count at all and that data isn't published until it's printed in an academic journal. It can fee…

Simply anticipating basic push backs from reviewers makes sure that you do a somewhat thorough job. Not 100% thorough and the reviews are sometimes frivolous and lazy and stupid. But just knowing that what you put out there has to pass the admittedly noisily gatekept gate of peer review overall improves papers in my estimation. There is also a negative side because people try to hide limitations and honest assessment…

I really am not sure about that: https://biologue.plos.org/wp-content/uploads/sites/7/2020/05...

The problem is that "optimizing for peer-review" is not the same thing as optimizing for quality. E.g., I like to add a few tongue-in-cheeks to entertain the reader. But then I have to worry endlessly about anal-retentive reviewers who refuse to see the big picture.

Re: ArXiv declares independence from Cornell

#124
post #104

Earlier quoted context omitted.

The reason the French can’t build these things is the same reason they shouldn’t be allowed to be in charge. It’s a preprint PDF host. Just make your own if you can run this one.

They do have their own: https://hal.science/ It is actually quite common to come across HAL in subfields of mathematics in my experience.

HAL is decidedly second-tier. Given the option, everyone would pick arXiv over HAL. Hence, HAL hosts lots of stuff that didn't (even) make it to arXiv => lots of subpar dredge.

Re: ArXiv declares independence from Cornell

#125

I fear their Mozilla-ification and Wikipedia-ification. Scope creep, various outreach feel-good programs, ballooning costs, lost focus etc. And other types of enshittification. Any change to the basic premise will be a negative step. They should just be boring quiet unopininionated neutral background infrastructure.

> Mozilla-ification All the Mozilla executives have done for the last 15+ years is * lay off developers * spend lots of money on stupid side projects nobody asked for or wants * increase their own salaries and all that with the backdrop of falling quality, market share, and relevance. I would happily donate to Firefox, but this fucked up organization will never see a single cent from me. They will spend it on anythin…

And it is a risk for Arxiv too that once they start to drink the koolaid and start going to the same cocktail parties that these kinds of nonprofit board members and execs go to and will feel the need to prance around with some fancy stuff.

"oh no, you see we are not a preprint server host anymore, our mission is a values driven blablabla to make a meaningful change in the blablabla, we have spent X dollars to promote the blablabla, take me seriously please I'm also fancy like you! "

Re: ArXiv declares independence from Cornell

#126
post #105

Earlier quoted context omitted.

Not sure about that. How would a university test scaling hypotheses in AI, for example? The level of funding required is just not there, as far as I know.

Universities are also not suited to test which race car is the fastest, but that does not obviate the need for academic research in mechanical engineering.

Perhaps but the fastest race car is not possibly marshalling in the end of human involvement in science, so you might consider these of considerably different levels of meriting the funding.

Re: ArXiv declares independence from Cornell

#127

> raised concerns about the proposed $300,000 salary for arXiv’s new CEO, saying it seemed high Is a mid-to-high engineering salary outlandish for a CEO of what is likely to be a fairly major non-profit? Even non-profits have to be somewhat competitive when it comes to salary, and the ideal candidate is likely someone who would be balancing this against a tenured position at a major university

For anybody outside the SV, and especially outside the US, this seems high, yes. arXiv does not need to and should not optimize for “shareholder value”, which is at least nominally the justification for outlandish CEO pay packages.

arXiv doesn't need much. All they do is host static pdfs uploaded by someone else with free CDN services from Fastly [0]. I'm sure they could get academics to volunteer moderation services as well.

In reality you could host the entire thing for well under $50k/year in hardware and storage if someone else is providing a free CDN. Their costs could be incredibly low.

But just like Wikipedia I see them very likely very quickly becoming a money hole that pretends to barely be kept afloat from donations. All when in reality whats actually happening is that its a ridiculous number of rent seekers managed to ride the coattails of being the defacto preprint server for AI papers to land themselves cushy Jobs at a place that spends 90+% of their money on flights and hotels and wages for their staff.

I'm already expecting their financial reports to look ridiculously headcount heavy with Personnel Expenses, Meetings and Travel blowing up. As well as the classic Wikipedia style we spend a ton of money in unclear costs [1].

Whats already sad is they stopped having a real broken down report that used to actually showed things. Like look at this beautiful screenshot of a excel sheet. Imagine if Wikipedia produced anything this clear. [2]

[0] https://blog.arxiv.org/2023/12/18/faster-arxiv-with-fastly/

[1] https://info.arxiv.org/about/reports/FY26_Budget_Public.pdf

[2] https://info.arxiv.org/about/reports/2020_arXiv_Budget.pdf

Re: ArXiv declares independence from Cornell

#128
post #105

Earlier quoted context omitted.

Universities (outside a few) just have much weaker PR machines so you never hear what they do. Also their work is not user facing products so regular people, even tech power users won't see them.

Not sure about that. How would a university test scaling hypotheses in AI, for example? The level of funding required is just not there, as far as I know.

There are a million other research things to do besides running huge pretraining runs and hyperparam grid search on giant clusters. To see what, you can start with checking out the best paper and similar awards at neurips, cvpr, iccv, iclr, icml etc.

Re: ArXiv declares independence from Cornell

#129
post #118

>Cornell, for example, had a limited capacity to pay software developers to maintain and upgrade the site, which still has a very no-frills look and feel. arXiv is doomed. It was nice while it lasted.

I am not a software engineer, although I do write programs. What is it about digital infrastructure that requires maintenance? In the natural world, there is corrosion, thermal fluctuation, radiation, seismic activity, vandalism, whathaveyou. What are the issues facing the arxiv demanding the attention of multiple people 'round the clock?

Re: ArXiv declares independence from Cornell

#130
post #82

Earlier quoted context omitted.

Salaries in the US are so bonkers. Everywhere else outside of the US, $300,000 is an outlandish high salary. To call it "mid to high" is insane.

Even in the states, it’s more a distortion caused by the big tech centres. A software engineer in Ohio doesn’t command that kind of salary, but in San Francisco or Seattle that’ll buy you a moderately-senior engineer. And while academic salaries are generally not great, tenured professors at big universities tend to make a fair bit (plus a lot more vacation time and perks than is normal in the US)

It's also caused by progressive tax rates. People take harder jobs based on net wage, not gross wage, so gross wage has to compensate.
Post reply on HN