Live data from Hacker News

ArXiv declares independence from Cornell

science.org

181–190 of 300 posts

Re: ArXiv declares independence from Cornell

#181
It's not that hard to make a mirror or arXiv. Basically, anybody who can pay for hosting (which, I suppose, isn't very cheap now when the whole world uses it). It's a problem to make users switch, because academia seems to have this weird tradition of resisting all practices that, god forbid, might improve global research capabilities and move forward the scientific progress. But then, if arXiv actually becomes unusable, I suppose they won't really have much choice than to switch?

And, FWIW, I do think that arXiv truly has a vast potential to be improved. It is currently in the position to change the whole process of how the research results are shared, yet it is still, as others have said, only a PDF hosting. And since the universities couldn't break out of the whole Elsevier & co. scam despite the internet existing for the 30 years, to me, breaking free from the university affiliation sounds like a good thing.

But, of course, I am talking only about the possibilities being out there. I know nothing about the people in charge of the whole endeavor, and ultimately in depends on them only, if it sails or sinks.

Re: ArXiv declares independence from Cornell

#182

Earlier quoted context omitted.

Make it an external service then, and leave the thing that's already working great to just be. The reason authors like and use arxiv is that it gives 1) a timestamp, 2) a standardized citable ID, and 3) stable hosting of the pdf. And readers like the no-nonsense single click download of the pdf and a barebones consistent website look. All else is a side show.

You have to keep in mind that an increasing portion of their time and labor is going towards moderation and filtering due to a mass influx of nonsensical AI generated papers, non-academic numerology-tier hackery, and other useless drivel. Spinning the service off forces other the labor out onto other universities rather than leaving them to solely Cornell

Is the problem the storage cost for hosting them, the HDDs? I'm sure they can be offloaded to cold storage because most of that slop won't be opened by anyone.

Arxiv doesn't need moderation. Nobody is asking for Arxiv moderation. It needs minimal checks to remove overtly illegal content.

Re: ArXiv declares independence from Cornell

#183

Earlier quoted context omitted.

Simply anticipating basic push backs from reviewers makes sure that you do a somewhat thorough job. Not 100% thorough and the reviews are sometimes frivolous and lazy and stupid. But just knowing that what you put out there has to pass the admittedly noisily gatekept gate of peer review overall improves papers in my estimation. There is also a negative side because people try to hide limitations and honest assessment…

I really am not sure about that: https://biologue.plos.org/wp-content/uploads/sites/7/2020/05... The problem is that "optimizing for peer-review" is not the same thing as optimizing for quality. E.g., I like to add a few tongue-in-cheeks to entertain the reader. But then I have to worry endlessly about anal-retentive reviewers who refuse to see the big picture.

Currently a kind of rule of thumb is that a PhD student can graduate after approximately 3 papers published in a good peer reviewed venue.

If peer review were to go away, this whole academic system would get into a crisis. It's dysfunctional and has many problems but it's kinda load bearing for the system to chug along.

Re: ArXiv declares independence from Cornell

#184

I'm not sure why we're so focused on filtering what gets into arxiv (which is an uphill battle and DOA at this point) vs fixing the indexing , i.e. the page rank of academia. Google "sorted out" a messy web with pagerank. Academic papers link to each others. What prevents us from building a ranking from there? I'm conscious I might be over-simplifying things, but curious to see what I am missing.

Page rank was inspired by bibliometrics and evaluation of science publications. It's messed up now because of the rankings. Further fiddling with ranking will not fix the problem.

Re: ArXiv declares independence from Cornell

#185
post #113
post #110

Earlier quoted context omitted.

I responded to the idea that $300,000/year is a "mid-to-high engineering salary". CEO salaries are absurdly high everywhere.

Oh right, well it depends on CoL doesn't it? You can reframe European salaries as 'obscene' by world standards too. Both the US and Europe have totally broken and unaffordable housing markets, for example, but at least the Bay Area compensates with salary. I would say that relative to costs it's more that other salaries are obscenely low, if anything. People in Europe should be rioting, but unfortunately only the hom…

> Oh right, well it depends on CoL doesn't it?

To some extent, maybe, but often not. For example, London has similar cost of living to the Bay Area, and when I was at Meta experienced folks like Dan Abramov over in London were making about the same as fresh college hires in Menlo Park...

Re: ArXiv declares independence from Cornell

#186

Earlier quoted context omitted.

When I was involved it was an x86 machine in a rack in Rhodes Hall. I had a copy of the whole thing under my desk though in Olin Library on a Pentium 3 machine from IBM that was built like a piece of military hardware. In April the sun would shine in the windows of my office, the HVAC system was unable to cool my office, and temperatures would soar above 100F and I'd be sitting there in a tank top and drinking a lot…

Thanks for confirming. We need to stop marketing for AWS by talking about the ability to use the internet in AWS branded product terms.

The S3 API/UX/cost model is so seductively simple for static hosting though. I kind of think they deserve their ubiquity. Not on 90% of their products though.

Re: ArXiv declares independence from Cornell

#187
This is exactly what happened last time when scientific publishing got cornered. Journals run by departments and research groups were spun out or sold off to publishers and independent orgs. And they continued to slowly boil the frog over 50 years with fees and gate keeping.

Its especially problematic because while ArXiv love to claim to be working for open science, they don't default to open licensing. Much of the publications they host are not Open Access, and are only read access. So there is definitely the potential to close things off at some point in the future, when some CEO need to increase value.

Re: ArXiv declares independence from Cornell

#188

I'm not sure why we're so focused on filtering what gets into arxiv (which is an uphill battle and DOA at this point) vs fixing the indexing , i.e. the page rank of academia. Google "sorted out" a messy web with pagerank. Academic papers link to each others. What prevents us from building a ranking from there? I'm conscious I might be over-simplifying things, but curious to see what I am missing.

I am of the same opinion, and ultimately ArXiv becoming a journal that can prevent one from publishing a paper — no matter how junk it is — would pretty much kill its purpose. But I suppose that now when flooding the interned with LLM-generated garbage is almost endorsed by some satanic people, it is pretty much a security issue to have some sort of filter on uploads.

Now, honestly, I have no idea why would one spend resources on uploading terabytes of LLM garbage to arXiv, but they sure can. Even if some crazy person is publishing like 2 nonsense papers daily, it is no harm and, if anything, valid data for psychology research. But if somebody actually floods it with non-human-generated content, well, I suppose it isn't even that expensive to make ArXiv totally unusable (and perhaps even unfeasible to host). So there has to be some filtering. But only to prevent the abuse.

Otherwise, I indeed think that proper ranking, linking and user-driven moderation (again, not to prevent anybody from posting anything, but to label papers as more interesting for the specific community) is the only right way to go.

Re: ArXiv declares independence from Cornell

#189
post #113

Earlier quoted context omitted.

Oh right, well it depends on CoL doesn't it? You can reframe European salaries as 'obscene' by world standards too. Both the US and Europe have totally broken and unaffordable housing markets, for example, but at least the Bay Area compensates with salary. I would say that relative to costs it's more that other salaries are obscenely low, if anything. People in Europe should be rioting, but unfortunately only the hom…

> Oh right, well it depends on CoL doesn't it? To some extent, maybe, but often not. For example, London has similar cost of living to the Bay Area, and when I was at Meta experienced folks like Dan Abramov over in London were making about the same as fresh college hires in Menlo Park...

Yeah I was talking more about the definition of obscene. Like is it obscene to make 300k if housing is so expensive? I say no, and that London salaries are just bad. Although it would be preferable to fix the housing market.

To be fair though, Dan specifically is kind of notorious for messing up his comp negotiation. Did you not see the Twitter pile on at the time?

Re: ArXiv declares independence from Cornell

#190

Earlier quoted context omitted.

> Unfortunately, over the years, arXiv has become something like a "venue" in its own right, ... In my experience as a publishing scientist, this is partly because publishing with "reputable" journals is an increasingly onerous process, with exorbitant fees, enshittified UIs, and useless reviews. The alternative is to upload to arXiv and move on with your life.

That’s true. But that’s separate than the use in ML in Blockchain circles as a form of a marketing - using academic appearances.

Every field and every publisher has this issue though.

I've read papers in the chemical literature that were clearly thinly veiled case studies for whatever instrument or software the authors were selling. Hell, I've read papers that had interesting results, only to dig into the math and find something fundamentally wrong. The worst was an incorrect CFD equation that I traced through a telephone game of 4 papers only to find something to the effect of "We speculate adding $term may improve accuracy, but we have not extensively tested this"

Just because something passed peer review does not make it a good paper. It just means somebody* looked at it and didn't find any obvious problems.

If you are engaged in research, or in a position where you're using the scientific literature, it is vital that you read every paper with a critical lens. Contrary to popular belief, the literature isn't a stone tablet sent from God. It's messy and filled with contradictory ideas.

*Usually it's actually one of their grad students

Post reply on HN