Live data from Hacker News

ArXiv declares independence from Cornell

science.org

271–280 of 300 posts

Re: ArXiv declares independence from Cornell

#271

The recent announcement to reject review articles and position papers already smelled like a shift towards a more "opinionated" stance, and this move smells worse. The vacuum that arXiv originally filled was one of a glorified PDF hosting service with just enough of a reputation to allow some preprints to be cited in a formally published paper, and with just enough moderation to not devolve into spam and chaos. It ha…

[deleted]

Re: ArXiv declares independence from Cornell

#272

I fear their Mozilla-ification and Wikipedia-ification. Scope creep, various outreach feel-good programs, ballooning costs, lost focus etc. And other types of enshittification. Any change to the basic premise will be a negative step. They should just be boring quiet unopininionated neutral background infrastructure.

My prediction exactly.

Maybe a bloated foundation (pursuing expensive objectives completely unrelated to ArXiv's core mission of hosting PDFs), new classes of unnecessary management staff, new and useless paid features that nobody wants, and obnoxious nag banners claiming "ArXiv is not for sale!" but demanding money anyway.

Re: ArXiv declares independence from Cornell

#273

Earlier quoted context omitted.

> They should just be quiet unopininionated neutral background infrastructure. Exactly. It should be a utility. Not quite dumb pipe, but not too far either.

We don't do 'utility' in America. Everything has S.V. brain rot - it's mixed with wall street brain rot, and now if you aren't extracting wealth out of what you have access to - you are failing.

I mean... someone needs to "unlock value" from ArXiv, right?

Re: ArXiv declares independence from Cornell

#274

I fear their Mozilla-ification and Wikipedia-ification. Scope creep, various outreach feel-good programs, ballooning costs, lost focus etc. And other types of enshittification. Any change to the basic premise will be a negative step. They should just be boring quiet unopininionated neutral background infrastructure.

> Mozilla-ification All the Mozilla executives have done for the last 15+ years is * lay off developers * spend lots of money on stupid side projects nobody asked for or wants * increase their own salaries and all that with the backdrop of falling quality, market share, and relevance. I would happily donate to Firefox, but this fucked up organization will never see a single cent from me. They will spend it on anythin…

> They will spend it on anything but Firefox, which is the only thing anybody wants them to spend it on.

;_;

Re: ArXiv declares independence from Cornell

#275

Earlier quoted context omitted.

> Mozilla-ification All the Mozilla executives have done for the last 15+ years is * lay off developers * spend lots of money on stupid side projects nobody asked for or wants * increase their own salaries and all that with the backdrop of falling quality, market share, and relevance. I would happily donate to Firefox, but this fucked up organization will never see a single cent from me. They will spend it on anythin…

And it is a risk for Arxiv too that once they start to drink the koolaid and start going to the same cocktail parties that these kinds of nonprofit board members and execs go to and will feel the need to prance around with some fancy stuff. "oh no, you see we are not a preprint server host anymore, our mission is a values driven blablabla to make a meaningful change in the blablabla, we have spent X dollars to promot…

Well, maybe they don't need to be a nonprofit. How about a public benefit corporation?

And maybe that public benefit thing, well we don't really need it do we? Now that we're deep into AI you know.

For-profit has a nice ring to it. We're delivering value to founders and shareholders, where it belongs.

Re: ArXiv declares independence from Cornell

#276

Earlier quoted context omitted.

> Ideally, that's how it is for all papers, but it isn't We require a method of filtering such that a given researcher doesn't have to personally vet in excruciating detail every paper he comes across because there simply isn't enough time in the day for that. Ideally such a system would individually for each paper provide a multi-dimensional score that was reputable. How can those be calculated in a manner such that…

Can't we do better than that? PageRank was a decent solution for websites. Can't we treat citations as a graph, calculate per-author and per-paper trustworthiness scores, update when a paper gets retracted, and mix in a dash of HN-style community upvotes/downvotes and openly-viewable commentary and Q&A by a community of experts and nonexperts alike?

You know that is what PageRank was originally for, right?

Re: ArXiv declares independence from Cornell

#277

Earlier quoted context omitted.

> Ideally, that's how it is for all papers, but it isn't We require a method of filtering such that a given researcher doesn't have to personally vet in excruciating detail every paper he comes across because there simply isn't enough time in the day for that. Ideally such a system would individually for each paper provide a multi-dimensional score that was reputable. How can those be calculated in a manner such that…

Can't we do better than that? PageRank was a decent solution for websites. Can't we treat citations as a graph, calculate per-author and per-paper trustworthiness scores, update when a paper gets retracted, and mix in a dash of HN-style community upvotes/downvotes and openly-viewable commentary and Q&A by a community of experts and nonexperts alike?

Of course we could! My tongue in cheek "exercise is left for the reader" comment was meant to convey that it's deceptively simple.

Just one example off the top of my head. How do you handle negative citations? For example a reputable author citing a known incorrect paper to refute it. You need more metadata than we currently have available.

tl;dr just draw the rest of the fucking owl.

Upvotes, downvotes, and commentary? That's extremely complicated. Long term data persistence? Moderation? Real names? Verification of lab affiliations? Who sets the rules? How do you cope with jurisdictional boundaries and related censorship requirements? The scientific literature is fundamentally an open and above all international collaboration. Any sort of closed, centralized, or proprietary implementation is likely to be a nonstarter.

Thus if your goal is a universal system then I'm fairly certain you need to solve the decentralized social networking problem as a more or less hard prerequisite to solving the decentralized scientific literature review problem. This is because you need to solve all the same problems but now with a much higher standard for data retention and replication.

Very topically I assume you'd need a federated protocol. It would need to be formally standardized. It would need a good story for data replication and archival which pretty much rules out ActivityPub and ATProto as they currently stand so you're back to the drawing board.

A nontrivial part of the above likely involves also solving the decentralized petname system problem that GNS attempts to address.

I think a fully generalized scoring or ranking system is exceedingly unlikely to be a realistic undertaking. There's no problem with isolated private venues (ie journals) we just need to rethink how they work. Services such as arxiv provide a DOI so there's nothing stopping "journals" that are actually nothing more than lightweight review platforms that don't actually host any papers themselves from being built.

Re: ArXiv declares independence from Cornell

#278

Earlier quoted context omitted.

I think there is a misunderstanding here. Does arXiv count as a publication? Yes, pretty much anything that gives you a DOI does, for example Zenodo. Does it function as a reputable anything? No. The paper you link to counts as a publication, but its reputation stands on its own, it has nothing to do with arXiv as a venue. Ideally, that's how it is for all papers, but it isn't, just by publishing in certain venues yo…

> Ideally, that's how it is for all papers, but it isn't We require a method of filtering such that a given researcher doesn't have to personally vet in excruciating detail every paper he comes across because there simply isn't enough time in the day for that. Ideally such a system would individually for each paper provide a multi-dimensional score that was reputable. How can those be calculated in a manner such that…

> We require a method of filtering such that a given researcher doesn't have to personally vet in excruciating detail every paper he comes across because there simply isn't enough time in the day for that.

We do require such a method. Isn't that what AI is for? Strictly working as a filter. You still need to personally vet in excruciating detail every paper you rely on for your work.

Re: ArXiv declares independence from Cornell

#279

Earlier quoted context omitted.

Can't we do better than that? PageRank was a decent solution for websites. Can't we treat citations as a graph, calculate per-author and per-paper trustworthiness scores, update when a paper gets retracted, and mix in a dash of HN-style community upvotes/downvotes and openly-viewable commentary and Q&A by a community of experts and nonexperts alike?

Of course we could! My tongue in cheek "exercise is left for the reader" comment was meant to convey that it's deceptively simple. Just one example off the top of my head. How do you handle negative citations? For example a reputable author citing a known incorrect paper to refute it. You need more metadata than we currently have available. tl;dr just draw the rest of the fucking owl. Upvotes, downvotes, and commenta…

> Upvotes, downvotes, and commentary? That's extremely complicated.

No, it is not. Don't throw the baby out with the bath water. Zenodo is centralized, and that is fine. A system hosted by CERN would be universal enough for most purposes.

The truth is, most papers cannot stand on their own, they need a reputable venue. While it is difficult to get into Nature, it is much more difficult to actually contribute something substantial to science. That's why we don't have a system like that.

Re: ArXiv declares independence from Cornell

#280

Earlier quoted context omitted.

Of course we could! My tongue in cheek "exercise is left for the reader" comment was meant to convey that it's deceptively simple. Just one example off the top of my head. How do you handle negative citations? For example a reputable author citing a known incorrect paper to refute it. You need more metadata than we currently have available. tl;dr just draw the rest of the fucking owl. Upvotes, downvotes, and commenta…

> Upvotes, downvotes, and commentary? That's extremely complicated. No, it is not. Don't throw the baby out with the bath water. Zenodo is centralized, and that is fine. A system hosted by CERN would be universal enough for most purposes. The truth is, most papers cannot stand on their own, they need a reputable venue. While it is difficult to get into Nature, it is much more difficult to actually contribute somethin…

I think you've misunderstood me. Did you read my final paragraph? I was agreeing with what you wrote there - that simply rethinking how centralized journals operate could accomplish the majority of the goal while sidestepping most of the complexity.

That said, I disagree that papers require a centralized venue in any fundamental sense. They currently need such a venue because we don't have a better process for vetting and filtering them at scale. The issue is that decentralizing such a process in an acceptable manner is a monstrously complicated prospect.

Post reply on HN