Live data from Hacker News

arXiv moving from Cornell servers to Google Cloud

info.arxiv.org

141–150 of 178 posts

Re: arXiv moving from Cornell servers to Google Cloud

#141
post #121
post #40

It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'd like to know if GCP is covering part of the bill? Or will Cornell be paying all of it? The new architecture smells of "[GCP] will pay/credit all of these new services if you agree to let one of our architects work with you". If GCP is hel…

> If GCP is helping, stay tuned for a blog post from google some time around the completion of the migration with a title like "Reaffirming our commitment to science" or something similarly self affirming. "Google pays to run an enormous intellectual resource in exchange for a self-congratulatory blogpost" seems like a perfectly acceptable outcome for society here.

It wasn't when it happened to Usenet.

Re: arXiv moving from Cornell servers to Google Cloud

#143

I would be curious if by any chance google has given a discount as long as they allow them to use latex sources of the papers to addrest their artificial intelligence models

Arxiv makes the latex source available for download since a while ago. I'm sure all of that data has long been used for training already.

Re: arXiv moving from Cornell servers to Google Cloud

#144
post #112

Odd but related question if anyone knows: Are all preprints on arXiv public? Or is there actually a private unlisted preprint queue? From behavior I've observed, I'm guessing maybe authors have the ability to hide papers and send private invites for select peer-review?

> I'm guessing maybe authors have the ability to hide papers

No they do not.

In fact, authors cannot even delete submitted papers after they have officially appeared in the (approximately) daily cycle. You can update your paper with a new version (which happens frequently) or mark it as withdrawn (which happens rarely). But in either case all the old versions remain available.

Re: arXiv moving from Cornell servers to Google Cloud

#145
post #40

It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'd like to know if GCP is covering part of the bill? Or will Cornell be paying all of it? The new architecture smells of "[GCP] will pay/credit all of these new services if you agree to let one of our architects work with you". If GCP is hel…

> If GCP is helping, stay tuned for a blog post from google some time around the completion of the migration with a title like "Reaffirming our commitment to science" or something similarly self affirming.

This is an odd criticism. If a company is footing the bill, it can’t even talk about it to gain some publicity/good will?

Re: arXiv moving from Cornell servers to Google Cloud

#146
post #124

Would love to see arXiv set up as a consortium of international academic libraries instead. Scientific publishing is where it is today because universities and scientific societies sold off or gave their journals to private enterprises. Letting Google in is a move in the wrong direction imo.

Some sort of federated preprint protocol where anyone could stand up a node and clone the existing data would be ideal. The current centralized operator then becomes "just" a curator (and competing curators are easy to set up).

Re: arXiv moving from Cornell servers to Google Cloud

#147
post #80
post #49

Earlier quoted context omitted.

That's true. I recently had to move a VM from gcp to hetzner because gcp would silently drop all packets to some countries, Iran included. And a Stack overflow question was the easiest way to learn about it, not gcp docs.

I have looked at it recently and it seems Iran is blocking GCP, not the other way around. Not sure if Google keep a doc up to date with who blocks them.

From my past experience, I can say that Google Cloud services (e.g load balancers) by default blocked traffic from ITAR sanctioned countries. Not just blocking people in those countries from becoming customers of GCP, but blocking them from accessing content hosted on GCP.

I didn't know how that situation had evolved since I last used GCP.

Re: arXiv moving from Cornell servers to Google Cloud

#148
post #96

Earlier quoted context omitted.

The people on both ends of that conversation (google vs cornell) are clever but the result will probably be enshittification.

I'm torn on flagging comments that throw out "enshittification". Do you feel that this stands in for actual thought?

[deleted]

Re: arXiv moving from Cornell servers to Google Cloud

#149
post #109

If it means I can download papers without spoofing my user-agent, then I am happy!

On the contrary, I'd expect Google to be much more proficient at blocking requests based on various factors. It wouldn't surprise me in the slightest if recaptcha made an appearance.

By all rights arxiv should be moving towards decentralization as opposed to being picked up by one of the largest centralized players.

Re: arXiv moving from Cornell servers to Google Cloud

#150
post #107
post #59

Earlier quoted context omitted.

> I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling Funny, I also work on academic sites (much smaller than arXiv) and we're looking at moving from AWS to bare metal for the same reason. The $90/TB AWS bandwidth exit tariff can be a budget killer if people write custom scripts to download all your stuff; better to slow down than 10x the monthly budget.…

The two are not comparable. The 1TB of transit at Amazon can be subdivided over many recipients, while the solid state drive is empty and only can be sent to one. That said, I agree that transit costs are too high.

So order multiple drives, transfer the data to them, and drop them in the mail to the client. That should always be the higher bandwidth option, but in a sane world it would also be less cost effective given the differences in amount of energy and sorts of infrastructure involved.

The reason to switch away from fiber should be sustained aggregate throughput, not transfer cost.

Post reply on HN