Live data from Hacker News

arXiv moving from Cornell servers to Google Cloud

info.arxiv.org

161–170 of 178 posts

Re: arXiv moving from Cornell servers to Google Cloud

#161
post #158

Earlier quoted context omitted.

So order multiple drives, transfer the data to them, and drop them in the mail to the client. That should always be the higher bandwidth option, but in a sane world it would also be less cost effective given the differences in amount of energy and sorts of infrastructure involved. The reason to switch away from fiber should be sustained aggregate throughput, not transfer cost.

The other guy was also comparing them based on transfer cost. Given that 1TB can be divided across billions of locations, shipping physical drives is not a feasible alternative to transit at Amazon in general.

I'm not trying to claim that it's generally equivalent or a viable alternative or whatever to fiber. That would be a ridiculous claim to make.

The original example cited people writing custom scripts to download all your stuff blowing your budget. A reasonable equivalent to that is shipping the interested party a storage device.

More generally, despite the two things being different their comparison can nonetheless be informative. In this case we can consider the up front cost of the supporting infrastructure in addition to the energy required to use that infrastructure in a given instance. The result appears to illustrate just how absurd the current pricing model is. Bandwidth limits notwithstanding, there is no way that the OPEX of the postal service should be lower than the OPEX of a fiber network. It just doesn't make sense.

Re: arXiv moving from Cornell servers to Google Cloud

#162
post #145
post #40

It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'd like to know if GCP is covering part of the bill? Or will Cornell be paying all of it? The new architecture smells of "[GCP] will pay/credit all of these new services if you agree to let one of our architects work with you". If GCP is hel…

> If GCP is helping, stay tuned for a blog post from google some time around the completion of the migration with a title like "Reaffirming our commitment to science" or something similarly self affirming. This is an odd criticism. If a company is footing the bill, it can’t even talk about it to gain some publicity/good will?

Footing the bill for how long?

Re: arXiv moving from Cornell servers to Google Cloud

#163
post #44

Cornell is currently in hiring freeze. These roles will not be filled. Source: I applied to a Cornell-related lab in March. A week after submitting my application the role was rescinded and my contact emailed me explaining the situation.

It looks like the links are broken to the software engineer and DevOps open roles.

Re: arXiv moving from Cornell servers to Google Cloud

#164
post #104

Earlier quoted context omitted.

Can you tell me more? I think my business needs some abhorrent sales practices. That's how it's done, right?

One example https://robindev.substack.com/p/cloudflare-took-down-our-web...

I suspect that is the result of this:

https://www.reddit.com/r/sales/comments/134u0mq/cloudflare_c...

They got rid of all of the “underperforming” sales people and hired new ones. That nightmare is the result. I suspect the higher the sales performance, the more likely they were doing things like this.

Re: arXiv moving from Cornell servers to Google Cloud

#165
post #158

Earlier quoted context omitted.

The other guy was also comparing them based on transfer cost. Given that 1TB can be divided across billions of locations, shipping physical drives is not a feasible alternative to transit at Amazon in general.

I'm not trying to claim that it's generally equivalent or a viable alternative or whatever to fiber. That would be a ridiculous claim to make. The original example cited people writing custom scripts to download all your stuff blowing your budget. A reasonable equivalent to that is shipping the interested party a storage device. More generally, despite the two things being different their comparison can nonetheless b…

That is true. I was imagining the AWS egress costs at my own work where things are going to so many places with latency requirements that the idea of sending hard drives is simply not feasible, even with infinite money and pretending the hard drives had the messages prewritten on them from the factory. Delivery would never be fast enough. Infinite money is not feasible either, but it shows just how this is not feasible in general in more than just the cost dimension.

Re: arXiv moving from Cornell servers to Google Cloud

#166
post #18

> We are already underway on the arXiv CE ("Cloud Edition") project... replace the portion of our backends still written in perl and PHP...re-architect our article processing to be fully asynchronous,.. we can deploy via Kubernetes or services like Google Cloud Run...improve our monitoring and logging facilities Why do I smell someone from G was there and sold them fancy cloud story (or they wanted VMs and reseller s…

I don't think it is that. I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling. As a primary source of information there is a lot of traffic. They do have technical issues from time to time due to this demand, and I think their stability is just due to the exceptional amount of effort they take to keep it going. They are also getting more submissions and i…

They don't need K8S, containerization yes, but not K8S.

Re: arXiv moving from Cornell servers to Google Cloud

#167
post #59

Earlier quoted context omitted.

> I work for an org with close ties to arXiv, and just like us they are getting a lot more demand due to AI crawling Funny, I also work on academic sites (much smaller than arXiv) and we're looking at moving from AWS to bare metal for the same reason. The $90/TB AWS bandwidth exit tariff can be a budget killer if people write custom scripts to download all your stuff; better to slow down than 10x the monthly budget.…

I don't understand, why don't you use cloudflare? Don't they have an unlimited egress policy with R1? Its way more predictable in my opinion that you only pay per month a fixed amount to your storage, it can also help the fact that its on the edge so users would get it way faster than lets say going to bare metal (unless you are provisioning a multi server approach and I think you might be using kubernetes there and…

Regardless, if you are delivering PDFs, you should be using a CDN.

If crawling is a problem, 1 it is pretty easy to rate limit crawlers, 2 point them at a requestor pays bucket and 3, offer a torrent with anti leech.

Re: arXiv moving from Cornell servers to Google Cloud

#168

Earlier quoted context omitted.

So it is not _ideal_ , considering that human knowledge is a common good.

Disagree. It is ideal. Iran has been behind many terrorist activities over the decades.

The point is that this is not a punishment against "them", it's a loss for everyone (human kind), who affects for the vast majority innocent people both in Iran and all over the world. I don't see a world where this is "ideal", even if you agree with the block.

We don't know where the people who will make scientific breakthrough will be. Imagine losing the cure for cancer, or a form of clean energy (or anything that could change the world for everyone) due to this.

Re: arXiv moving from Cornell servers to Google Cloud

#169
post #145
post #40

It may be that it was time for the hardware that was previously running Arxiv to be retired and this is just another Capex -> Opex decision being made by so many tech companies. I'd like to know if GCP is covering part of the bill? Or will Cornell be paying all of it? The new architecture smells of "[GCP] will pay/credit all of these new services if you agree to let one of our architects work with you". If GCP is hel…

> If GCP is helping, stay tuned for a blog post from google some time around the completion of the migration with a title like "Reaffirming our commitment to science" or something similarly self affirming. This is an odd criticism. If a company is footing the bill, it can’t even talk about it to gain some publicity/good will?

How much is the bill for running Arxiv? $1000 - $3000/month? Yeah, I don't think Google deserves any recognition for footing that bill. Likely just another self-congratulatory bullshit move on behalf of big G.

Re: arXiv moving from Cornell servers to Google Cloud

#170

Earlier quoted context omitted.

Generally speaking all companies are capable of stopping services at a whim, unless there are contractual obligations for otherwise. Singling out Google here isn't helpful unless there is a unique provision in their contract that others don't have. Also worth noting that gcp has over a decade of continuous service with no indication to think it should disappear any time soon. It's not clear why Google's consumer prod…

I wonder why they chose not to use in-house tools instead. In any case, are there any documented instances of Google Cloud discontinuing service or terminating a client's hosting for ?

Bazel RBE and IOT Core come to mind.
Post reply on HN