Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

21–30 of 207 posts

Re: Thank you for helping us increase our bandwidth

#21
post #16
post #9

The wayback machine has been essential through COVID as many government just publish the numbers and data "of the day" and the only way to compare to the day before is to look at IA.

Governments aren't the only ones overwriting old information with new either. The BBC has developed an annoying habit of completely overwriting old articles about Covid-19 with new, semi-related ones, meaning that the only way to see what they were saying earlier in the pandemic is through sites like the Internet Archive. There's also at least one outright correction that only seems to exist on the Internet Archive n…

Here's another one: "some of our old tweets were wrong so we just quietly deleted them lol":

https://twitter.com/voxdotcom/status/1242537366620966912

Re: Thank you for helping us increase our bandwidth

#22
post #5

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

Please don't make the whole internet basically cloudflare. they've banned my VPN endpoint (Hetzner server) and as a result a huge chunk of websites already don't work for me despite me having done nothing wrong. I've heard reports of Tor users being restricted as well.

It's more likely that website owners are blocking the ASNs of hosting providers since those are often used for content scraping and exploit type attacks [since they're cheap]. There are public and private lists of hosting provider ASNs you can use to block them all if you generally only want visitors with a business/residential IP.

Re: Thank you for helping us increase our bandwidth

#23

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

The geographic distribution benefits for archive.org seem minimal to me - for the actual content high latency isn't a big deal.

Re: Thank you for helping us increase our bandwidth

#24
post #9

The wayback machine has been essential through COVID as many government just publish the numbers and data "of the day" and the only way to compare to the day before is to look at IA.

Do you have any examples of government sites that are doing this?

I have a side-hobby of setting up scrapers which pull scraped data into a git repository, precisely for this kind of thing. I'd be happy to set a few up.

Some of my posts about this technique (which I call "git scraping"): https://simonwillison.net/tags/gitscraping/

Re: Thank you for helping us increase our bandwidth

#25

For the video and audio content. Could they take advantage of something like Webtorrent.

Every item has a torrent file, and can be retrieved with a BitTorrent client.

https://help.archive.org/hc/en-us/articles/360004715251-Arch...

Re: Thank you for helping us increase our bandwidth

#26
post #12

Earlier quoted context omitted.

I’m sorry to hear this. I will make an additional donation on your behalf. I can appreciate how authors feel, but they must also realize copyright laws are asymmetrical (tilted heavily towards copyright owners) and in these times, many have no access to their local public library. They are essentially robbed of access to the book materials their tax dollars (which pays an author) have paid for during this time. It’s…

I also find it amusing that sofware developers do not recognize any more how we are on the same side in this fight. Imagine if we were still selling boxed software and they decided that in light of COVID-19 they just hand it out with a two week limit. It's an imperfect comparison but the current situation is not a reason to just discard the law. It might not be the best law, the place to fight that is in Congress and…

> Imagine if we were still selling boxed software and they decided that in light of COVID-19 they just hand it out with a two week limit.

I think that is a fair comparison. As a software developer, I would JUMP at the opportunity to do that. Partly because it would be an opportunity to help sustain the world through this time of crisis at no real cost to myself (perhaps some opportunity cost). And partly because it would serve as free advertising.

Re: Thank you for helping us increase our bandwidth

#27

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

I don't foresee use of a CDN saving them much or any money. Cache hit ratio would probably be relatively low and the amount of storage required at the edge to cover an appreciable portion of traffic would be very expensive. E.g. their Akamai bill would be in the five figures per month if not six figures, depending on what kind of volume discount they could arrange, just based on outbound bandwidth. They're not the cheapest but also not the most expensive.

Serving up that large of a media library at mid scale just isn't really a great use case for a CDN, those that have to due to transit costs becoming truly enormous (e.g. Netflix) make an enormous investment in hardware at the edge that probably isn't affordable to IA (or necessary at this point).

Re: Thank you for helping us increase our bandwidth

#28
post #12

Earlier quoted context omitted.

I also find it amusing that sofware developers do not recognize any more how we are on the same side in this fight. Imagine if we were still selling boxed software and they decided that in light of COVID-19 they just hand it out with a two week limit. It's an imperfect comparison but the current situation is not a reason to just discard the law. It might not be the best law, the place to fight that is in Congress and…

You assume authors are losing revenue from this effort. It is likely this revenue would never have been realized regardless of the Archive’s efforts. A piece of content copied doesn’t mean someone would’ve paid for it. As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH. https://www.entrepreneur.com/article/347840

"As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH."

I'm sure a lot of authors would have contributed their work to the effort, if they'd been asked. But they weren't asked. It's difficult to imagine how you'd similarly force SaaS companies to give away their products for free during the pandemic -- lucky for them -- but if you found a way to do it technically, how do you think they'd react?

Re: Thank you for helping us increase our bandwidth

#29

Earlier quoted context omitted.

You assume authors are losing revenue from this effort. It is likely this revenue would never have been realized regardless of the Archive’s efforts. A piece of content copied doesn’t mean someone would’ve paid for it. As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH. https://www.entrepreneur.com/article/347840

"As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH." I'm sure a lot of authors would have contributed their work to the effort, if they'd been asked. But they weren't asked. It's difficult to imagine how you'd similarly force SaaS companies to give away their products for free during the pandemic -- lucky for them -- but if you found a way to do it technical…

Honest question: how do you propose a non profit online library ask every author for written permission, while also considering the legal interests of their publishers?

Re: Thank you for helping us increase our bandwidth

#30
post #17

60 Gbit/s of continuous traffic is a lot. If I'm reading the graphs right, Wikipedia "only" has 13.4 Gbit/s of outbound traffic [1]. Of course still well below the single-digit Tbit/s traffic of large internet exchanges [2] [3], but still unexpectedly large. [1]: adding up the outbound numbers for each datacenter: 1.888 + 8.003 + 810 + 1.958 + 807: https://grafana.wikimedia.org/d/000000605/datacenter-global-... [2]:…

I wonder what the comparison between served data is for IA vs Wikipedia - anecdotally, I feel like most of the data for Wikipedia is text based (plus some number of images), whereas IA includes (as mentioned in the article) audio, video, etc.
Post reply on HN