Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

121–130 of 207 posts

Re: Thank you for helping us increase our bandwidth

#121
post #17

60 Gbit/s of continuous traffic is a lot. If I'm reading the graphs right, Wikipedia "only" has 13.4 Gbit/s of outbound traffic [1]. Of course still well below the single-digit Tbit/s traffic of large internet exchanges [2] [3], but still unexpectedly large. [1]: adding up the outbound numbers for each datacenter: 1.888 + 8.003 + 810 + 1.958 + 807: https://grafana.wikimedia.org/d/000000605/datacenter-global-... [2]:…

60 Gbps is about 1/3 of the traffic served by a single Netflix CDN node.

does archive.org use cdn or is it all served from one location?

Re: Thank you for helping us increase our bandwidth

#122
post #70
post #21

Earlier quoted context omitted.

Here's another one: "some of our old tweets were wrong so we just quietly deleted them lol": https://twitter.com/voxdotcom/status/1242537366620966912

Twitter ought to have some kind of strikethrough feature. Allow users to mark that they no longer stand behind a tweet without completely deleting it.

What would be the incentive to use that feature over just deleting the tweet?

Re: Thank you for helping us increase our bandwidth

#123
post #5

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

Please don't make the whole internet basically cloudflare. they've banned my VPN endpoint (Hetzner server) and as a result a huge chunk of websites already don't work for me despite me having done nothing wrong. I've heard reports of Tor users being restricted as well.

So many of the VPS companies have bots galore doing credit card attacks on smaller sites that it’s very easy and convenient to simply block those VPS providers.

Re: Thank you for helping us increase our bandwidth

#124
post #24

Earlier quoted context omitted.

Do you have any examples of government sites that are doing this? I have a side-hobby of setting up scrapers which pull scraped data into a git repository, precisely for this kind of thing. I'd be happy to set a few up. Some of my posts about this technique (which I call "git scraping"): https://simonwillison.net/tags/gitscraping/

Not sure if you’re asking for: (1) US coronavirus data in general, (2) examples of sources that do not log prior days data, or (3) source that “edit” prior data without noting the edits. (4) something else If (1) this page in the table under the column “sources” links to where the data came from: https://www.worldometers.info/coronavirus/country/us/

2 and 3.

Re: Thank you for helping us increase our bandwidth

#125

Earlier quoted context omitted.

At the bandwidth levels they are using they would need to use Cloudflare Enterprise, and in my experience that is way more expensive than other CDN providers. Also, since archive.org has so much content, the caching ratio is going to be very bad and kill CDN efficiency while still requiring lots of direct bandwidth. Cheap direct bandwidth in their case looks best(which is what they seem to be doing).

Cloudflare does not have bandwidth limits.

[deleted]

Re: Thank you for helping us increase our bandwidth

#126

Earlier quoted context omitted.

You assume authors are losing revenue from this effort. It is likely this revenue would never have been realized regardless of the Archive’s efforts. A piece of content copied doesn’t mean someone would’ve paid for it. As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH. https://www.entrepreneur.com/article/347840

"As an aside, many SaaS products have given away their product for free due to COVID and widespread forced WFH." I'm sure a lot of authors would have contributed their work to the effort, if they'd been asked. But they weren't asked. It's difficult to imagine how you'd similarly force SaaS companies to give away their products for free during the pandemic -- lucky for them -- but if you found a way to do it technical…

I doubt many authors would have been able to do so, given that they've signed contracts with their publishers that likely prevent them from giving their work to anyone else.

Re: Thank you for helping us increase our bandwidth

#127
What's the endgame for archive.org? Are they expecting for there to be a breakthrough in storage technology?

I read somewhere that data creation is exceeding storage solutions' pace. Is this true?

What about a mesh of some kind where every person who install an application hosts bits and pieces of random data and serves it to whoever asks for it?

Re: Thank you for helping us increase our bandwidth

#128
post #66

Earlier quoted context omitted.

ipfs: https://betanews.com/2018/08/09/decentralized-archive-org

I would love to donate bandwidth/storage. I haven't a clue how. I wish there was software I could throw in my server, set how much storage and bandwidth I can donate and run 24/7.

This should be the future of internet media as a whole. You set your limits, and billions of people do the same, suddenly no more single point of failure. Someday

Re: Thank you for helping us increase our bandwidth

#129
post #81

Archive.org works surprisingly well as a general purpose web proxy. Just prefix the URL, e.g., http://example.com , with https://web.archive.org/save/ , e.g., https://web.archive.org/save/http://example.com The aesthetic intrusiveness of the archive.org header and footer are minimal since I use a text-only browser that has no Javascript engine. Sometimes I get "This url is not available on the live web or can not be…

I use a custom browser keyword search to find existing archived pages before saving one, personally. Eg: ar for:

  https://wayback.archive.org/web/*/%S
I'd imagine it would be useful for IA to implement some message for scenarios where a page has already been saved within a certain timespan and provide both a link to the already saved version and offer to save again. As this would mitigate mass savings of an identical page that can occur when some popular link is accidentally shared with the /save/ URL instead of the static URL or when it's a popular page that people want to archive.

Archive.is displays such a message (to the effect of, 'this page was archived , if it looks outdated click save') and also redirects to the most recent copy.

Re: Thank you for helping us increase our bandwidth

#130
post #66

Earlier quoted context omitted.

ipfs: https://betanews.com/2018/08/09/decentralized-archive-org

I would love to donate bandwidth/storage. I haven't a clue how. I wish there was software I could throw in my server, set how much storage and bandwidth I can donate and run 24/7.

Every item on archive.org has a bit-torrent file, I believe. This doesn't solve the problem of working out which items are the most popular, but if you can figure out that then you would be able to manage storage and bandwidth. I suppose the real question is: how many other people use bit-torrent to download from archive.org, as opposed to direct downloads. I (sadly) suspect it's a very small fraction.

Semi-unrelated, but if you're looking for ways to help and have a spare server, Archive Team [1] is always looking for additional capacity. Although Archive Team != archive.org, they do grabs of at-risk content which (almost always) get uploaded to archive.org. [disclaimer: I help out with various Archive Team projects, the most recent of which was the backup of Yahoo Groups).

[1] https://www.archiveteam.org/

Post reply on HN