Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

71–80 of 207 posts

Re: Thank you for helping us increase our bandwidth

#71
post #64
post #61

Earlier quoted context omitted.

I'm very worried about its backups and what happens when the next Big One hits SF. As far as I've ever been able to determine from talking to anyone at IA (e.g., Kahle, Scott), they don't really have any sort of backups that could actually be restored from in a disaster situation.

There is http://iabak.archiveteam.org , but it’s not exactly large.

The IA is about 50 PB.

IABAK stores 100 TB, or 0.2% of it.

Re: Thank you for helping us increase our bandwidth

#72
post #64

Earlier quoted context omitted.

There is http://iabak.archiveteam.org , but it’s not exactly large.

If I'm reading that correctly, it would only cost a bit over 500 bucks a month to host that whole archive on BackBlaze B2. Furthermore it would not be so hard to translate Archive.org items to IPFS objects, if there were an effort to pin a significant number of them to storage and network.

You're reading it correctly, but IABAK backs up 0.2% of the Internet Archive.

Re: Thank you for helping us increase our bandwidth

#73
post #24
post #9

The wayback machine has been essential through COVID as many government just publish the numbers and data "of the day" and the only way to compare to the day before is to look at IA.

Do you have any examples of government sites that are doing this? I have a side-hobby of setting up scrapers which pull scraped data into a git repository, precisely for this kind of thing. I'd be happy to set a few up. Some of my posts about this technique (which I call "git scraping"): https://simonwillison.net/tags/gitscraping/

Not sure if you’re asking for:

(1) US coronavirus data in general,

(2) examples of sources that do not log prior days data, or

(3) source that “edit” prior data without noting the edits.

(4) something else

If (1) this page in the table under the column “sources” links to where the data came from:

https://www.worldometers.info/coronavirus/country/us/

Re: Thank you for helping us increase our bandwidth

#74

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

At the bandwidth levels they are using they would need to use Cloudflare Enterprise, and in my experience that is way more expensive than other CDN providers. Also, since archive.org has so much content, the caching ratio is going to be very bad and kill CDN efficiency while still requiring lots of direct bandwidth. Cheap direct bandwidth in their case looks best(which is what they seem to be doing).

Cloudflare does not have bandwidth limits.

Re: Thank you for helping us increase our bandwidth

#75

Earlier quoted context omitted.

> single Netflix CDN node. Looks like they max out at 100 Gbps, or as low as 40 Gbps, depending on appliance model and link aggregation configuration. No argument either way, just thought it was cool info. https://openconnect.zendesk.com/hc/en-us/articles/3600345383...

That's out of date. Here's a presentation from last year where they talk about >190 Gbps: https://people.freebsd.org/~gallatin/talks/euro2019.pdf

[deleted]

Re: Thank you for helping us increase our bandwidth

#77
post #62

Earlier quoted context omitted.

I mean, imagine they did. So a huge group of people get to try your software for free. There's a system to try and cut them off at two weeks but it might fail. This is basically free marketing. All sorts of people who may never have tried your software try it. Nearly all of them don't buy it, but some do. It doesn't cost you any money and you're able to help people in need. There's evidence that, in general, piracy l…

> . I think that a crisis like this is exactly the time to try new and experimental ways of making things available. Yeah a crisis when the stress level of anyone is already much higher is the best time to crank the stress of authors even higher by experimenting with their livelihood without consulting them!

There is a huge difference between a book and code that I write. A book can be read once and the reader benefits. Useful software is usually useful for far more than one use. Free trials and demos are how we sell more software. A book? It's hard to have a free trial.

That said, I really have a hard time assailing libraries. Ebooks are sold to libraries at 2-5x the price I can buy the same ebook for as a consumer. Print books are sold at similar prices to consumers and libraries.

Re: Thank you for helping us increase our bandwidth

#78
post #5

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

Please don't make the whole internet basically cloudflare. they've banned my VPN endpoint (Hetzner server) and as a result a huge chunk of websites already don't work for me despite me having done nothing wrong. I've heard reports of Tor users being restricted as well.

It is usually nothing specific to your VPN setup or IP. It is mainly either for (D)DoS protection or Content Licensing related ( VPNs are easy way to bypass Geo-locked content). Many CDNs and services will allow you to access using only non commercial IP range

Re: Thank you for helping us increase our bandwidth

#79
post #43

The article doesn't mention why archive.org demand went up so much after COVID. Does anyone know?

Not just reading, general increase in internet use, for example they have a lot of old archived games, many official covid health pages only track current numbers so wayback machine is the only way to track first / second derivative data points etc.

Re: Thank you for helping us increase our bandwidth

#80
post #13

Earlier quoted context omitted.

Culture wants to be free.

Writing wants to be uncompensated?

I guess it is more information wants to be free? protecting the price of information is harder when making copies is very cheap.
Post reply on HN