Live data from Hacker News

ArchiveBox/ArchiveBox: open-source self-hosted web archiving

github.com

1–10 of 27 posts

Re: ArchiveBox/ArchiveBox: open-source self-hosted web archiving

#3

What is the advantage of this over something like Kiwix, or just using Playwright CLI? It seems useful but a bit unnecessary if just using to create Archive.org links.

This does _a lot more_ than just creating archive.org links. It saves the entire page contents (in multiple formats, including nicely searchable PDF and some embedded media) locally.

Re: ArchiveBox/ArchiveBox: open-source self-hosted web archiving

#4

What is the advantage of this over something like Kiwix, or just using Playwright CLI? It seems useful but a bit unnecessary if just using to create Archive.org links.

Archivebox is also an awesome tool to create copies of a website. Whether you want to demonstrate a phishing attack or do a POC to integrate your product.

Re: ArchiveBox/ArchiveBox: open-source self-hosted web archiving

#5
I really need to try this out soon. I keep bumping into it online. What do people on HN use it for?

What I do right now "to collect, save, and view sites you want to preserve offline" is by use of a Firefox plugin called WebScrapBook. Click-click-done, and I have a local searchable (!) copy of a webpage exactly as it looked in the browser. With styles and all, in one file. WebScrapBook is pretty highly configurable.

In the future I would like to have a solution that doesn't require some Firefox plugin.

Re: ArchiveBox/ArchiveBox: open-source self-hosted web archiving

#7

What is the advantage of this over something like Kiwix, or just using Playwright CLI? It seems useful but a bit unnecessary if just using to create Archive.org links.

I use it like permanent bookmarks. I can go back to it and trust it'll still be there. And this isn't just "things disappear eventually" - for a specific example, I was working on something rather last-minute recently, and wanted to refer to a vendor's whitepaper - and their entire site was "down for maintenance" all weekend. So I went onto my archive and I still had a copy from my first pass over the topic. It's a free resource, I can't blame them and I can't complain - but if I can make sure it doesn't impact me, even better.

I know other people would still have that tab open from 3 weeks ago, but I just don't work like that.

I'm not going to complain about wayback/archive.org at all, but the nature of the beast is that there's certain requests they have to obey - and with my own offline, non-exposed equivalent, I don't (well, I do, but I simply don't receive them)

Re: ArchiveBox/ArchiveBox: open-source self-hosted web archiving

#8

I really need to try this out soon. I keep bumping into it online. What do people on HN use it for? What I do right now "to collect, save, and view sites you want to preserve offline" is by use of a Firefox plugin called WebScrapBook. Click-click-done, and I have a local searchable (!) copy of a webpage exactly as it looked in the browser. With styles and all, in one file. WebScrapBook is pretty highly configurable.…

I’ve used it to save a lot of pages related to ham radio. I have several 30-40 year old radios and I’m afraid one day information about them will just drop off the internet.
Post reply on HN