Live data from Hacker News

ArchiveBox is evolving: the future of self-hosted internet archives

docs.sweeting.me

111–120 of 166 posts

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#111

You really should add timestamping to ArchiveBox. The easiest way to do that would be via my OpenTimestamps protocol, https://opentimestamps.org It's open source and free to use, and uses Bitcoin for the actual timestamps. Users of it do not need to make Bitcoin transactions themselves as a set of community calendar servers do that for you. You also don't need a Bitcoin node to create an OTS timestamp, and you can va…

We're going to add TLSNotary support for real cryptographic signing, see my comments below :) Timestamping is also on my roadmap, definitely as a plugin (and likely paid) as it's more corporate users that really need it. We need to keep some of the really advanced attestation features paid to be able to support the rest of the business.

> We're going to add TLSNotary support for real cryptographic signing, see my comments below :)

Last I checked TLSNotary requires a trusted third party. I would strongly suggest timestamping TLSNotary evidence, to be able to prove that evidence was created prior to any of these trusted third parties being compromised.

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#112
post #87

Earlier quoted context omitted.

Let chat more. I'm almost ready to raise some seed money, hire a second staff dev or find a cofounder, and I'm looking for people that care deeply about the space. It's only been during the last few months that I decided to go all in on the project, so this is still just the first few pages of a new chapter in the project's history. (I should also mention that if you're a commercial entity relying on ArchiveBox, you…

"I too would like commit access to your promising looking project's git repo and CI/CD pipeline. Thanks, Jia Tan"

[flagged]

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#113

As someone who was archiving a doomed website earlier today using wget, I was reminded that really need to get ArchiveBox working... I used to rely on my Pinboard subscription, but apparently archive exports haven't worked for years, so those days are over.

Oh, writing my own Pinboard archive exporter is somewhere on my too-long to-do list. I should find out what would be good for importing into Archivebox. (WARC?)

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#114
post #92

So, after reading through the comments and website, I just realized I used ArchiveBox a month or two ago for a very specific purpose. You see, I inherited a boat. This boat belonged to my father. He was not materialistic but he took very good care of the things he cared about, and he cared about this boat. It's an old 18' aluminum fishing/cruising boat built in the early 1960's. It's not particularly valuable as a co…

>because 10 or 20 years ago, there were quite a few active web forums containing informational/tutorial threads from the proud owners of these old boats. ... But the forums are gone, so a large chunk of knowledge on these boats is too, probably forever. These days, that kind of info would be locked up in a closed Discord chat somewhere, so you can forget about people 20 years from now ever seeing it.

Or people today ever discovering it.

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#115
post #52

Earlier quoted context omitted.

Even when I‘m logged out I expect at least information on my geographical location to seep into the archive via URLs addressing specific CDN endpoints or similar mechanisms.

Yup, this is why the ArchiveBox browser extension sends URLs to a separate server for archiving with an isolated burner profile. I should write a full article on the security implications at some point, there aren't many good top-down explanations of why this is a hard problem.

I know it’s a lot of work but this would be great and it may give readers a deeper understanding into security in general.

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#116
I have no programming skill at all and I don’t know a ton about ArchiveBox except I set it up and ran it for myself for a while, so I’m asking as an innocent, ignorant and curious geek, but is this something that could be adapted to peer to peer distribution or some other means of making it simultaneously as private and local as you want it and as distributed and bulletproof, uptime wise, as possible?

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#117

Earlier quoted context omitted.

Let chat more. I'm almost ready to raise some seed money, hire a second staff dev or find a cofounder, and I'm looking for people that care deeply about the space. It's only been during the last few months that I decided to go all in on the project, so this is still just the first few pages of a new chapter in the project's history. (I should also mention that if you're a commercial entity relying on ArchiveBox, you…

I love this project. I "independently" "invented" it in my head the other day, and happy to see it already exists! I'd love to see blockchain proof/notary support. The ability to say "content matching this hash existed at this time. I'm exceptionally busy now but that being said, I may choose to contribute nonetheless. I'd love to connect directly, and will connect to the Zulip instance later. If we align on values,…

Can you please explain what you mean by “blockchain proof/notary support”?

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#118

Earlier quoted context omitted.

I love this project. I "independently" "invented" it in my head the other day, and happy to see it already exists! I'd love to see blockchain proof/notary support. The ability to say "content matching this hash existed at this time. I'm exceptionally busy now but that being said, I may choose to contribute nonetheless. I'd love to connect directly, and will connect to the Zulip instance later. If we align on values,…

Can you please explain what you mean by “blockchain proof/notary support”?

Motivation: Have evidence that some content existed at a particular time. For example, let's say a major website publishes an article, and later they remove it, and there is no record of it ever existing. If I host an ArchiveBox, I can look at it and see "Oh here is that article. Looks line it was published after all." However, why should you believe me I didn't just make it up?

If when I initially archived it, I computed a cryptographic hash of the content and posted that on a blockchain, then at a future date I can at least claim "As of block N, approximately corresponding to this time UTC, content that hashes to this hash exited."

If multiple unrelated parties also make the same claim, it is stronger evidence.

Is this sufficient explanation? I can expand on this more later.

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#119

Earlier quoted context omitted.

I've been using the Single File extension to save self-contained html files of pages I want to keep for posterity. I like it because any browser can open the files it creates. Is it easy to view the archive files from readeck? I haven't looked at fancier alternatives to my existing solution. https://addons.mozilla.org/en-US/firefox/addon/single-file/

Singlefile is excellent, Gildas is a great developer. ArchiveBox has had singlefile as one of its extractors built in for years :)

Thank you so much Niki :). The P2P sharing is a great idea. I really hope this feature will get things moving in the archiving field.

Re: ArchiveBox is evolving: the future of self-hosted internet archives

#120

Earlier quoted context omitted.

Can you please explain what you mean by “blockchain proof/notary support”?

Motivation: Have evidence that some content existed at a particular time. For example, let's say a major website publishes an article, and later they remove it, and there is no record of it ever existing. If I host an ArchiveBox, I can look at it and see "Oh here is that article. Looks line it was published after all." However, why should you believe me I didn't just make it up? If when I initially archived it, I com…

There's no reason to believe that the hashed and timestamped content was hosted at a particular domain, however (unless the content was signed by the author of course, then there's no Blockchain necessary). sure multiple peers could make some attestation that they saw it at that URL, but then you're back at square one of the reputation problem

Internet archive as an institution with a reputation that holds up to a judge is actually more valuable than a cryptographic proof that x bytes existed at y time

Post reply on HN