it would be nice to have a bit of explanation on how it works and why we can be confident that we can rely upon it
OP here. Definitely, great idea :) Briefly: Sites are archived using a system written in Golang and uploaded to a Google Cloud bucket. More: The system downloads the remote HTML, parses it to extract the relevant dependencies ( , , etc) and then downloads these as well. Tesoro is even parsing CSS files to extract the url('...') file dependencies from here as well, meaning most background images and fonts should conti…
The page URI is a bit obscure though. I think a tresoro.io/example.tld/page/foobar/timestamp would look good.
What about big media content and/or small differences between them ?