Live data from Hacker News

To preserve their work journalists take archiving into their own hands

niemanlab.org

61–70 of 94 posts

Re: To preserve their work journalists take archiving into their own hands

#61
post #30

A nice social attack is to create an internet archive looking website call it archive.newtld and use it to create social proof of things you didn't actually do. "Oh yeah the Washington Post did a redesign but here are my past 10 posts which I saved in archive: link " In post truth internet, proving archives is going to be tough and unless there's some other form of verification it's going to be useless fast for "prov…

Could this be solved by digital signatures on web content? (Or, a way to store those)

I designed it back in 2017. Called it the “permissionless timestamping network”.

https://intercoin.org/technology.pdf

You’d have to store the actual content. The network would just store the hash.

It’s just a Merkle DAG. The innovation is all about forcing nodes to timestamp everything if they timestamp anything.

Blockchains are overkill for this. Blockchains are for when content that was timestamped changes. Not when it just accumulates.

Also, I wrote this back in 2017 as an aspirational roadmap: https://github.com/Qbix/architecture/wiki/Internet-2.0

Re: To preserve their work journalists take archiving into their own hands

#62
post #58

Earlier quoted context omitted.

> the half-life of many URLs (5-10 years?) makes them unreliable. "Simple" enough experiment, on this very site: Use the "past" feature on the main menu to go back a step at a time, and tally the number of broken links from the external submissions the further you go back.- The amount of dead projects, expired domains, broken links, 404s, etc. is sad.-

Someone did that and found that only 5% of submitted URLs (200k/4M) were dead: https://blog.wilsonl.in/hackerverse/ Submitted 3 months ago: https://news.ycombinator.com/item?id=40307519

Thanks for the data. That's actually not bad. Then again, 5% of a "webscale" number of sites is still a lot.-

Thanks for pointing to that study ...

Re: To preserve their work journalists take archiving into their own hands

#63

Earlier quoted context omitted.

Archive.org is such a godsend.- The entire information 'substrate' of society is ephemeral , if digital, and none (at least not enough) seem to have noticed .-

I wrote a book in 2010. It had a references section with links to about 100 websites. When I wrote the second edition only about five years later, 50% of those links no longer worked. What we're doing right now is borderline insane. We're putting all of this information on the web, but almost each individual bit of information is dependent on either a company or a human being keeping it online. It's inevitable that c…

> so almost all of the information that is online right now will just disappear in the next 80 years.

> And we essentially only have one single entity that tries to retain that information.

Will future ages find ours a dark age, a gap in their records, a void ...

... up until the point - if ever - where a sufficiently advanced solution for permanence is found and comes online?

Re: To preserve their work journalists take archiving into their own hands

#64
post #32

Earlier quoted context omitted.

Archive.org is such a godsend.- The entire information 'substrate' of society is ephemeral , if digital, and none (at least not enough) seem to have noticed .-

Two big issues with Archive.org are that 1. it's a single point of failure, they don't encourage mirror sites to emerge, and 2. they keep using the "brand" to fight unwinnable battles like hosting books they don't own online, risking the whole endeavor. I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

> Now that Google no longer shows cached results,

That was also the "end of an era" of sorts right there.-

> they don't encourage mirror sites to emerge,

Something over BitTorrent or blockchain would work well here, methinks. As a baseline substrate.-

Re: To preserve their work journalists take archiving into their own hands

#65
One of the great ironies of this situation is many of the now defunct websites had contracts and writing agreements that were absolutely egregious. Often the boilerplate would say that they owned the article (which they paid you a pittance for) until the end of all time.

Prior, in the print era, the standard agreement was they'd have the rights to your story upon publication then after a reasonable amount of time the rights would revert to the author.

Re: To preserve their work journalists take archiving into their own hands

#66

Earlier quoted context omitted.

I wrote a book in 2010. It had a references section with links to about 100 websites. When I wrote the second edition only about five years later, 50% of those links no longer worked. What we're doing right now is borderline insane. We're putting all of this information on the web, but almost each individual bit of information is dependent on either a company or a human being keeping it online. It's inevitable that c…

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

certainly not GitHub

Re: To preserve their work journalists take archiving into their own hands

#67

Earlier quoted context omitted.

I wrote a book in 2010. It had a references section with links to about 100 websites. When I wrote the second edition only about five years later, 50% of those links no longer worked. What we're doing right now is borderline insane. We're putting all of this information on the web, but almost each individual bit of information is dependent on either a company or a human being keeping it online. It's inevitable that c…

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

But wouldn't that be a big waste if everyone who references the same thing is then keeping a copy of it.

Re: To preserve their work journalists take archiving into their own hands

#68

Earlier quoted context omitted.

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

But wouldn't that be a big waste if everyone who references the same thing is then keeping a copy of it.

Better many copies than none. References usually mean written text and maybe some figures, cost of storage is going down, we can afford the duplication.

Re: To preserve their work journalists take archiving into their own hands

#69
post #32

Earlier quoted context omitted.

Archive.org is such a godsend.- The entire information 'substrate' of society is ephemeral , if digital, and none (at least not enough) seem to have noticed .-

Two big issues with Archive.org are that 1. it's a single point of failure, they don't encourage mirror sites to emerge, and 2. they keep using the "brand" to fight unwinnable battles like hosting books they don't own online, risking the whole endeavor. I still appreciate it, but just imagine if it goes down due to a lawsuit. Now that Google no longer shows cached results, an entire historical record would be gone.

Its surprising that archive.org is the only such outfit I have encountered. Just like we have had libraries since ancient times, why are there so few digital libraries? There must be others, but nowhere near the number (or awareness) that we should have.

Heck, existing paper-based libraries should probably each include a digital archiving department.

Maybe this is already happening or already exists, and is trivial to those studying library science or something. I can hope, anyway.

Re: To preserve their work journalists take archiving into their own hands

#70

Earlier quoted context omitted.

> references section with links to about 100 websites. Books deserve a github repo with PDF web archives of referenced links, the same way that Wikipedia mirrors the content of cited links.

But wouldn't that be a big waste if everyone who references the same thing is then keeping a copy of it.

One person’s waste is another person’s resilience.
Post reply on HN