Live data from Hacker News

The People Behind the Wayback Machine

motherjones.com

11–20 of 33 posts

Re: The People Behind the Wayback Machine

#11
post #3

The team at the Internet Archive, responsible for Wayback Machine, ArchiveTeam, the TV News Archive, and many other projects, are true gems. I had the chance to meet Brewster and many others whilst last in San Francisco and their passion is infectious. If anyone is interested, they have an open lunch on Fridays[1] where you get to see the church, the tech, and meet the team. Each team member and guest gives a few sen…

FYI your two footnotes are swapped :)

Re: The People Behind the Wayback Machine

#12
I love what the archive team does. I used their VM when the posterous backup effort was happening last year, and today I sent a link to my friend's now defunct posterous blog.

But I know that the owner of days posterous page had no intent on keeping the page a going concern, and was happy to see it disposed of. In light of the recent Google "right to be forgotten" ruling, will there come a day when the right to be forgotten will extend to archive.org?

Re: The People Behind the Wayback Machine

#13

I love what the archive team does. I used their VM when the posterous backup effort was happening last year, and today I sent a link to my friend's now defunct posterous blog. But I know that the owner of days posterous page had no intent on keeping the page a going concern, and was happy to see it disposed of. In light of the recent Google "right to be forgotten" ruling, will there come a day when the right to be fo…

>will there come a day when the right to be forgotten will extend to archive.org?

Sites can at any time opt out of being archived via a robots.txt exclusion (IA still keep their previous archives privately). However for public blogging sites operated by a third-party that's another matter.

Re: The People Behind the Wayback Machine

#14
post #2

All this great content and their website is designed in a way that discourages people to peruse it. They really need a re-design and "relaunch" of their brand to flaunt the great things that they're doing.

They have limited resources and already have a giant redesign in the works, according to Jason Scott.

Re: The People Behind the Wayback Machine

#16
post #2

All this great content and their website is designed in a way that discourages people to peruse it. They really need a re-design and "relaunch" of their brand to flaunt the great things that they're doing.

Unlike most of the web 2.0 world, there is substantially more value in their content than in their design wizardry. It's a team with limited resources whose can barely keep up with the information they archive and is doing an impressive work, not at all devalued by the absence of some precious yetanothercrap.js.

Re: The People Behind the Wayback Machine

#18

I love the Wayback Machine, I wish they'd archive all pages though even those that don't wish to be archived.... keeping them away from public view until copyrights expire someday.

I don't know about this, to me archiving everything seems like a gross inefficiency. Most of the internet is spam and advertising, and of the rest, less than 5% is actually useful information or knowledge.

Archiving books, scientific journals and the likes would seem much more useful, but obviously you'd run into copyright issues.

Re: The People Behind the Wayback Machine

#19
post #9

Earlier quoted context omitted.

Among many many other differences, assuming your angle isn't to denigrate the Archive team, one of the important differences is that TPB makes no garantees about content availability. Information that you can find through TPB (they host magnet links, not content. magnet links can be used to find other who host content) is only available as long as those who are interested in the content are interested in hosting it.…

That and they primarily archive public domain material and abandonware (apart from their web archiving project). They really couldn't be more different.

I think simply disregarding the web archiving is a bit of a cop out. It's interesting though that for the most part, nobody minds them redistributing loads of copyrighted material. Here's some reasons that come to mind:

They web material was distributed for free in the first place. They're redistributing ad-ware, not stuff behind a paywall. (The same can be said of some TV shows and indeed I think TV show piracy if often met with a comparatively cavalier attitude.)

It's used as a measure of last resort. If I want to read an article from Wired, I'm going to try to find it on Wired -- or more likely, I'm going to Google it and get a link to Wired, and not the archive. It's only when it's unavailable from the original publisher or when I have specific historic interest that I end up using the web archive. The result is that publishers aren't denied their ad revenues as long as they host their material. Your abandonware argument translates neatly to the web archiving efforts.

They're archiving. This gives them a touch of academia and altruism that's casts them in a totally different light.

Post reply on HN