Live data from Hacker News

The People Behind the Wayback Machine

motherjones.com

21–30 of 33 posts

Re: The People Behind the Wayback Machine

#21
post #2

All this great content and their website is designed in a way that discourages people to peruse it. They really need a re-design and "relaunch" of their brand to flaunt the great things that they're doing.

You can't get to every single piece of functionality, but there's this -- https://archive.org/help/json.php -- and this -- https://archive.org/advancedsearch.php, which you can output as JSON/JSONP.

For instance, this broken piece of shit was my attempt to do something about the LMA interface -- http://www.archive-ui.org/#/. (select a show, reload the page, then it'll work. got busy and lost interest).

The interface is as bad as you say, but at least they give us the ability to do something about it ourselves.

Re: The People Behind the Wayback Machine

#22
post #18

I love the Wayback Machine, I wish they'd archive all pages though even those that don't wish to be archived.... keeping them away from public view until copyrights expire someday.

I don't know about this, to me archiving everything seems like a gross inefficiency. Most of the internet is spam and advertising, and of the rest, less than 5% is actually useful information or knowledge. Archiving books, scientific journals and the likes would seem much more useful, but obviously you'd run into copyright issues.

The trick is, of course, that it's nearly impossible to predict what will be useful to someone ahead of time. While you can probably sort out some of the spam, a comprehensive archiving project should probably avoid false positives when throwing things away.

Seems like a hard problem to solve. The low-hanging fruit would probably be detecting duplicates and combining them, which loses redundancy but handles all of those identical landing pages.

Re: The People Behind the Wayback Machine

#23
post #18

I love the Wayback Machine, I wish they'd archive all pages though even those that don't wish to be archived.... keeping them away from public view until copyrights expire someday.

I don't know about this, to me archiving everything seems like a gross inefficiency. Most of the internet is spam and advertising, and of the rest, less than 5% is actually useful information or knowledge. Archiving books, scientific journals and the likes would seem much more useful, but obviously you'd run into copyright issues.

Agree that highest priority should go to the "serious" stuff. However the most interesting part of a really old magazine or newspaper, for me, is the advertising. For example an early 80s computer ad, or a 50s railroad or airline ad. I find that stuff really fascinating, and it gives more of the flavor of the era. It might have a surprising amount of value to a historian or anthropologist.

Re: The People Behind the Wayback Machine

#25
post #3

The team at the Internet Archive, responsible for Wayback Machine, ArchiveTeam, the TV News Archive, and many other projects, are true gems. I had the chance to meet Brewster and many others whilst last in San Francisco and their passion is infectious. If anyone is interested, they have an open lunch on Fridays[1] where you get to see the church, the tech, and meet the team. Each team member and guest gives a few sen…

[deleted]

Re: The People Behind the Wayback Machine

#26
A lot of people here are confusing the Archive Team with the Internet Archive. They are not the same. IA is more polite, they always respect robots.txt and they will sometimes remove data if you ask politely. AT are self-described "rogue archivists" and their motto is "we are going to rescue your shit".

Re: The People Behind the Wayback Machine

#27

I love what the archive team does. I used their VM when the posterous backup effort was happening last year, and today I sent a link to my friend's now defunct posterous blog. But I know that the owner of days posterous page had no intent on keeping the page a going concern, and was happy to see it disposed of. In light of the recent Google "right to be forgotten" ruling, will there come a day when the right to be fo…

>will there come a day when the right to be forgotten will extend to archive.org? Sites can at any time opt out of being archived via a robots.txt exclusion (IA still keep their previous archives privately). However for public blogging sites operated by a third-party that's another matter.

The right to be forgotten is way different than a robots.txt file...

Re: The People Behind the Wayback Machine

#28

I love what the archive team does. I used their VM when the posterous backup effort was happening last year, and today I sent a link to my friend's now defunct posterous blog. But I know that the owner of days posterous page had no intent on keeping the page a going concern, and was happy to see it disposed of. In light of the recent Google "right to be forgotten" ruling, will there come a day when the right to be fo…

The "right to be forgotten" ruling was new because it talked about _linking to_ stuff. But Archive Team and the Internet Archive both make actual copies, so they are covered by copyright. If your friend want his Posterous posts gone he can serve an DMCA notice on them.

Re: The People Behind the Wayback Machine

#29
post #19

Earlier quoted context omitted.

That and they primarily archive public domain material and abandonware (apart from their web archiving project). They really couldn't be more different.

I think simply disregarding the web archiving is a bit of a cop out. It's interesting though that for the most part, nobody minds them redistributing loads of copyrighted material. Here's some reasons that come to mind: They web material was distributed for free in the first place. They're redistributing ad-ware, not stuff behind a paywall. (The same can be said of some TV shows and indeed I think TV show piracy if o…

And they also respect robots.txt.

None of these things in isolation necessarily makes what the IA does entirely legit under current copyright law; they effectively operate in something of a legal grey area. But add it all together and not many people are going to get upset--especially given that they'll remove material if asked to do so.

There have been a few legal cases http://en.wikipedia.org/wiki/Wayback_Machine but not many considering the scope of what they archive.

Re: The People Behind the Wayback Machine

#30

I love the Wayback Machine, I wish they'd archive all pages though even those that don't wish to be archived.... keeping them away from public view until copyrights expire someday.

Quite coincidentally, I was just now reading an interview with Brewster Kahle, from NewScientist (23 November 2002) - back when the Wayback Machine had only 100 terabytes archived.

He said: "I guarantee that in the future researchers will curse us for having missed something absolutely critical. But only people using the archive can tell us about mistakes in what we collect. There is a cheaper alternative concept, called 'dark archiving', which means that we should not give people access to them. But preservation without access is dangerous - there's no way of reviewing what's in there."

But later on, he mentioned that: "AltaVista was the first Internet search engine that tried to be a complete index of all the pages. But what really got me was that they threw away the original pages. That grated, no end."

Aside: Kahle was one of the founders, with Danny Hillis, of Thinking Machines - the company that created the fabulous 'Connection Machine'.

Post reply on HN