Live data from Hacker News

There’s a simple alternative to the current web

hapgood.us

91–100 of 148 posts

Re: There’s a simple alternative to the current web

#91
post #55
post #18

Earlier quoted context omitted.

> And disappearing is the default, natural state of things. Books, clay tablets, scrolls, engraved stone, to which humans owe their entire knowledge of their premodern history, seem to have put up pretty well against entropy. The same is not the case for information disseminated in a controlled manner from privately owned servers. > If I see some people playing music on a corner and return the next night to see they'…

"Books, clay tablets, scrolls, engraved stone, to which humans owe their entire knowledge of their premodern history, seem to have put up pretty well against entropy." Only the ones that have survived. For every book or tablet we have, there are certainly tens of THOUSANDS of which every copy ever published has been lost - most of those are ephemera that wouldn't mean much to us anyways, but the lost also include thi…

You're absolutely right. Maybe it's a fool's errand to try and hold onto the past.

But many people consider those losses to be an immeasurable tragedy.

Re: There’s a simple alternative to the current web

#92

I'd rather a page go offline than have it taken out of context. As if plagiarism wasn't bad enough already. (Yahoo Answers cough ) These crooks will even steal your copyright notice. It's quite possible the original content producers are offline because scraper thieves stole so much content that it's no longer possible to earn a living. As an artist, this reminds me of the condescending attitude that gave us fake Rol…

As someone with a teensy bit of film background, I have to disagree. The number of early Hollywood films that were lost is astounding. This is a massive part of our visual history that is completely gone. It will never be restored. With the current environment on the internet, with DRM'd video, music and text, I have to assume that we will lose far more from this time period than we ever had before. While I don't pir…

I'm not against reproduction if care is taken to preserve what can be preserved, as close to the original as possible while giving credit, compensating creators, etc. In your example, I think reproduction/conservation technology was available but the studios couldn't justify spending the time/money it would cost to preserve their entire library. Who would have paid for that? I don't know. Supposedly, half of Van Gogh's entire lifetime output was burned because he was too poor to find a place to store it. At the same time, I'd rather see a bug-eaten Van Gogh with fugitive reds long faded away, than a flat, lifeless high-def copy. I suppose it's a complex subject and each work is unique. Sometimes I'm thrilled when I can find an old page in "Wayback Machine" but what they manage to save is typically broken and low quality.

Re: There’s a simple alternative to the current web

#93
post #80

Earlier quoted context omitted.

Slightly related to this, the other day I tried to download every youtube video in my watch history. Turns out it's pretty much impossible. The Youtube API only delivers about 20 results, which is a bug that has existed for about 2 years. I tried manually loading the watch history page, and was only able to get about 1000 out of ~8000 results. When I selected those 1000 results and tried to add them to a playlist, th…

> The Youtube API only delivers about 20 results, which is a bug that has existed for about 2 years. That may not be a bug. It might be a way to limit the traffic created by scrapers who submit requests and then download all the videos listed on the result page.

Hrm, then why offer the feature?

I'm kind of still a beginner when it comes to APIs, but it seems disingenuous to offer an API feature and then intentionally break it so that it can't be used.

Re: There’s a simple alternative to the current web

#94

Earlier quoted context omitted.

As someone with a teensy bit of film background, I have to disagree. The number of early Hollywood films that were lost is astounding. This is a massive part of our visual history that is completely gone. It will never be restored. With the current environment on the internet, with DRM'd video, music and text, I have to assume that we will lose far more from this time period than we ever had before. While I don't pir…

I'm not against reproduction if care is taken to preserve what can be preserved, as close to the original as possible while giving credit, compensating creators, etc. In your example, I think reproduction/conservation technology was available but the studios couldn't justify spending the time/money it would cost to preserve their entire library. Who would have paid for that? I don't know. Supposedly, half of Van Gogh…

I'm having a little trouble understanding your perspective. Copying films is good, letting bugs eat Van Gogh is good, but copying websites is bad?

Well, luckily, with digital technology, we can copy things flawlessly with very little cost. Unfortunately, most content creators are still stuck trying to adapt physical distribution models to the information age, which is why we're stuck with DRM. You can't say "The Medium is the Message" and then get angry because you're producing content for a medium that is infinately copyable.

We need to move to a model of perceiving data as holographic. Especially with the advent of blockchain technology, we're increasingly moving to a model where every node in a network contains the entire network. Trying to adapt 20th century Disney copyright to that paradigm is stupid.

Re: There’s a simple alternative to the current web

#95
post #78

Earlier quoted context omitted.

My second reply, but, I think that this is really important. We only need to look at early film history to know how easy it is to lose massive parts of our history. Going back to old pages, I frequently get 404 results. For politically sensitive documents, the problem is much more widespread. I would like something that not only archives pages I visit, but also versions them and tracks changes. If there was a bookmar…

It clearly hasn't occurred to you that private collections evaporate over time, just like public ones. > I would like something that not only archives pages I visit, but also versions them and tracks changes. That's a huge storage requirement, you must realize this. If you're an avid Web browser, and if every archived page had to look as it originally looked (i.e. all the linked resources) you could accumulate severa…

Several terabytes per week is a huge overestimate. Right now I'm getting 12 Mbps download speed (sadly typical for US broadband). If I saturated the connection, and if no one throttled me, I could download 900 gigabytes in a week. That would be some intense web surfing.

My practical experience is that you need about 1.5 MB per URL for storing large numbers of web pages, if you exclude video.

Re: There’s a simple alternative to the current web

#96
post #80

Earlier quoted context omitted.

> The Youtube API only delivers about 20 results, which is a bug that has existed for about 2 years. That may not be a bug. It might be a way to limit the traffic created by scrapers who submit requests and then download all the videos listed on the result page.

Hrm, then why offer the feature? I'm kind of still a beginner when it comes to APIs, but it seems disingenuous to offer an API feature and then intentionally break it so that it can't be used.

> Hrm, then why offer the feature?

Because people can still get 20 results. I mean, 20 results is way better than no results.

> I'm kind of still a beginner when it comes to APIs, but it seems disingenuous to offer an API feature and then intentionally break it so that it can't be used.

It's not broken. It returns 20 results. Maybe Google decided that was enough of a hit on their database. One could also argue that most people wouldn't want more than 20 results on their small-screen Android device served by a slow connection.

Re: There’s a simple alternative to the current web

#97

I will pay good money for a Chrome extension that does the following: 1) I can select (or do select all) Chrome bookmarks that I want to keep offline page backups/archives of (saved to google drive or dropbox or some such). 2) Whenever I want, instead of seeing the current online version of that bookmarked page, I can look up the originally bookmarked archived page. 3) It allows me to choose the level of links to the…

Pinboard archiving costs about $25 a month. Not sure it does the deep link archiving.

Re: There’s a simple alternative to the current web

#98
post #96

Earlier quoted context omitted.

Hrm, then why offer the feature? I'm kind of still a beginner when it comes to APIs, but it seems disingenuous to offer an API feature and then intentionally break it so that it can't be used.

> Hrm, then why offer the feature? Because people can still get 20 results. I mean, 20 results is way better than no results. > I'm kind of still a beginner when it comes to APIs, but it seems disingenuous to offer an API feature and then intentionally break it so that it can't be used. It's not broken. It returns 20 results. Maybe Google decided that was enough of a hit on their database. One could also argue that m…

https://code.google.com/p/gdata-issues/issues/detail?can=2&s...

The number of results returned is random. Some people report seeing videos watched on computers but not locally. One person says it works.

That's not a limit. That's broken.

Re: There’s a simple alternative to the current web

#99
Link rot is a serious problem: http://www.gwern.net/Archiving%20URLs#link-rot

>In a 2003 experiment, Fetterly et al. discovered that about one link out of every 200 disappeared each week from the Internet. McCown et al. (2005) discovered that half of the URLs cited in D-Lib Magazine articles were no longer accessible 10 years after publication [the irony!], and other studies have shown link rot in academic literature to be even worse (Spinellis, 2003, Lawrence et al., 2001). Nelson and Allen (2002) examined link rot in digital libraries and found that about 3% of the objects were no longer accessible after one year.

>Bruce Schneier remarks that one friend experienced 50% linkrot in one of his pages over less than 9 years (not that the situation was any better in 1998), and that his own blog posts link to news articles that go dead in days; the Internet Archive has estimated the average lifespan of a Web page at 100 days. A Science study looked at articles in prestigious journals; they didn’t use many Internet links, but when they did, 2 years later ~13% were dead. The French company Linterweb studied external links on the French Wikipedia before setting up their cache of French external links, and found - back in 2008 - already 5% were dead. (The English Wikipedia has seen a 2010-2011 spike from a few thousand dead links to ~110,000 out of ~17.5m live links.) The dismal studies just go on and on and on (and on). Even in a highly stable, funded, curated environment, link rot happens anyway. For example, about 11% of Arab Spring-related tweets were gone within a year (even though Twitter is - currently - still around).

Re: There’s a simple alternative to the current web

#100
post #82

Earlier quoted context omitted.

You can use youtube-dl to do this: $ youtube-dl ':ythistory' -u USER -p PWD --write-pages Of course, you won't get your entire history since g00gle would rather mine all your personal data itself, and never share it back with you again.

I tried this earlier. The result: [youtube:history] playlist Youtube Watch History: Collected 0 video ids (downloading 0 of them)

Just a thought, do you have google two factor auth turned on? Maybe you need to generate an application password and use that to get access.
Post reply on HN