Live data from Hacker News

There’s a simple alternative to the current web

hapgood.us

71–80 of 148 posts

Re: There’s a simple alternative to the current web

#71
post #66
post #61

Earlier quoted context omitted.

Many people just don't have the clue to make usable local archives. You're also proposing mass copyright infringement which - stupidly in this example - is not legal.

If the system itself is federated the same way as the rest of the data, then it doesn't matter if it's legal or not. You can't make an omelette without breaking a few eggs and you can't fix the world without breaking a few laws.

>You can't make an omelette without breaking a few eggs and you can't fix the world without breaking a few laws.

Heh. The NSA should put this on a t-shirt and sell it.

Re: There’s a simple alternative to the current web

#72
I'd rather a page go offline than have it taken out of context. As if plagiarism wasn't bad enough already. (Yahoo Answers cough)

These crooks will even steal your copyright notice. It's quite possible the original content producers are offline because scraper thieves stole so much content that it's no longer possible to earn a living.

As an artist, this reminds me of the condescending attitude that gave us fake Rolexes, Facebook & North Korea's 28 state-approved haircuts. Either it's "just content" to stuff in a database somewhere or you understand the medium is the message too.

Re: There’s a simple alternative to the current web

#73

Clearly I'm a biased observer, but I really think people should take steps to archive stuff that is important to them. Of course it's terrible when large sites go offline and take vast swaths of the Internet with them, and we should continue to shame the ones that do it. At the same time, if something is really important to you, you shouldn't store it in the form of links to random third-party servers. One problem we…

My second reply, but, I think that this is really important.

We only need to look at early film history to know how easy it is to lose massive parts of our history.

Going back to old pages, I frequently get 404 results. For politically sensitive documents, the problem is much more widespread.

I would like something that not only archives pages I visit, but also versions them and tracks changes. If there was a bookmarking tool that did this, you could easily have an opt-in feature that shared content. This type of system would be a huge boost to something like the wayback machine.

Re: There’s a simple alternative to the current web

#74
post #61

Clearly I'm a biased observer, but I really think people should take steps to archive stuff that is important to them. Of course it's terrible when large sites go offline and take vast swaths of the Internet with them, and we should continue to shame the ones that do it. At the same time, if something is really important to you, you shouldn't store it in the form of links to random third-party servers. One problem we…

Many people just don't have the clue to make usable local archives. You're also proposing mass copyright infringement which - stupidly in this example - is not legal.

I don't think copyright infringement is the issue. Making a copy to view offline is always legal (well, fair use clauses sort-of make it mostly legal) and you are doing that with a web browser anyway, with cached pages, and cacheing proxies and similar devices also perform a similar function. The protections afforded by copyright all pertain to public reproduction and dissemination.

The WSJ point about restricting the number of free article views made in another reply is similarly nothing to do with copyright. This is more of a contract issue, and anyway, if I can view only N articles for free, I can only save those N, so it is really moot.

Re: There’s a simple alternative to the current web

#75

I'd rather a page go offline than have it taken out of context. As if plagiarism wasn't bad enough already. (Yahoo Answers cough ) These crooks will even steal your copyright notice. It's quite possible the original content producers are offline because scraper thieves stole so much content that it's no longer possible to earn a living. As an artist, this reminds me of the condescending attitude that gave us fake Rol…

As someone with a teensy bit of film background, I have to disagree. The number of early Hollywood films that were lost is astounding. This is a massive part of our visual history that is completely gone. It will never be restored.

With the current environment on the internet, with DRM'd video, music and text, I have to assume that we will lose far more from this time period than we ever had before.

While I don't pirate things (I'd rather just consume Creative Commons and Public Domain content), I wholehartedly support people who are trying to archive the things that are part of our collective culture. When I have kids, I'd like to be able to show them where they came from.

Re: There’s a simple alternative to the current web

#76
post #52

It’s interesting that Andreessen can’t see the solution, but perhaps expected. What a weird dig. It's neither expected, nor established that he can't see a solution. I'm not as smart as Andreessen and I could come up with half a dozen solutions. Author's favourite is fine but far, far from obvious. How viable is it to run your own federated wiki anyway? Are there packages for popular systems? Are there plugins for ma…

What he seems to be getting at is a much larger, more revolutionary approach to not just "the web", but "the internet" as we know it: https://en.wikipedia.org/wiki/Named_data_networking

Thats a very academic and static view of content - I don't see how that would work in todays hyper-dynamic environment, where the Ads that are displayed on a site are priced by millisecond real-time auctions before they are delivered to the user, and websites are single-page apps with REST APIs in the background. How would that work?

Re: There’s a simple alternative to the current web

#77
post #65

Earlier quoted context omitted.

Is saving a webpage to your local harddrive copyright infringement? The data is already on your local system when you view a webpage.

Copyright is a legal construct, not a technical construct. It doesn't matter where the data is. If a judge decides it's copyright infringement to save webpages, it will be. I can't imagine that the WSJ or any other paywalled institution wouldn't consider saving pages locally to be copyright infringement; how would they enforce a limit on article views?

The same way they do now? I can't save articles if I can't view them.

Re: There’s a simple alternative to the current web

#78

Clearly I'm a biased observer, but I really think people should take steps to archive stuff that is important to them. Of course it's terrible when large sites go offline and take vast swaths of the Internet with them, and we should continue to shame the ones that do it. At the same time, if something is really important to you, you shouldn't store it in the form of links to random third-party servers. One problem we…

My second reply, but, I think that this is really important. We only need to look at early film history to know how easy it is to lose massive parts of our history. Going back to old pages, I frequently get 404 results. For politically sensitive documents, the problem is much more widespread. I would like something that not only archives pages I visit, but also versions them and tracks changes. If there was a bookmar…

It clearly hasn't occurred to you that private collections evaporate over time, just like public ones.

> I would like something that not only archives pages I visit, but also versions them and tracks changes.

That's a huge storage requirement, you must realize this. If you're an avid Web browser, and if every archived page had to look as it originally looked (i.e. all the linked resources) you could accumulate several terabytes per week.

> This type of system would be a huge boost to something like the wayback machine.

Here's the road to madness. Someone, aware of the rapidly declining cost of storage, rebuilds the Wayback machine based on your scheme, with the intent of archiving every Web page in existence, including all required resources so the pages look just as they originally did. Then, as the project approaches completion, this genius says, "For the next phase, I need to archive the Wayback machine itself." At that point, as the implications of what he's said occur to him, a strange look crosses his face and his imagination begins writing checks his intellect can't cash.

Re: There’s a simple alternative to the current web

#79
post #15

The fact that he thinks a federated wiki would be "simple" or "easy" leads me to believe he has not actually thought through the details of how it would work in practice.

It would not be simple or easy, but crypto-currency blockchains make it more possible than ever.

What does this mean? I can't understand the relationship between the two ideas?

Re: There’s a simple alternative to the current web

#80

Clearly I'm a biased observer, but I really think people should take steps to archive stuff that is important to them. Of course it's terrible when large sites go offline and take vast swaths of the Internet with them, and we should continue to shame the ones that do it. At the same time, if something is really important to you, you shouldn't store it in the form of links to random third-party servers. One problem we…

Slightly related to this, the other day I tried to download every youtube video in my watch history. Turns out it's pretty much impossible. The Youtube API only delivers about 20 results, which is a bug that has existed for about 2 years. I tried manually loading the watch history page, and was only able to get about 1000 out of ~8000 results. When I selected those 1000 results and tried to add them to a playlist, th…

> The Youtube API only delivers about 20 results, which is a bug that has existed for about 2 years.

That may not be a bug. It might be a way to limit the traffic created by scrapers who submit requests and then download all the videos listed on the result page.

Post reply on HN