Live data from Hacker News

There’s a simple alternative to the current web

hapgood.us

111–120 of 148 posts

Re: There’s a simple alternative to the current web

#111

Earlier quoted context omitted.

To run the experiment, I am going to be drawing a few thousand links at random from the entire pool of Pinboard bookmarks. This will include private bookmarks, which make up about half the Pinboard collection. You chose to include everyone's private bookmarks in your research without asking their consent? What? I will publish some aggregate information about what I find, and use it to seek glory, and persuade people…

Why would I need consent to study the global link rot rate? Publishing it reveals no information about users, either individually or in the aggregate. I've made an effort to let anyone who wants to opt out of the research, because I know people can have strong feelings about privacy. I agree with you that publishing an 'anonymized' dataset would be a violation of privacy guarantees. I wouldn't even do it for public b…

"Private" means "private," not "private unless Maciej wants to study them."

I didn't get an email saying "There's a chance I might select your private bookmarks and examine them." A blog post doesn't count when you're messing with people's private data. Certainly it should be opt-in, and not opt-out?

You're doing this for a noble purpose, but for what it's worth, this is the first time my trust in you has ever felt violated.

People have entrusted you with years worth of private data, and you just asked, "Why should I ask permission to study their private links?"

Actually, as far as I can tell, your comment seems to implicitly assume that you already have consent to examine all private links, and that asking consent would only be necessary if you were planning on publishing something that might reveal some of their private links. Isn't that the opposite of privacy?

Re: There’s a simple alternative to the current web

#112

Earlier quoted context omitted.

My second reply, but, I think that this is really important. We only need to look at early film history to know how easy it is to lose massive parts of our history. Going back to old pages, I frequently get 404 results. For politically sensitive documents, the problem is much more widespread. I would like something that not only archives pages I visit, but also versions them and tracks changes. If there was a bookmar…

What do you think of this generalized architecture? HARDWARE: differs depending on whether you want local search/analytics or just network storage. For mobile use, either a VPN back to your personal home/cloud server, or a hackable wifi hard drive proxy, e.g. Seagate Wirless Plus + HackGFS. For non-analytics home use, hackable router with USB3 storage and Linux software RAID, connected to a USB3 drive chassis with ro…

Holy living God.

Solves my problem, then some. Also provides an alternative to search engines as the goto for internet browsing. Brilliant stuff. I can't tell you how grateful I am for your work on this.

The only thing left is FLOSS version control for sound and video editing, and an effective "publish to BitTorrent" feature, and then we can pretty much put this "Web 2.0" crap to bed.

EDIT: Out of curiosity, is there any reason that analytics couldn't be done with a dedicated PC, or do you think it requires server hardware to run effectively?

Re: There’s a simple alternative to the current web

#113
post #77
post #65

Earlier quoted context omitted.

Copyright is a legal construct, not a technical construct. It doesn't matter where the data is. If a judge decides it's copyright infringement to save webpages, it will be. I can't imagine that the WSJ or any other paywalled institution wouldn't consider saving pages locally to be copyright infringement; how would they enforce a limit on article views?

The same way they do now? I can't save articles if I can't view them.

You can view all wsj articles if you use google as referer

Re: There’s a simple alternative to the current web

#114

Earlier quoted context omitted.

Why would I need consent to study the global link rot rate? Publishing it reveals no information about users, either individually or in the aggregate. I've made an effort to let anyone who wants to opt out of the research, because I know people can have strong feelings about privacy. I agree with you that publishing an 'anonymized' dataset would be a violation of privacy guarantees. I wouldn't even do it for public b…

"Private" means "private," not "private unless Maciej wants to study them." I didn't get an email saying "There's a chance I might select your private bookmarks and examine them." A blog post doesn't count when you're messing with people's private data. Certainly it should be opt-in, and not opt-out? You're doing this for a noble purpose, but for what it's worth, this is the first time my trust in you has ever felt v…

I don't share your outrage. The blog post linked clearly states the scope and purpose, the researcher says aggregated data will not be released and individual user data is unnecessary for the purposes described. He then offers those that are still wary of the research an option to not have their data used. Maybe I have a blind spot but this seems pretty straightforward and harmless to individual privacy.

Re: There’s a simple alternative to the current web

#115
post #114

Earlier quoted context omitted.

"Private" means "private," not "private unless Maciej wants to study them." I didn't get an email saying "There's a chance I might select your private bookmarks and examine them." A blog post doesn't count when you're messing with people's private data. Certainly it should be opt-in, and not opt-out? You're doing this for a noble purpose, but for what it's worth, this is the first time my trust in you has ever felt v…

I don't share your outrage. The blog post linked clearly states the scope and purpose, the researcher says aggregated data will not be released and individual user data is unnecessary for the purposes described. He then offers those that are still wary of the research an option to not have their data used. Maybe I have a blind spot but this seems pretty straightforward and harmless to individual privacy.

Is it okay for any owner of a website to go through their userbase's private data, simply because they own that website?

I don't mean "okay" in a legal sense, but rather a moral sense.

It reminds me quite a lot of http://i.imgur.com/5quY1Iq.png except that the difference is that Maciej is a researcher and isn't disclosing the data to people. However, he's still going through people's private stuff. Notifying them that you're planning to go through their private stuff is the most basic common courtesy; it's why landlords can't simply walk in to a tenant's house whenever they feel like it, even though they own the property.

Let's put it another way: I didn't know Maciej was the type of person to trawl through people's private information that they trusted him with. If I did, I would've investigated other options for a bookmarking site a couple years ago, or would've written my own, and I wouldn't have breathlessly recommended Pinboard to whoever would listen. The recommendation would be more like "Pinboard is great, but the owner likes to look at your stuff, even if it's marked 'private,' so keep that in mind."

Re: There’s a simple alternative to the current web

#116

Earlier quoted context omitted.

Why would I need consent to study the global link rot rate? Publishing it reveals no information about users, either individually or in the aggregate. I've made an effort to let anyone who wants to opt out of the research, because I know people can have strong feelings about privacy. I agree with you that publishing an 'anonymized' dataset would be a violation of privacy guarantees. I wouldn't even do it for public b…

"Private" means "private," not "private unless Maciej wants to study them." I didn't get an email saying "There's a chance I might select your private bookmarks and examine them." A blog post doesn't count when you're messing with people's private data. Certainly it should be opt-in, and not opt-out? You're doing this for a noble purpose, but for what it's worth, this is the first time my trust in you has ever felt v…

I think the tension between us is that you think 'private' means 'visible only to me', and I think private means 'never displayed to any other user, or on a public page'.

There are a thousand routine tasks that require me to have unrestricted access to bookmarks and URLs. I try to be as uninvasive as I can about it, but you have no way of verifying that.

If you want something to remain truly private to you—and I say this in full sympathy to your feelings—don't put it on a stranger's computer. Where there's a server, there's an admin.

Re: There’s a simple alternative to the current web

#117
post #78

Earlier quoted context omitted.

My second reply, but, I think that this is really important. We only need to look at early film history to know how easy it is to lose massive parts of our history. Going back to old pages, I frequently get 404 results. For politically sensitive documents, the problem is much more widespread. I would like something that not only archives pages I visit, but also versions them and tracks changes. If there was a bookmar…

It clearly hasn't occurred to you that private collections evaporate over time, just like public ones. > I would like something that not only archives pages I visit, but also versions them and tracks changes. That's a huge storage requirement, you must realize this. If you're an avid Web browser, and if every archived page had to look as it originally looked (i.e. all the linked resources) you could accumulate severa…

For compressed static html and maybe images, you could store your entire history on a flash drive nowdays. See this: http://memkite.com/blog/2014/04/01/technical-feasibility-of-...

Re: There’s a simple alternative to the current web

#118

Earlier quoted context omitted.

"Private" means "private," not "private unless Maciej wants to study them." I didn't get an email saying "There's a chance I might select your private bookmarks and examine them." A blog post doesn't count when you're messing with people's private data. Certainly it should be opt-in, and not opt-out? You're doing this for a noble purpose, but for what it's worth, this is the first time my trust in you has ever felt v…

I think the tension between us is that you think 'private' means 'visible only to me', and I think private means 'never displayed to any other user, or on a public page'. There are a thousand routine tasks that require me to have unrestricted access to bookmarks and URLs. I try to be as uninvasive as I can about it, but you have no way of verifying that. If you want something to remain truly private to you—and I say…

It seems like "private" could be defined as, "This is my stuff, and if you want access to it, then come ask me. It's okay if you accidentally access it, but if you want to intentionally look through it, ask me first." It seems hard to argue that most people wouldn't feel that way about their own private stuff.

I fully understood the implications of giving you the data. I'm a fan of your work and your writing, and I had full confidence in your stewardship of my data. Essentially, I was totally okay with you being the admin, or anyone you decided to hire, and I trusted you to take reasonable steps not to look through your users' private data unless it was to track down some bug, test some new feature, or some other incidental task that was unrelated to analyzing that private data.

What I didn't expect was that you'd specifically and intentionally create a program whose sole purpose was to analyze private user data and report on the results.

Why didn't I expect that? The only answer is that I should have expected that. I just didn't realize you were that type of developer. It was a bit shocking that someone who has trumpeted the benefits of sticking with businesses that haven't taken VC investment would explicitly break their users' trust like this.

In this case, you have both the legal right and the moral high ground. But intentionally seeking through your own users' private data without getting consent isn't something that can easily be forgotten.

Re: There’s a simple alternative to the current web

#119
post #114

Earlier quoted context omitted.

I don't share your outrage. The blog post linked clearly states the scope and purpose, the researcher says aggregated data will not be released and individual user data is unnecessary for the purposes described. He then offers those that are still wary of the research an option to not have their data used. Maybe I have a blind spot but this seems pretty straightforward and harmless to individual privacy.

Is it okay for any owner of a website to go through their userbase's private data, simply because they own that website? I don't mean "okay" in a legal sense, but rather a moral sense. It reminds me quite a lot of http://i.imgur.com/5quY1Iq.png except that the difference is that Maciej is a researcher and isn't disclosing the data to people. However, he's still going through people's private stuff. Notifying them tha…

I like the image of myself sitting at the computer with a box of bon-bons, lazily hitting 'next' on the special admin page that selects only the juiciest private bookmarks for my delectation.

The reality is less fun. I have to look at (potentially) private bookmarks when:

- someone's import file fails to parse, or has an encoding problem

- there are garbled or missing results for a search query

- I need to answer questions like 'how much disk space does a typical bookmark use', so I can provision what I need

- there's a bug in the fulltext parser

- the twitter API client misses some tweets or mutilates a URL

- the pinboard API is misbehaving in one of a thousand ways

- I need to verify that backups I make actually contain everything they're supposed to

- I want to find and fix privacy bugs!

Along with a thousand other scenarios that will be familiar to anyone who has ever had the misfortune to import, format and store user-provided data.

Anywhere bookmarks come into or leave the site, or are displayed on the site, there will be bugs. If I tried to enforce some kind of viewing restrictions on myself, it would just introduce an additional layer of bugs while making my job completely intractable.

I'm not a landlord walking into a tenant's house without permission. I'm a hotel manager, doing my best to be discreet, but ultimately requiring full access to everything in order to do my job. I'm going in to check the sprinklers and fire alarm even if you've left the 'do not disturb' sign on.

This will be the case on any outside site you use, even one that makes sweeter promises to you than I ever did. Please think twice, and then three times, before uploading your data anywhere if you have these kinds of expectations.

Re: There’s a simple alternative to the current web

#120

Earlier quoted context omitted.

Is it okay for any owner of a website to go through their userbase's private data, simply because they own that website? I don't mean "okay" in a legal sense, but rather a moral sense. It reminds me quite a lot of http://i.imgur.com/5quY1Iq.png except that the difference is that Maciej is a researcher and isn't disclosing the data to people. However, he's still going through people's private stuff. Notifying them tha…

I like the image of myself sitting at the computer with a box of bon-bons, lazily hitting 'next' on the special admin page that selects only the juiciest private bookmarks for my delectation. The reality is less fun. I have to look at (potentially) private bookmarks when: - someone's import file fails to parse, or has an encoding problem - there are garbled or missing results for a search query - I need to answer que…

I'm scratching my head wondering if I'm being unclear. Let me try again:

You created a program whose sole purpose was to analyze private user data and report on the results. (In fact, not merely "private user data," but "data which users explicitly marked as 'private.'")

What you did was equivalent to a hotel manager sending employees to peep into 1,000 random rooms and compile a detailed report of what those rooms contained and what its occupants were doing, and then claiming it was for the betterment of all hotels. Yes, that may be true, and the data may be quite helpful, but people still expected their rooms to be private.

Intentions matter. You weren't accessing the private bookmarks in order to fix a bug or test a new feature.

Post reply on HN