Live data from Hacker News

Online tracking: A 1-million-site measurement and analysis

webtransparency.cs.princeton.edu

241–250 of 272 posts

Re: Online tracking: A 1-million-site measurement and analysis

#241

Earlier quoted context omitted.

How about not developing applications in the browser? Its about linked documents. Not angular-17 MVVM async session persistence in indexdb with websql and asm.js rendering webgl for a spinning teapot.

OK. So you get your 'document only' internet. Where do we put all the other stuff? I have a bunch of 'non-document' websites that are essential to me now. What happens under your new regime? Someone reimplements them all as native apps? The point I'm making is just because your internet is document-only please don't assume mine or other people's are.

Because it is so, it must be so.

Every web app is written in perl/php, therefore we will keep using perl/php. - What will you do, rewrite the apps? internet != web.

This cludge of in-browser tech everyone is pursuing already comes with so much suffering for the developer. But the enthusiasm to embrace garbage like religion is just unbelievable. (yeah js everything, spotify lol)

I think the current webapp tech stack is not doing our generation justice. Yes google has very strong interests in maintaining status quo since it dominates it and so do many other giants. But good, user friendly and maintainable software looks different.

Why not break the completely misused model of documents for apps ? there is no document semantics in 1000x elements riddled with js callbacks.

But dont let my cynicism annoy you, its the resignation talking. Imagine what cool tech we would have if something was started in the 90s (no, not java applets) and was all grown up now.

But every big company is now a walled garden provider. Just think about UI toolkits, would spotify be in Chromeframe otherwise?.

Re: Online tracking: A 1-million-site measurement and analysis

#242
post #218
post #211

Earlier quoted context omitted.

I'm not familiar with WebRTC. What's the use case there? I can't remember ever wanting to create an in-browser p2p connection on my local network. What would it be used for?

Please read the Chrome 48 release notes, WebRTC's default behavior has changed. https://groups.google.com/d/msg/discuss-webrtc/_5hL0HeBeEA/H...

But it still ignores the proxy settings and will use STUN to discover your "external IP". Thus users that think they are using a proxy end up not actually doing so.

Re: Online tracking: A 1-million-site measurement and analysis

#243

Earlier quoted context omitted.

What can be done by the browser vendors such as Mozilla, Google, and Microsoft? To prevent fingerprinting, your browser has to disable all sorts of useful modern JavaScript API's (e.g., WebRTC) by default, prevent spurious HTTP requests (e.g., to prevent abusing @font-face to find out which fonts are installed), and pretend you are an American using the most popular web browser of the moment (i.e., hide the user's pr…

Personally I think there are so many of these APIs that for the browser to try to prevent the ability to fingerprint is putting the genie back in the bottle. But there is one powerful step browsers can take: put stronger privacy protections into private browsing mode, even at the expense of some functionality. Firefox has taken steps in this direction https://blog.mozilla.org/blog/2015/11/03/firefox-now-offers-... Tr…

That's not effective, because very few people use private mode.

Re: Online tracking: A 1-million-site measurement and analysis

#244

Coauthor here. I lead the research team at Princeton working to uncover online tracking. Happy to answer questions. The tool we built to do this research is open-source https://github.com/citp/OpenWPM/ We'd love to work with outside developers to improve it and do new things with it. We've also released the raw data from our study.

In your paper you say,

"When using the headless configuration, we are able to run up to 10 stateful browser instances on an Amazon EC2 “c4.2xlarge” virtual machine."

Also it seems like you ran the crawl only in the month of January this year, and crawled about 90 million pages. Were you able to do that on the single AWS instance, using Firefox via Selenium? What do you think the performance would have been just issuing raw requests?

Just interested because I'm currently building a crawler and am trying to decide if Selenium would be worth it performance wise.

Re: Online tracking: A 1-million-site measurement and analysis

#245

Coauthor here. I lead the research team at Princeton working to uncover online tracking. Happy to answer questions. The tool we built to do this research is open-source https://github.com/citp/OpenWPM/ We'd love to work with outside developers to improve it and do new things with it. We've also released the raw data from our study.

nice Arvind!

Re: Online tracking: A 1-million-site measurement and analysis

#246

Earlier quoted context omitted.

Consider the implications of what this means though. If sites are not free to innovate, things like Github and Gmail wouldn't exist . They only reason we aren't stuck with a Hotmail interface circa 2002 is because people were able to innovate on the web. To lock down CSS (or Javascript, there's no reason I can think of you would lock CSS and not Javascript) to a specific set of capabilities is both a statement that i…

The fact of standard templates needn't prevent the possibility of novel templates. But it ought make the prospect slightly more user-controllable. Design-by-committee isn't the alternative to design-by-fuckwits, the present mode. Github and Gmail are both tools which now face the dilemma of gratuitous changes -- many of the recent innovations haven't done much for usability, for numerous reasons (familiarity itself i…

> The fact of standard templates needn't prevent the possibility of novel templates.

If your stance is "provide well established default templates, but don't enforce their use", then I have no disagreement. That's not how I interpreted "I've considered what might be necessary to dispose of server-side CSS."

> Github, Gmail, Google Maps, etc., are largely the exception to long-form informational content pages. I'm OK with an explicit "app mode" for such sites. But 99.999999% of what I read would do vastly better with uniform presentation.

I think that depends heavily on what you use the web for. You and I likely read a lot on the web. Some people might stick largely to Facebook and Gmail. There are people that spend a lot of time in Github, and others that spend very little. Some people use a lot of online organizational and collaboration tools, others none.

> More attention to content and semantic construction. Less to layout frippery.

What you call layout frippery, someone else desires. This sounds suspiciously like remaking the web for your use cases, not for general use cases (which are always changing). But I'm not sure there's even a problem to address, you already addressed through referencing "readability modes" as an example of presentation styles that do work well. Why isn't that your solution to this perceived problem?

It feels like you're trying to achieve the equivalent of forcing all the printers to agree to not print magazines that don't conform to someone's opinion of what a good magazine is. I'm just not sure why that's even desirable.

> Something tells me you'll not be convinced.

No, not yet, if I understand your position correctly.

Re: Online tracking: A 1-million-site measurement and analysis

#247
post #230

Earlier quoted context omitted.

Just imagine, if the audio stack exposes the volume level, that's roughly 7.5 bits of uniqueness to contribute to the 33 required to uniquely identity any person on Earth (not that you can expect it to be uniformly distributed, and thus fully usable).

that doesn't quite work since the volume level changes frequently.

Yes, it's a poor example in that respect. I meant it more as an explanation of how different attributes of a source all contribute a little but to providing a unique identifier, but you are correct that it's much less useful if the attribute is not static.

Re: Online tracking: A 1-million-site measurement and analysis

#248
post #70

Earlier quoted context omitted.

That doesn't sound too far from Desktop Computing as a Service. It will be a sad day when I have to pay $9/month to be able to log into my desktop.

I mostly agree. Ideally I'd like to have a minimal OS and file set on my local machine (for offline and poor connectivity scenarios), that automatically syncs with my own, encrypted cloud system, such that I can (at my own discretion) update the OS from controlled sources (e.g. git). But I don't think there is enough interest from others for such a system, and I'm occupied with enough other projects that I won't be a…

I think that it depends on your use-case. I use Google Photos to store my media files, Github to store application configuration & source code of my applications, Chrome to store bookmarks, passwords, Spotify to save & listen music etc. Even if I lost my computer now, I would easily setup my desktop environment again.

Re: Online tracking: A 1-million-site measurement and analysis

#249

Earlier quoted context omitted.

As cm3 notes, usability is as much if not more a concern than privacy and security . Though I'd not dislodge any of these three from a position of high primacy. There's a risk / frequency trade off with all of these. Privacy can be quite possibly costly or fatal, though slightly more rare. Not so rare though that 20% of all Web users in a US Department of Commerce survey (see my recent comments history) report known…

> My most common response when landing on a website is to sigh, roll my eyes, and dump it to something more readable. Firefox's Reader Mode. Pocket. Straight ASCII text. w3m. > My half-serious response to this is to create a new web browser embodying these and a few other principles. In all seriousness, I wonder if spoofing a mobile client (easily done through most browser developer console's or an extension) might i…

The majority of my browsing is mobile these days. 10" tablet.

Even sites which are otherwise well-designed (Aeon and Medium come to mind) insist on dark-pattern behavior such as fixed headers/footers. Again: straight to reader-mode for that.

Except for the sites which break that. Violet Blue's Peerlyst comes to mind: https://plus.google.com/104092656004159577193/posts/PWuVmx2r...

(Screenshots contrasting site and a Reader Mode session included.)

I've written directly with the site designer who seems utterly insensate to why 14pt font isn't in fact a majickal solution to all readability problems.

HN itself is only barely usable.

Re: Online tracking: A 1-million-site measurement and analysis

#250
post #211

Earlier quoted context omitted.

> For WebRTC, browsers could block local addresses. That would defeat a huge selling point of WebRTC, the ability to create in-browser p2p connections over the user's local network.

I'm not familiar with WebRTC. What's the use case there? I can't remember ever wanting to create an in-browser p2p connection on my local network. What would it be used for?

WebRTC is a secure real-time protocol for audio, video, and data.

You'd want to use it any time you want a high-speed network connection with another user. For example, a multiplayer game or video teleconference.

Post reply on HN