Live data from Hacker News

Invisible Click Tracking Using “Empty” UTF-8 Characters

kaspars.net

21–30 of 38 posts

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#22

i don't see any variables in the url when i run the inspector.

Copy and paste the link into python, like so:

    >>> repr('http://kaspars.net/blog/web-development/invisible-click-tracking')
    "'http://kaspars.net/blog/web-development/invisible-click-tracking\\xe2\\x80\\x8b'"
EDIT: Yeah, they didn't show up anywhere in the Safari inspector for me either.

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#23
post #3

I actually like this. Assuming you agree there are non-creepy reasons to track clicks (not a given on HN) then this lets you keep your urls nice and clean when people just want to copy and paste them. I imagine it might break in some scenarios (the url ends up with junk in it from bad unicode conversion) but it's up to you to be permissive in the urls you accept.

I usually remove trailing spaces from URLs, and a lot of auto-linkifying processes (if they're Unicode-aware) will likely strip them too, so I predict this method won't be robust enough to survive a lot of the handling that URLs are often subject to.

If I'm filling out some form fields with URLs, I strip trailing (and leading) spaces too.

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#24

Isn't this exploiting the same kind of Unicode ambiguity that allowed phishing sites to impersonate trusted domains by substituting certain latin characters with identical-looking cyrillic equivalents? I would expect this capability to last long in the wild.

These domains are displayed in their punycode notation, at least in Firefox, so that doesn't work anymore.

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#25
post #3

I actually like this. Assuming you agree there are non-creepy reasons to track clicks (not a given on HN) then this lets you keep your urls nice and clean when people just want to copy and paste them. I imagine it might break in some scenarios (the url ends up with junk in it from bad unicode conversion) but it's up to you to be permissive in the urls you accept.

There are non-creepy reasons for click-tracking--e.g., you're a nonprofit supporting a cause, trying to identify preferences of your supporters. But caution: it might become creepy when you /appear/ to be hiding the tracking from your supporters. Why not give them readable URLs and be totally upfront about what you're doing?

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#26
post #16

Anyone have a suggestion of how this would be used in practice for uniquely identifying something? So there are a certain number of non-width space characters. As far as I can find in the links in the OP, U+FEFF, U+180E, U+200B, U+200C, U+200D would make 5. So we have at least 5 values to work with, which would make... 120 combinations if they're ordered differently? Surely we would need more if we want to uniquely i…

5 5 5 5 5 if you limit to a 5 character word!

You have 5 choices for each character. If you have limit to 5 character words you can assume that all 5 char, 4 char, 3, 2, 1 and 0 char words are included too.

I get:

    5^0+5^1+5^2+5^3+5^4+5^5 = 3,906
But I may have totally missed something.

Edit: looking a little closer I see some italics in your 55555 so I guess you have 5 * 5 * 5 * 5 * 5

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#27
post #19

Just to point out that hovering over the link shows URL-encoded codepoints, i.e. %u200B (in Firefox, for me, at least). I do think it's a better solution than the existing approach with GET parameters which does look clumsy.

Interesting. Chrome on Windows doesn't.

Here neither. But even then, for very long URLs (can easily happen with that SEO crap), just hide your tracking byte so that when Chrome/FF truncates the URL for display, the bytes will be inside the hidden area.

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#28

Just to point out that hovering over the link shows URL-encoded codepoints, i.e. %u200B (in Firefox, for me, at least). I do think it's a better solution than the existing approach with GET parameters which does look clumsy.

the %u200B is not being displayed while hovering the link by Firefox 35.0.1 / Linux (but it is displayed with beta 36 on the same computer)

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#29

Isn't this exploiting the same kind of Unicode ambiguity that allowed phishing sites to impersonate trusted domains by substituting certain latin characters with identical-looking cyrillic equivalents? I would expect this capability to last long in the wild.

would->wouldn't

Re: Invisible Click Tracking Using “Empty” UTF-8 Characters

#30

i don't see any variables in the url when i run the inspector.

Copy and paste the link into python, like so: >>> repr('http://kaspars.net/blog/web-development/invisible-click-tracking') "'http://kaspars.net/blog/web-development/invisible-click-tracking\\xe2\\x80\\x8b'" EDIT: Yeah, they didn't show up anywhere in the Safari inspector for me either.

I'm not seeing this in Python either. Perhaps the URL on the article has been "fixed"?
Post reply on HN