Live data from Hacker News

Rides of Glory – Uber Blog (2012)

web.archive.org

71–80 of 116 posts

Re: Rides of Glory – Uber Blog (2012)

#71

To me, it's not even the use of data per se that is most creepy about this post. Really, the tone of the essay seems to revel in "having 'fun' with user data," as if a sophomore at a university wrote it. I mean, I found the idea behind the post interesting: of course you can analyze trends in ridership to draw interesting conclusions. At the end of the day, however, it's a horrible idea to say "Hey, we know which of…

Since I know author (he had the desk next to me in the lab) I feel I should add something here. Let me just say that he was someone who had recently finished his PhD and was taking a summer data science sabbatical with Uber. I know that they were also working on how to optimize the distribution of cars to best serve a market but that would be pretty boring to talk about. I think what you are seeing in "having fun" with the data is actually more of the neuroscientist/psychologist coming out in him combined with his propensity for making science topics interesting and maybe a little sensational. This is the guy who also brought us zombie neuroscience remember (check it out if you aren't familiar). At best we are daily trying to infer subjects internal processing/motivations from the sorts of actual behavioral measures we collect in the lab. This is one where behavioral inferences are perhaps better than self-report. As to your other point, I can even see how knowing how many people are being frisky and where they live being used to improve public health policy although that was obviously not the point (and it being kept anon of course). But let's not forget that a cohort of uber riders is not a very uniform sample of the population and you are going to be somewhat limited in your ability to use it as a tool for social justice. He did a similar analysis although more in line with our lab work where he combined a big brain training app's anon user data with their state geography and its demographics to draw some interesting conclusions that you wouldn't have expected. These are just the application of techniques that we use in the lab to make inferences about brain activity and unravel its complexity and weave an understandable story. I am certainly not trying to defend Uber's other data privacy issues which are in the news and very concerning but I don't think this is one of them. I hope this helps a little knowing more of the backstory.

Re: Rides of Glory – Uber Blog (2012)

#72

I gotta say, I'm not really seeing the creepy / cringey / evil / whatever-else here... Anyone (especially the HN crowd) should know they have the data, and if you think they're not carefully analyzing it behind the scenes (like every other tech company who has your data), I've got things to sell you. I personally think a tiny peek like this into the data, much like the usage posts that OKCupid, YouPorn, and others gi…

A limousine service that uses business records to work out when passengers are fucking and then writes articles about doing this, complete with fucking graphs and even some fucking maps, I think qualifies as pretty fucking creepy.

Re: Rides of Glory – Uber Blog (2012)

#73

I gotta say, I'm not really seeing the creepy / cringey / evil / whatever-else here... Anyone (especially the HN crowd) should know they have the data, and if you think they're not carefully analyzing it behind the scenes (like every other tech company who has your data), I've got things to sell you. I personally think a tiny peek like this into the data, much like the usage posts that OKCupid, YouPorn, and others gi…

The problem here (for me personally, at least) is that Uber is not in the business of selling dates/"encounters" and that people don't expect a ridesharing company to go right for the sexual data. Even OKCupid is straddling the line here with http://blog.okcupid.com/index.php/we-experiment-on-human-bei... noting that: To test this, we took pairs of bad matches (actual 30% match) and told them they were exceptionally…

The idea that statistics have a moral imperative to be "decent" is as fascinating as it is ridiculous. Anonymized data is not a privacy breach, and Uber probably doesn't have any data that can help with "hunger, poverty, illiteracy, etc.".

I'm sorry if the idea that "people's short overnight stays are evident in their travel data" makes you blush, but that isn't anyone else's problem.

Re: Rides of Glory – Uber Blog (2012)

#74
One more thing--

Would google publish data that shows how searches for porn spike during different times of the day or times of the year, as if it's some "cool and hip and edgy!" insight?

I don't think so.

And for the same reason they don't (whatever reason that is), it would probably also be wise for Uber not to post stuff like this.

I really don't care, nor am I offended. I'm just speculating that Uber doesn't have the brightest team of execs and still have a lot of "growing up" to do.

Re: Rides of Glory – Uber Blog (2012)

#75

Earlier quoted context omitted.

The problem here (for me personally, at least) is that Uber is not in the business of selling dates/"encounters" and that people don't expect a ridesharing company to go right for the sexual data. Even OKCupid is straddling the line here with http://blog.okcupid.com/index.php/we-experiment-on-human-bei... noting that: To test this, we took pairs of bad matches (actual 30% match) and told them they were exceptionally…

A little off-topic, but I don't see why OKCupid's actions here are unethical. Their matching algorithm isn't perfect, so they shouldn't treat it as an oracle of truth. How else would they discover false negatives in their algorithm? Especially since, in this case, a false negative is worse than a false positive (not meeting someone you'll like vs having one unsuccessful date).

Don't they sell people on their super-accurate-awesomesauce-state-of-the-art matching algorithm? Were people warned that they may be guinea pigs?

Re: Rides of Glory – Uber Blog (2012)

#76
post #30

Earlier quoted context omitted.

You have to start with the original data, which is obviously de-anonymized. Full data -> [gender, time, origin neighborhood, destination neighborhood] leaves you with a pretty anonymous dataset, and is all that would be required for this analysis. Internal metrics teams nearly always have access to complete data. The issue is sharing non-anonymized data externally.

I agree it's possible that the name-to-gender mapping was done before the full ride data was handed over to this analyst. (Though just removing real names would still leave a lot to be desired in the anonymizing process). However there's no mention in these posts of such safeguards, and subjectively the post reads more like the analyst is just fishing around in the full raw dataset of ride times, start and end locati…

> Do you think that your average Uber rider would be OK with Uber employees analyzing their ride patterns (with their real names attached) to try to figure out where and when they are having sex?

Sure, as long as Uber isn't broadcasting that information with their name attached. The average person really doesn't care about (or understand the extent of) data analysis (from companies or the government) -- what they care about is public disclosure which may mean personal embarrassment or a lawsuit or other form of inconvenience. People who want to control all their data are hoping for a fantasy world where observations and inferences by third parties are magically made impossible. The reasonable thing to focus lawmaking efforts on is limiting legal forms of disclosure and standardizing safe storage requirements for the raw data -- indeed such laws already exist, with the HIPPA privacy rule perhaps being the best known in the US.

Re: Rides of Glory – Uber Blog (2012)

#77

To me, it's not even the use of data per se that is most creepy about this post. Really, the tone of the essay seems to revel in "having 'fun' with user data," as if a sophomore at a university wrote it. I mean, I found the idea behind the post interesting: of course you can analyze trends in ridership to draw interesting conclusions. At the end of the day, however, it's a horrible idea to say "Hey, we know which of…

The author is Bradley Voytek ( http://darb.ketyov.com ), who is now a professor of neuroscience and should probably know better.

He acknowledges the questionable nature of his blog posts on his website:

> Between June 2011 and August 2011 I worked with my friends over at Uber as their data scientist, writing (what I thought were) amusing, data-driven blog posts (among other, more serious roles).

Re: Rides of Glory – Uber Blog (2012)

#78

Earlier quoted context omitted.

The problem here (for me personally, at least) is that Uber is not in the business of selling dates/"encounters" and that people don't expect a ridesharing company to go right for the sexual data. Even OKCupid is straddling the line here with http://blog.okcupid.com/index.php/we-experiment-on-human-bei... noting that: To test this, we took pairs of bad matches (actual 30% match) and told them they were exceptionally…

A little off-topic, but I don't see why OKCupid's actions here are unethical. Their matching algorithm isn't perfect, so they shouldn't treat it as an oracle of truth. How else would they discover false negatives in their algorithm? Especially since, in this case, a false negative is worse than a false positive (not meeting someone you'll like vs having one unsuccessful date).

> How else would they discover false negatives in their algorithm?

This is exactly why research that deals with humans at Universities invariably must pass a human subjects review process. "How else would we discover X?" is certainly not reason to subject anyone to an unethical experiment. Subjecting people to what you likely believe to be a bad date should very definitely raise red flags, even if the details in practice would pass a human subjects review.

And that's the trouble: there's a tremendous space of research that just isn't ethical to carry out on actual living humans. As such, we have to find methods to determine answers to those questions that don't breach ethical standards. The burdens of discovery must lie squarely on the researchers, not on the (often unwitting) experimental subjects.

Re: Rides of Glory – Uber Blog (2012)

#79
post #58

Earlier quoted context omitted.

The Internet Archive should implement some sort of digital signature system to allow website owners with foresight to prevent this.

They could just timestamp different versions of robots.txt (which they probably do already), and respect it depending on date (which is more of a hassle, because you have to build it in your UI logic).

That would not solve the problem they're trying to solve.

Let's say I post something that I shouldn't have posted -- insider stock information, nude photos, whatever. Perhaps something illegal for me to post. I need to make it go away.

I need to be able to create a robots.txt today which affects stuff I posted yesterday.

This is why archive.org respects the current robots.txt for access to past content.

Re: Rides of Glory – Uber Blog (2012)

#80
post #36
post #10

Earlier quoted context omitted.

Is it cringey because of the data they present, or because of the conclusions they draw from that data?

I think probably the language and tone of the post more than either the data or the conclusions.

Interesting. In the more socially-aware fora I'm familiar with, what's called "tone policing" or "tone trolling" is seen as a major faux pas, an aggressive act designed to shut people out of the conversation.
Post reply on HN