Off-topic: Because of situations like these, I'm surprised that part of the checklist when launching a PR blog is not: "Block googlearchive/archive.org robots" There have been very, very few times when a company's webpage was down and I needed to go to google-archive or archive.org to refer to some innocuous information. However, the times that I've used those sites to gather evidence of possible whitewashing? Many,…
Rides of Glory – Uber Blog (2012)
21–30 of 116 posts
Re: Rides of Glory – Uber Blog (2012)
#22Re: Rides of Glory – Uber Blog (2012)
#23In another deleted post [0] the author talks about using a name-to-gender API to look at ride locations by gender, which implies that these analyses were not done using anonymized data. [0] https://web.archive.org/web/20140827195715/http://blog.uber....
Re: Rides of Glory – Uber Blog (2012)
#24Stuff like this is one of the many reasons I love archive.org. I think i's really important to capture historical artifacts for future analysis. The service they provide doesn't allow the "Ministry of Truth"[1] to doctor historical documents to meet their present day narrative. [1] https://en.wikipedia.org/wiki/Ministry_of_Truth
Sadly, it does. archive.org respect the robots.txt of the current website owner. This can mean that they have the data but choose not to give you access to them. I have seen cases in the past where a website I once frequented became defunct, then the domain expired, then someone parked a holding page on that domain including a robots.txt that keeps archive.org from displaying the old data (which do not even belong to…
Re: Rides of Glory – Uber Blog (2012)
#25Stuff like this is one of the many reasons I love archive.org. I think i's really important to capture historical artifacts for future analysis. The service they provide doesn't allow the "Ministry of Truth"[1] to doctor historical documents to meet their present day narrative. [1] https://en.wikipedia.org/wiki/Ministry_of_Truth
And yes I don't care what you think, but a company with a billion(ish) of funding is more powerful than YOU.
Re: Rides of Glory – Uber Blog (2012)
#26Anyone (especially the HN crowd) should know they have the data, and if you think they're not carefully analyzing it behind the scenes (like every other tech company who has your data), I've got things to sell you. I personally think a tiny peek like this into the data, much like the usage posts that OKCupid, YouPorn, and others give, is neat.
Re: Rides of Glory – Uber Blog (2012)
#27Stuff like this is one of the many reasons I love archive.org. I think i's really important to capture historical artifacts for future analysis. The service they provide doesn't allow the "Ministry of Truth"[1] to doctor historical documents to meet their present day narrative. [1] https://en.wikipedia.org/wiki/Ministry_of_Truth
Sadly, it does. archive.org respect the robots.txt of the current website owner. This can mean that they have the data but choose not to give you access to them. I have seen cases in the past where a website I once frequented became defunct, then the domain expired, then someone parked a holding page on that domain including a robots.txt that keeps archive.org from displaying the old data (which do not even belong to…
Re: Rides of Glory – Uber Blog (2012)
#28Is it sad that in the realm of Uber blunders I find this relatively tame?
One could do an analysis like this while still working with anonymized data. Still a bit creepy, but not that different from reports and blog posts you see from other startups and tech companies.
Re: Rides of Glory – Uber Blog (2012)
#29Uber employees are the kind of people who kept a telescope in their bedroom window to peep on girls down the street. I'll never understand why they're still in business.
Re: Rides of Glory – Uber Blog (2012)
#30In another deleted post [0] the author talks about using a name-to-gender API to look at ride locations by gender, which implies that these analyses were not done using anonymized data. [0] https://web.archive.org/web/20140827195715/http://blog.uber....
Internal metrics teams nearly always have access to complete data. The issue is sharing non-anonymized data externally.