Live data from Hacker News

Scaling Machine Learning at Uber with Michelangelo

eng.uber.com

11–20 of 61 posts

Re: Scaling Machine Learning at Uber with Michelangelo

#11

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

would love to know what is your model development iteration. especially how you do testing, etc

See my comment here, but I can answer other questions if you have them

https://news.ycombinator.com/item?id=18376567

Re: Scaling Machine Learning at Uber with Michelangelo

#12
post #5

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

Can you elaborate a bit more about your usage of GitLab CI/CD for model management/development. I am currently working on a platform [1] that tries to solve some of the issues mentioned in the article, i.e. improving data scientists' productivity and velocity, compare models, solve reproducibility issues... [1] https://github.com/polyaxon/polyaxon

Polyaxon looks nice but we don’t admin the majority of the GPU resources we use (which is why being able to tell GitLab-runner to invoke Slurm is cool)

Pachyderm is another one I’ve looked at but we don’t have the sys admin bandwidth for that stuff right now.

Re: Scaling Machine Learning at Uber with Michelangelo

#13
post #3

It's kinda funny they tout their usage of GPS. I use Uber on a near daily basis and drivers by an large use Google maps. They have out right said "Uber sucks for directions" And if you use express pools it will always say to go the wrong side of an intersection. I like uber because of the drivers, but their fancy technology is flawed.

I believe they’re using GPS data here more for analytics, rather than navigation.

They can use GPS data to chart usage metrics, plan pool rides, check for anomalies, and harass journalists, for example.

Re: Scaling Machine Learning at Uber with Michelangelo

#14

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

Why would you do this instead of using pymc3?

Re: Scaling Machine Learning at Uber with Michelangelo

#15
post #3

It's kinda funny they tout their usage of GPS. I use Uber on a near daily basis and drivers by an large use Google maps. They have out right said "Uber sucks for directions" And if you use express pools it will always say to go the wrong side of an intersection. I like uber because of the drivers, but their fancy technology is flawed.

Please do not conflate GPS with navigation. There is a massive set of problems you can solve with high fidelity GPS Data (Uber knows it is a driver in a car, verifies it with another GPS entity (rider app reports GPS also), etc). There is not that much overlap between great GPS data and great maps - no amount of great GPS data will give you a good basemap. Please let me know if I am not making sense, I am more than happy to provide examples / explain further!

Re: Scaling Machine Learning at Uber with Michelangelo

#16

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

Why would you do this instead of using pymc3?

PyMC3 didn’t run well on GPUs last I tried. That may have changed but I find PyTorch easier to work with than Theano or TensorFlow.

Re: Scaling Machine Learning at Uber with Michelangelo

#17
post #5

Earlier quoted context omitted.

Can you elaborate a bit more about your usage of GitLab CI/CD for model management/development. I am currently working on a platform [1] that tries to solve some of the issues mentioned in the article, i.e. improving data scientists' productivity and velocity, compare models, solve reproducibility issues... [1] https://github.com/polyaxon/polyaxon

We uh treat models as code, but also have NFS shares setup for the storage and GitLab runner talking to a Slurm cluster to run the models. Results and cross validation upload to GitLab. Main thing we haven’t built out yet are performance dashboards for showing improvement across commits, but with the GitLab APIs that’s a script away (currently we do it by hand)

Thanks for your reply.

Actually the question was more around "how do you create your models and what do you mean treating them as code", "why slurm and not something like airflow" , "what is the test/performance setup - backtesting, smoke test" etc etc

The Gitlab stuff is easier to understand.

Re: Scaling Machine Learning at Uber with Michelangelo

#18
post #5

Earlier quoted context omitted.

Can you elaborate a bit more about your usage of GitLab CI/CD for model management/development. I am currently working on a platform [1] that tries to solve some of the issues mentioned in the article, i.e. improving data scientists' productivity and velocity, compare models, solve reproducibility issues... [1] https://github.com/polyaxon/polyaxon

We uh treat models as code, but also have NFS shares setup for the storage and GitLab runner talking to a Slurm cluster to run the models. Results and cross validation upload to GitLab. Main thing we haven’t built out yet are performance dashboards for showing improvement across commits, but with the GitLab APIs that’s a script away (currently we do it by hand)

What do you think are the best dashboard options for showing improvements?

Re: Scaling Machine Learning at Uber with Michelangelo

#19

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

What does Uber currently do to people that you hate? Uber currently provides people with more than 2 Million jobs [1]. Uber drivers/couriers made almost $13 Billion in the US alone last year [2].

[1] https://medium.com/@gc/ubers-path-forward-b59ec9bd4ef6 [2] https://www.sfchronicle.com/business/article/Uber-drivers-in...

Re: Scaling Machine Learning at Uber with Michelangelo

#20

I love what Uber does with machines (ML), hate what it (currently) does to people. We recently potted some models from Stan to Pyro (SVI on PyTorch), and it’s been reallly exciting (except for the dark corner of poutines), it really has the performance of something being used in production, except the occasional nan explosion. edit we are lazy and use our GitLab CI/CD to drive model development iteration. It’s not as…

I like what Uber (currently) does to people. Gets passengers from point A to point B efficiently while saving them significant money in the process over alternatives. Metaphorically puts dinner on the table of hundreds of thousands of drivers. Literally puts dinner on the table of millions (UberEats). Has a business model that doesn't rely exposing more eyeballs to more ads, corrupting the press, media, and privacy in the process. Reduces car ownership and dependence. Moving towards encouraging people to ride green vehicles. Literally saves lives (reducing DUI). Yeah, I'm okay with the Uber of 2018.*

*Disclaimer: I work at Uber, and my opinions are solely my own. We're hiring.

Post reply on HN