Live data from Hacker News

Microsoft Releases Open Source ML Library for Distributed Search Engine Creation

aka.ms

11–12 of 12 posts

Re: Microsoft Releases Open Source ML Library for Distributed Search Engine Creation

#11

Earlier quoted context omitted.

On the point about autograd tools, it’s unfortunately not always helpful for types of models that are not amenable to that type of framework, like gradient-free methods, custom Bayesian inference models, or modified versions of some traditional models (like bias-corrected logistic regression). Here again, if you tie machine learning to a big system like Spark, which is typically a huge IT cost in a lot of companies,…

I definitely would agree that you should pick the best tool for the job and not limit yourself to one ecosystem if it's too difficult. One way to make Spark a bit easier to work with is through kubernetes or a tool like databricks that provides it as a service. Kubernetes, in particular, provides you a really nice amount of flexibility and composability when designing systems. One thing that we created to try to fill…

This is actually the part I most disagree with. The overhead of connectors to Spark, particularly any use of py4j, is far too limiting except in cases when the data workload is so large that it effectively amortizes the overhead. For small scale prototypes, it’s a disaster, and then separate there are concerns for data type marshaling through the JVM when you may have a Python-only data model.

At the time of evaluation for me, I also found Databricks had extremely limited support for runtime environments defined by arbitrary containers. You have to select cluster nodes according to their prescribed images and choices.

Say you need Tensorflow or CUDA compiled with a weird set of optimization flags, or you need other special provisions in the runtime environment. In fact, variations of the runtime environment may even be part of some reproducible experiments, so you need to execute across a variety of parameters that govern how the runtime is built.

Anything that can’t support this type of stuff out of the box is just not worth it. Anybody can hook a notebook environment up to analyze data from some data warehouse or distributed file system.

The hard part is always how to make that setup configurable and parameterizable across the needs of different projects, especially arbitrary runtime environments.

Re: Microsoft Releases Open Source ML Library for Distributed Search Engine Creation

#12

Earlier quoted context omitted.

Having evaluated Spark extensively for my company’s ML use cases, I came away deeply disappointed. One thing that bugs me in particular is that it essentially presumes all workflows involve huge, distributed datasets. But most model development work, especially for projects that will eventually be trained for production using huge, distributed data sets, must begin their life cycle as small data prototypes with extre…

Spark is intended to work on big datasets. It's machine learning capability is very limited and it's primary strength is processing huge amounts of data. I think it's unfair to blame it for 'failing' on small datasets

I agree that Spark may be a fine choice for ETL and generic pipeline tasks.

But lots of companies will choose it as a data warehouse computation layer and then enforce a policy to standardize everything around it, including tasks like machine learning that are poorly suited for Spark.

Worse, companies like Databricks will encourage this standardization and act like yes-man consultants, promising Spark ML offerings can solve all the problems, and you quickly end up with some brittle monster of a data warehouse system that is oriented to be convenient for Spark (which can’t effectively be used to solve the problems) and everything is deeply inconvenient to pipe to non-Spark systems, and nobody is sympathetic to any budgetary needs for other systems, since they spent it all on Spark.

Post reply on HN