Live data from Hacker News

Data Wrangling at Slack

slack.engineering

21–30 of 76 posts

Re: Data Wrangling at Slack

#21
With Qubole you can offload data engineering to their platform. Cluster management is super simple. Hand rolled solutions in my experience are a pain and elastic cloud features take up time to build. Qubole's offering provides out of the box experience for most big data engines out there. Presto/ Spark/ Hive/ Pig - what have you - all work with your data living in S3 (or any other object storage). I believe they have offerings in other clouds too.

Some amount of S3 listing optimisation is done by Qubole's engineering team for: https://www.qubole.com/blog/product/optimizing-s3-bulk-listi...

They also have features that allow you to auto-provision for additional capacity in your compute clusters as your query processing times increase.

Re: Data Wrangling at Slack

#22
This is off-topic, but I can't help myself:

Slack, Hive, Presto, Spark, Sqooper, Kafka, Secor, Thrift, Parquet.

I sometimes can't tell the difference between real Silicon Valley product names and parodies. I'm starting to miss the days when it was all just letters and numbers.

Re: Data Wrangling at Slack

#23

We're actually having a debate now as we're starting to process larger datasets as to whether or not we should keep everything on S3 or start using HDFS w/ Hive. I'm curious if you guys considered HDFS and why you decided to go strictly with S3, and additionally, are there any issues you encounter with S3.

I would recommend S3.

Using S3 with EMR in production was breeze for us. Even cost effective, since you can play with spot instances depending on your jobs. You also improve utilization of your resources.

With recent Athena it is possible also to do ad hoc queries directly :) Before it required starting "QA" cluster.

Re: Data Wrangling at Slack

#24
post #21

With Qubole you can offload data engineering to their platform. Cluster management is super simple. Hand rolled solutions in my experience are a pain and elastic cloud features take up time to build. Qubole's offering provides out of the box experience for most big data engines out there. Presto/ Spark/ Hive/ Pig - what have you - all work with your data living in S3 (or any other object storage). I believe they have…

When Amazon Athena actually matures, wouldn't it solve at least the interactive query needs, probably at a much lower/elastic price point than Qubole?

Re: Data Wrangling at Slack

#25
Seems like a pretty typical set of problems. Dependency conflicts hard. Schema evolution hard. Upgrades hard.

The big data space still feels like an overengineered, fractured, buggy mess to me. I was hoping spark would simplify the user experience but it's as much of a clusterf*ck as anything else.

How hard can fast, reliable distributed computation and storage for petabytes of data be? He said ironically.

Re: Data Wrangling at Slack

#27
post #24
post #21

With Qubole you can offload data engineering to their platform. Cluster management is super simple. Hand rolled solutions in my experience are a pain and elastic cloud features take up time to build. Qubole's offering provides out of the box experience for most big data engines out there. Presto/ Spark/ Hive/ Pig - what have you - all work with your data living in S3 (or any other object storage). I believe they have…

When Amazon Athena actually matures, wouldn't it solve at least the interactive query needs, probably at a much lower/elastic price point than Qubole?

True, I've tried Athena and it's great at cost, performance and ease of use. However, most Data Engineering teams have lots of custom tweaks they need and certain level of control to add jars, applications, UDFs to their queries. I don't see this available through Athena today.

Re: Data Wrangling at Slack

#28
post #25

Seems like a pretty typical set of problems. Dependency conflicts hard. Schema evolution hard. Upgrades hard. The big data space still feels like an overengineered, fractured, buggy mess to me. I was hoping spark would simplify the user experience but it's as much of a clusterf*ck as anything else. How hard can fast, reliable distributed computation and storage for petabytes of data be? He said ironically.

IMO one major problem is integration between different projects. Like you said, its a hard problem, and any solution typically depends on many many different open source projects because of the scope of challenges. All of those projects go forward without much coordination between the teams because they're open source. Then we end up in this fun, fun clusterfuck.

Re: Data Wrangling at Slack

#29
post #22

This is off-topic, but I can't help myself: Slack, Hive, Presto, Spark, Sqooper, Kafka, Secor, Thrift, Parquet. I sometimes can't tell the difference between real Silicon Valley product names and parodies. I'm starting to miss the days when it was all just letters and numbers.

I give up. Which one's the real one?

Re: Data Wrangling at Slack

#30
post #22

This is off-topic, but I can't help myself: Slack, Hive, Presto, Spark, Sqooper, Kafka, Secor, Thrift, Parquet. I sometimes can't tell the difference between real Silicon Valley product names and parodies. I'm starting to miss the days when it was all just letters and numbers.

Yeah can't we go back to naming companies with a color and an animal like we did in the glory days?
Post reply on HN