Live data from Hacker News

Databricks response to Snowflake's accusation of lacking integrity

databricks.com

51–60 of 168 posts

Re: Databricks response to Snowflake's accusation of lacking integrity

#51

so, alternatives? Aside from the Azure/GCP/AWS internal offeringa I know about Snowflake and Firebolt, Databricks is new to me.

https://en.m.wikipedia.org/wiki/Databricks

"Databricks is an enterprise software company founded by the creators of Apache Spark. [...] Databricks develops a web-based platform for working with Spark, that provides automated cluster management and IPython-style notebooks."

Re: Databricks response to Snowflake's accusation of lacking integrity

#52
post #19

As much as I love seeing competition in the space and am enjoying my popcorn, I really don't understand what Databricks is doing here: this feels like a childish foodfight rather than an obsession with the customer...

:) That is a good question. Why spend eng cycles to submit results to the TPC council - why not just focus on customers?

I believe the co-founders have addressed this in the blog.

> Our goal was to dispel the myth that Data Lakehouse cannot have best-in-class price and performance. Rather than making our own benchmarks, we sought the truth and participated in the official TPC benchmark.

I'm sure anybody seriously looking at evaluating data platforms would want to look at things holistically. There are different dimensions like open ecosystem, support for machine learning, performance etc. And different teams evaluating these platforms would stack rank them in different orders.

These blogs, I believe, show that Databricks is a viable choice for customers when performance is a top priority (along with other dimensions). That IMO is customer obsession.

Re: Databricks response to Snowflake's accusation of lacking integrity

#53
Ive been following this and it’s kind of embarrassing to watch.

I love working with Databricks and Snowflake. They both knock it out of the park for their respective use case. They’re amazing products.

It makes no sense to fall out about this though.

For a 100TB dataset with a funky calculation, Spark will trounce Snowflake. For a 1 row dataset, Snowflake will return before the spark job has been serialised.

Re: Databricks response to Snowflake's accusation of lacking integrity

#54
post #45

Earlier quoted context omitted.

Ahh, fair enough there. That said, snowflake would help their case if they would actually do at least one actual tpc.org third party result

My guess is that they know the result won't look terrific. And they also know Snowflake works well in production for people despite that. So, little upside.

For sure.

All leaders in a space take this approach. Little be gained, a fair bit to lose if you are ALREADY leading without having to debate / do a benchmark etc.

Anyways, the benchmark is only one part of the overall story for these solutions.

Re: Databricks response to Snowflake's accusation of lacking integrity

#55

The irony here is that what Databricks is doing to Snowflake is exactly what Snowflake did to AWS and Redshift. Same playbook - show that you’re better in a key metric that’s easy to understand (performance) to get the attention, but then pitch the paradigm change. In Snowflake’s case, that was separation of storage and compute. In Databrick’s case, it’s the Lakehouse Architecture. I think the reason why Snowflake is…

> I think the reason why Snowflake is so nervous because they know they can’t win this game.

Isn't Databricks' delta.io, which their Data Lakehouse product builds on top of, open source? Snowflake could take the best parts from and run with it?

Re: Databricks response to Snowflake's accusation of lacking integrity

#57

The irony here is that what Databricks is doing to Snowflake is exactly what Snowflake did to AWS and Redshift. Same playbook - show that you’re better in a key metric that’s easy to understand (performance) to get the attention, but then pitch the paradigm change. In Snowflake’s case, that was separation of storage and compute. In Databrick’s case, it’s the Lakehouse Architecture. I think the reason why Snowflake is…

In what way is lakehouse architecture beneficial over something like Snowflake or BigQuery?

I understand the appeal over having lake and warehouse as separate components, but with those native cloud warehouses, you can already do everything a lake does.

Re: Databricks response to Snowflake's accusation of lacking integrity

#58

Serious question: Databricks, Snowflake, Dremio. All these "Data" platform companies => which one do you have for your Data Lake and Data Warehouse solution? I'm sick and tired of these companies Snake Oiling the Data industry by offering "the easiest" platform to satisfy your Data Lake + Warehouse solution only to fall hard whenever you hook it up with your production data (big dataset). PS: Anyone selling Data Lake…

Please read up on Lakehouse. Data Lake + Merge support + DW performance is now possible. That is the game changer.

Do you work for Databricks?

Re: Databricks response to Snowflake's accusation of lacking integrity

#59
post #36

Earlier quoted context omitted.

Databricks is way more than hadoop or spark. A great analogy - Spark is a great engine but you need to design and build all of the other subsystems. Databricks is an F1 car - everything is built out. You get in and drive - FAST.

Databricks is a shit platform that encourages terrible data practices and accretion of technical debt.

Seems pretty good to us! Can you give more information?

Re: Databricks response to Snowflake's accusation of lacking integrity

#60

The irony here is that what Databricks is doing to Snowflake is exactly what Snowflake did to AWS and Redshift. Same playbook - show that you’re better in a key metric that’s easy to understand (performance) to get the attention, but then pitch the paradigm change. In Snowflake’s case, that was separation of storage and compute. In Databrick’s case, it’s the Lakehouse Architecture. I think the reason why Snowflake is…

To be fair Apache Spark, which started long before either company existed, was built on the assumption that compute and storage should be separate. Unlike Hadoop, Spark did not come with any storage system and could read from any source.
Post reply on HN