Viewing profile — rxin
rxin
HN member- Joined
- Wed, Mar 11, 2009, 10:26 PM UTC
- HN karma
- 2,259
- Public activity
- 274 items
- HN profile
- View on Hacker News ↗
About rxin
Recent public activity
- comment
- story
-
comment
Comment #43044595
This is coming to Serverless SQL warehouses soon!
- story
- story
-
comment
Comment #35542386
(OP / Databricks cofounder here) I'm not sure what happened but I just sent an email out internally to ask people not to do this. The team might have gotten overly excited by this …
-
comment
Comment #35542305
(Databricks cofounder here) All datasets are biased, including this specific one. However, we believe it's still very valuable to open source, for a few reasons: - This dataset is …
- story
-
comment
Comment #34464399
Databricks founder here. Doesn’t seem like DM work on HN anymore. Do you mind shooting me an email? Would love to follow up and understand the issues. rxin@databricks.com
-
comment
Comment #30596822
You should try Databricks, especially the new Photon engine powering Spark. In general more performant than Snowflake in SQL and a lot more flexible. (There are some cases in which…
-
comment
Comment #29232796
Isn’t that what the official TPC does?
- story
-
comment
Comment #29208293
Totally. Simplicity is critical. That’s why we built Databricks SQL not based on Spark. As a matter of fact, we took the extreme approach of not allowing customers (or ourselves) t…
-
comment
Comment #29208003
Geometric mean is commonly used in benchmarks when the workloads consists of queries that have large (often orders of magnitude) differences in runtime. Consider 4 queries. Two run…
-
comment
Comment #29207991
Ah ok. Wasn't clear. I think some repro scripts will be available soon.
- comment
-
comment
Comment #29207965
Check my reply, Leo.
-
comment
Comment #29207946
Exactly. Not sure about Netflix special, but there are experts that have dedicated their professional careers to creating fair benchmarks. Snowflake should just participate in the …
-
comment
Comment #29207940
There's an official TPC process to audit and review the benchmark process. This debate can be easiest settled by everybody participating in the official benchmark, like we (Databri…
-
comment
Comment #28535395
Co-author of the paper here. I don't think your argument holds here at all. It's a common misconception to think high performance would require tight coupling of storage and query …
-
comment
Comment #27318227
We have taken a very different approach with the Unity Catalog. It is designed opposite to the "cluster-centric" access control model, and will be user and role centric. Disclosure…
- story
-
comment
Comment #22600131
You should take a look at Databricks, in all three dimensions.
- story
-
comment
Comment #22101378
Disclaimer: I designed the system that won the 2014 GraySort based on Apache Spark, and the same system was extended by a different team to win the 2016 record in cloud sort. "Merg…