Live data from Hacker News

The Rise and Fall of the OLAP Cube

holistics.io

51–54 of 54 posts

Re: The Rise and Fall of the OLAP Cube

#51
If you're a hacker interested in SQL and OLAP, you might enjoy Greenspun's (he of of Greenspun's tenth rule fame) writings on these subjects:

https://philip.greenspun.com/wtr/data-warehousing.html https://philip.greenspun.com/sql/

Although technically obsolete (as in talking about 90s database systems that have bitten the dust since then), that's a minor defect. He spends most effort on teaching timeless principles.

Re: The Rise and Fall of the OLAP Cube

#52
post #36

For anyone wondering, OLAP != OLAP Cubes OLAP = A category of databases meant for analyzing data. These are eventually consistent db's, and not OLTP db's. OLAP db's include Redshift, Teradata, Snowflake, BigQuery, and others. Generally what makes a database an MPP database is partitioning compute and storage. Generally what differentiates one MPP db from another is whether or not data and compute are colocated. OLAP…

Redshift is not eventually consistent, it's fully ACID compliant.

You are correct, and eventually consistent is a hand-wavy description of Redshift. While Redshift is ACID compliant, it squarely belongs in the category of OLAP db's and not OLTP db.s

Re: The Rise and Fall of the OLAP Cube

#53
Cost is a big factor the author underestimated in this big data era. Precalculated cube is not only faster but also times cheaper in the cloud, thanks to the reuse of precalculated result.

Dynamic query services in the cloud basically charge by processed data volume, like Google BigQuery and Amazon Redshift/Athena. For small and medium dataset, this works well. But for big data close to or above billions of rows, the cost will make you reconsider.

In the recent Apache Kylin Meetup in Berlin, OLX Group shared their comparison between OLAP cube and dynamic query in real case. Given 0.1 billion rows, cube technology (Apache Kylin and SSAS) prevails over MPP+Columnar (Redshift) easily. Especially Apache Kylin is 3.8x faster and 4.4x cheaper than Redshift for their business. (https://www.slideshare.net/TylerWishnoff/apache-kylin-meetup...)

For me, a mix of precalculation (80%) and dynamic calculation (20%) should hit the sweet point between cost effectiveness and query flexibility.

Re: The Rise and Fall of the OLAP Cube

#54
I have one question for you guys. If my company is focused on Spark and Vertica, and I want to learn data modelling on top of those, does Kimball still make sense? The article says yes in general but I'd like to know your opinions.

Currently the BI team doesn't do much dimensional modelling as far as I see. Every thing is taken from Kafka and dumped into some wide tables with all columns that we the analysts need. Actually there is no data modelling at all.

Post reply on HN