Live data from Hacker News

Show HN: MetricFlow – open-source metric framework

github.com

21–28 of 28 posts

Re: Show HN: MetricFlow – open-source metric framework

#21
post #20

It's a bit unclear how would I go with integrating MetricFlow with https://uxwizz.com for example that uses a MySQL database to store analytics data. From the docs, I don't really understand how it actually "understands" the underlying SQL database and how to retrieve the data I need. It feels like I have to write the query to get the data I want, but in a different syntax. Is there any point to use MetricFlow if you…

Three things:

First, MetricFlow does not currently support MySQL. We launched with support for BigQuery, Redshift, and Snowflake. I have opened an issue to add support for MySQL (and similar issues for other SQL engines are coming): https://github.com/transform-data/metricflow/issues/27

Second, what we call a data source is more similar to a table in a database, rather than the underlying database service itself. Metricflow itself is useful when you're using a single SQL engine - indeed, that's all we support today - but it is most useful when you're in a world where joins are a thing. That said, if you have one big data table you might still find it useful to have declarative metric definitions defined in Metricflow. Suppose, for example, you had a big NoSQL style table filled with JSON objects. You might define a few data sources that normalize those JSON objects into top level elements (identifiers, dimensions, aggregated measures) using the sql_query data source config attribute, and then that'd allow you to support structured queries on the data consumption end while pushing unstructured blobs from your application layer. This will be slow at query time, and only as reliable as the level of discipline exerted in your application development workflow, but it's possible.

Third, if we did support MySQL you'd basically connect to it via standard connection parameters - we have a config file where you can store the required information and then we'll manage the connections for you. However, I'm not familiar with uxwizz, and a quick perusal of their documentation did not turn up how one goes about connecting to the underlying DB. It's likely I just missed this, but at any rate I don't know how it is done. If they don't support standard MySQL client connections you'd need to write an adapter of some kind against whatever DB connection APIs they provide, in which case you'd likely need to roll a custom implementation of MetricFlow's SqlClient interface and initialize the MetricFlowEngine with that.

Re: Show HN: MetricFlow – open-source metric framework

#22
post #7
post #2

Cool. Like open source Looker. We adopted Looker at $previous_job. Then they got bought by Google, which was great for us as we were becoming a big GCP customer. I strongly encouraged google / looker team to at least open source their LookMl (looker modeling language - equivalent to MQL). They couldn’t figure it out. This type of metric definition is so empowering for businesses. Not enough engineers grok why this is…

Interested in why it's useful too. If the same SQL is used in multiple places and cause confusion, isn't it an organizational problem instead of a technical problem? For instance, what if two teams create two different and conflicting metric definitions to answer the same question? It's like turtles all the way down and how could we prevent diversion of query definitions even if we have a perfect metric definition sy…

I agree this is an organizational problem, but a lot of the tools we use every day - from bug trackers to distributed version control to collaborative document authoring tools - are technical solutions to organizational problems. One key organizational problem data people routinely deal with is exactly this issue of multiple metric definitions.

Imagine a company that wants to know how much time its users spend on its application, and wants that broken down by country in order to attract investors or something. They probably don’t have one metric for time spent on the application. They have like 12, all of them are different, and although all of them are wrong they each work sort of ok for a given context. To make matters worse, those 12 metrics are defined in different places, computed in different ways, and generated from different data sources. Maybe 3 of them are documented somewhere in the company wiki while the details of the rest live in the heads of 9 current or former employees who are off doing.... something.

That is a mess. Investors don’t care about all that nuance on use cases for time spent measures. Neither does anybody talking to the investors. What they don’t want to hear, when they ask for this simple metric, is “well..... what are you planning to use this for?” On the other hand, things get even worse if somebody grabs the wrong time spent metric and presents that publicly without realizing one of the other 11 was the one they really needed.

It would be better if our hypothetical company had its people define and document a set of consistent metrics and ensure they’re always computed the same way. By centralizing those definitions into a repository (ideally managed by source control) everybody can share the understanding of what a given metric represents. More importantly, they know where to look for pre-existing metric definitions. If this hypothetical company really needs 12 different time spent metrics (they don’t, but bear with me here) then each of those metrics can be named and defined in a way that explains what it’s good for, and they can simply point the makers of their investment pitch at the “baseline_user_time_spent” metric and have done with it.

For any of that to even be possible you need, at a minimum, a place to store metric definitions and a metric layer like MetricFlow that understands those metric definitions, translates them into the relevant queries, and executes them with the required filters and grouping attributes.

Re: Show HN: MetricFlow – open-source metric framework

#23
post #21
post #20

It's a bit unclear how would I go with integrating MetricFlow with https://uxwizz.com for example that uses a MySQL database to store analytics data. From the docs, I don't really understand how it actually "understands" the underlying SQL database and how to retrieve the data I need. It feels like I have to write the query to get the data I want, but in a different syntax. Is there any point to use MetricFlow if you…

Three things: First, MetricFlow does not currently support MySQL. We launched with support for BigQuery, Redshift, and Snowflake. I have opened an issue to add support for MySQL (and similar issues for other SQL engines are coming): https://github.com/transform-data/metricflow/issues/27 Second, what we call a data source is more similar to a table in a database, rather than the underlying database service itself. Met…

Thanks for the response!

With data sources, I mean if you have multiple services/products delivering data (e.g. an analytics platform, a CRM, all hosted on different servers).

> uxwizz, and a quick perusal of their documentation did not turn up how one goes about connecting to the underlying DB

UXWizz is self-hosted, so you just have a basic MySQL/MariaDB database that you have full control over (so you can connect remotely with the host/db/username/password you create). This is the database structure: https://docs.uxwizz.com/guides/database-querying

> you have one big data table you might still find it useful to have declarative metric definitions defined in Metricflow

Oh, so this is like saving queries? Why would I write the MetricFlow config instead of saving the SQL query directly? I had look over https://docs.transform.co/docs/metricflow/guides/introductio... but I found the concepts and config file a bit hard to understand (but maybe I'm not in the target audience).

Re: Show HN: MetricFlow – open-source metric framework

#24
post #22
post #7

Earlier quoted context omitted.

Interested in why it's useful too. If the same SQL is used in multiple places and cause confusion, isn't it an organizational problem instead of a technical problem? For instance, what if two teams create two different and conflicting metric definitions to answer the same question? It's like turtles all the way down and how could we prevent diversion of query definitions even if we have a perfect metric definition sy…

I agree this is an organizational problem, but a lot of the tools we use every day - from bug trackers to distributed version control to collaborative document authoring tools - are technical solutions to organizational problems. One key organizational problem data people routinely deal with is exactly this issue of multiple metric definitions. Imagine a company that wants to know how much time its users spend on its…

Totally agree with this ^

Re: Show HN: MetricFlow – open-source metric framework

#25

It looks like MetricFlow shines in constructing SQL queries on-demand, which means that it should be directly used by a BI tool, am I right with this?.. Generation of the static SQL (with CLI) for each report doesn't seem very usable on practice. In other words, BI tools needs to have a special connector that automatically utilizes MetricFlow Python API (or CLI). What BI tools already can use MetricFlow in this way (…

Looks like it supports GraphQL APIs[1], and downstream BI applications should be able to consume metric results from MetricFlow through GraphQL. [1] https://github.com/transform-data/metricflow#features

That's right! That's a potential option for an integration.

The other options are: - Transform's JDBC can be used to connect to tools that have SQL interfaces like Mode, Hex, Deepnote, etc.: https://docs.transform.co/docs/api/sql/sql-overview - Materializations can be exported as constructed data marts to tools like Tableau / Looker that take in constructed data sources: https://docs.transform.co/docs/metricflow/reference/material...

Re: Show HN: MetricFlow – open-source metric framework

#26
Thank you for open sourcing this. More competition in the budding metrics ecosystem is good for end users.

It seems like you think MetricFlow should be the data mart layer and not just the metrics layer. If that's true...why? Why would I join my fact and dimension tables in metricflow instead of in dbt? One of the value adds of dbt is that it centralizes business logic in a single place. Joins are business logic. The industry seems to be moving towards creating very wide data mart tables in dbt and surfacing them to the semantic layer 1:1, or building the metrics layer on top of them.

Re: Show HN: MetricFlow – open-source metric framework

#27
post #23
post #21

Earlier quoted context omitted.

Three things: First, MetricFlow does not currently support MySQL. We launched with support for BigQuery, Redshift, and Snowflake. I have opened an issue to add support for MySQL (and similar issues for other SQL engines are coming): https://github.com/transform-data/metricflow/issues/27 Second, what we call a data source is more similar to a table in a database, rather than the underlying database service itself. Met…

Thanks for the response! With data sources, I mean if you have multiple services/products delivering data (e.g. an analytics platform, a CRM, all hosted on different servers). > uxwizz, and a quick perusal of their documentation did not turn up how one goes about connecting to the underlying DB UXWizz is self-hosted, so you just have a basic MySQL/MariaDB database that you have full control over (so you can connect r…

Ah, I see!

> With data sources, I mean if you have multiple services/products delivering data (e.g. an analytics platform, a CRM, all hosted on different servers).

MetricFlow does not support this today. The model we are working with is of an analytics team relying on a centralized data warehouse service - hence the initial support for Redshift/Snowflake/BigQuery instead of Postgres/MySQL.

Deriving data from multiple input services is actually very complicated, because at some point you need to solve the cross-service join (or union) problem. This requires us to either merge everything into a single service layer (which isn't really appropriate for MetricFlow, there are entire service packages dedicated to just this problem) or keep track of input data lineages throughout the metric model.

My recommendation, if you need this, is to get an ETL service to transfer data from these different data service layers into a unified warehouse and then use that as your MetricFlow input source. This will be cleaner and easier to manage for you in the long run, even though it adds a bit of cost up front.

> UXWizz is self-hosted, so you just have a basic MySQL/MariaDB database

Oh, nice, that was the bit I was missing. In this case we will support it as soon as someone adds support for a MySQL client, but you'll likely have to bypass all of the UXWizz bindings and connect to the MySQL instance directly.

> Oh, so this is like saving queries? Why would I write the MetricFlow config instead of saving the SQL query directly?

Yes, it is, and you probably wouldn't want to do this in MetricFlow.

MetricFlow's data source config allows you to specify a SQL query inline:

https://docs.transform.co/docs/metricflow/guides/best-practi...

The data source itself essentially gives MetricFlow a base "table"-like construct that we can query on behalf of the user. So you define measures (which are basically aggregations), dimensions (attributes used for grouping and filtering), and identifiers (which you can use to link to other dimensions).

Ideally this would all be stored in a table in your data service already, and you can just provide us with a pointer to the table identifier and use the measure/dimension/identifier elements to provide what amounts to a very simple view over the underlying SQL table. You'd do this less for its own sake and more because that is the basis of the consistent metric computation and simplified dimension access MetricFlow provides.

The SQL query construct is in place to allow you to do some lightweight manipulation of the input tables in MetricFlow, because sometimes people need to do a little bit of filtering or transformation and either can not or prefer not to do this at a lower layer. When placed inside of MetricFlow proper these tend to be extremely simple.

The example I gave about splitting a massive JSON blob table is quite extreme, and I should have been more clear about this - I would not recommend using MetricFlow to do it. That's better handled by something end users of the metric system never have to see in any way at all, whether it be handled by traditional ETL processing on ingestion from the telemetry system into the warehouse, or by a data mart definition layer like dbt, or by something in between like a stack of Airflow operators. If you're in MySQL and your volumes are small you could even use views and then reference the view in MetricFlow (at least in theory, we've never tried anything like this).

Re: Show HN: MetricFlow – open-source metric framework

#28
post #26

Thank you for open sourcing this. More competition in the budding metrics ecosystem is good for end users. It seems like you think MetricFlow should be the data mart layer and not just the metrics layer. If that's true...why? Why would I join my fact and dimension tables in metricflow instead of in dbt? One of the value adds of dbt is that it centralizes business logic in a single place. Joins are business logic. The…

I'd say we think MetricFlow should be able to provide consistent, correct answers to reasonable queries end users of the metric model might ask. To do this across the various data warehouse layouts our users are likely to encounter we must necessarily provide support for dimensional joins. This doesn't mean MetricFlow should displace data mart services - to the contrary, I contend MetricFlow works best when layered on top of a warehouse built on centralized logic for managing its data layout. As an example, we generally push our customers to rely on the sql_table data source definition and push any sql_query constructs down to whatever warehouse management layers they have in place.

That said, you need to support joins, at least in some limited scope, in the semantic metric layer for it to be broadly useful. Consider this scenario - you have your dbt models producing wide tables for reasonable measure/dimension queries, and you have MetricFlow configs for the metric and dimension sets available in your data mart. Now imagine you've also got your finance team hooked up to a Google Sheets connector, and they're looking at revenue and unique customers by sales region. Cool, your wide table has that built in, no joins needed.

But what if they want something new? Let's say they want to know how they're doing against the target addressable market in each country. Should they have to submit a ticket to the data engineering team to add customer.country.market_size to your revenue table? Or should they be able to do "select revenue by customer__country__market_size" and get the report they need?

Our position is that we want to facilitate the latter - people getting what they need and knowing, as long as it's been defined properly in the model, that it's going to produce reasonable results. If your particular organization wants all of those joins run through a data mart ticket queue and surfaced as fully denormalized underlying tables that's fine by us, but most likely that's not what you want. You'd rather have some visibility into the types of joins people are requesting and then build out your data mart to more efficiently serve the requests people have on the ground, while still allowing them to ask new questions of the data without a long development feedback loop.

Post reply on HN