The mean means nothing: data visualization to debug a latency problem
1–10 of 24 posts
Re: The mean means nothing: data visualization to debug a latency problem
#2Re: The mean means nothing: data visualization to debug a latency problem
#3It's complicated, but really worth learning how prometheus and grafana heatmaps combine if you want dashboards for real time service data. Multimodal distributions are basically the expected outcome given all the caching done in distributed systems.
Re: The mean means nothing: data visualization to debug a latency problem
#4https://en.wikipedia.org/wiki/Anscombe%27s_quartet
Four sets of X,Y datapoints that have exactly (or very close) common statistical parameters (mean, variance, correlation, linear regression, R^2), but with vastly different spatial distributions and “behavior” when looked at visually.
Re: The mean means nothing: data visualization to debug a latency problem
#5Re: The mean means nothing: data visualization to debug a latency problem
#6With only having read the headline (as it is customary), I dare say: Aggregation without distribution always means nothing.
If I am to record something like queue times (as in latency), I'd have multiple samplings of both time enqueued and entries (or approx entries size), provided the queues have a reasonable size() method (e.g. not O(n) and not blocking). Even having max entries increased with lower overall latency, it'd speak of bursty behaviour.
In that regard - Recording percentiles en masse (production) is rather expensive when there are thousands of metrics. The avg/mean + max require few bytes (sub cacheline) and it can be implemented lock-free.
Re: The mean means nothing: data visualization to debug a latency problem
#7An interesting statistical example if you haven’t seen it before is the Anscombe Quartet : https://en.wikipedia.org/wiki/Anscombe%27s_quartet Four sets of X,Y datapoints that have exactly (or very close) common statistical parameters (mean, variance, correlation, linear regression, R^2), but with vastly different spatial distributions and “behavior” when looked at visually.
Statistics on the other hand is more accepting of discontinuous data due to it's basis in measure spaces.
Hence In general, you should know what your data should look like before you aim to detect anamolies. But also, it might be a good practice with new automated research actors to always use both approximations (measure theory based and otherwise), to figure out what's going on.
Re: The mean means nothing: data visualization to debug a latency problem
#8See this piece from Jeff Dean [0]
Quote:
""" Component-Level Variability Amplified By Scale
A common technique for reducing latency in large-scale online services is to parallelize sub-operations across many different machines, where each sub-operation is co-located with its portion of a large dataset. Parallelization happens by fanning out a request from a root to a large number of leaf servers and merging responses via a request-distribution tree. These sub-operations must all complete within a strict deadline for the service to feel responsive.
Variability in the latency distribution of individual components is magnified at the service level; for example, consider a system where each server typically responds in 10ms but with a 99th-percentile latency of one second. If a user request is handled on just one such server, one user request in 100 will be slow (one second). The figure here outlines how service-level latency in this hypothetical scenario is affected by very modest fractions of latency outliers. If a user request must collect responses from 100 such servers in parallel, then 63% of user requests will take more than one second (marked “x” in the figure). Even for services with only one in 10,000 requests experiencing more than one-second latencies at the single-server level, a service with 2,000 such servers will see almost one in five user requests taking more than one second (marked “o” in the figure).
"""
Re: The mean means nothing: data visualization to debug a latency problem
#9An interesting statistical example if you haven’t seen it before is the Anscombe Quartet : https://en.wikipedia.org/wiki/Anscombe%27s_quartet Four sets of X,Y datapoints that have exactly (or very close) common statistical parameters (mean, variance, correlation, linear regression, R^2), but with vastly different spatial distributions and “behavior” when looked at visually.
Re: The mean means nothing: data visualization to debug a latency problem
#10Think they made a typo - "before" curve is higher than the "after" curve to the right of 140ms, which means more requests take longer than before to finish.
Or am I the one misunderstanding?