Really respectable writing and perspective. Questdb blog posts that get posted here never disappoint
Lies, Damn Lies and Database Benchmarks
11–20 of 30 posts
Re: Lies, Damn Lies and Database Benchmarks
#12Re: Lies, Damn Lies and Database Benchmarks
#13The database wars of the late 1990s were full of this kind of stuff. Oracle, Sybase, IBM etc invested heavily in tuning specifically for benchmarks like TPC-C just so they could post ads in the Wall St Journal saying theirs was faster. I do sympathize with OP, though, their objection to measuring cold-start queries is incomplete without also describing how often cold start needs to happen. If you restart once every f…
My personal opinion is that you need a massive amount of data and massive number of different variables to test for separately. For example you might want to monitor how many cache misses/hits there were, p99 latency etc. And you want to do it under full load, expected load etc. And you want to compare the different versions of the same database because comparing different databases makes things combinatorially more difficult, unless you have a real production use case that you are optimizing for ofc.
The swisstable talk on cppcon is a good example of a useful benchmark and optimization that shows how difficult it is to really asses performance effects of even "small" changes. [3]
[1] https://github.com/ClickHouse/ClickBench#data-loading
Re: Lies, Damn Lies and Database Benchmarks
#14Re: Lies, Damn Lies and Database Benchmarks
#15ClickBench is of very limited utility already because it doesn't have a single join in it. Which is maybe less weird in the context of ClickHouse not being great at joins.
Re: Lies, Damn Lies and Database Benchmarks
#16The database wars of the late 1990s were full of this kind of stuff. Oracle, Sybase, IBM etc invested heavily in tuning specifically for benchmarks like TPC-C just so they could post ads in the Wall St Journal saying theirs was faster. I do sympathize with OP, though, their objection to measuring cold-start queries is incomplete without also describing how often cold start needs to happen. If you restart once every f…
The dataset they use is I don't think this is an oversight but it is just what they found to be feasible. This is explicitly written in [1]. Also the guy who setup this benchmark is very serious about benchmarking under difficult conditions [2] My personal opinion is that you need a massive amount of data and massive number of different variables to test for separately. For example you might want to monitor how many…
Re: Lies, Damn Lies and Database Benchmarks
#17The database wars of the late 1990s were full of this kind of stuff. Oracle, Sybase, IBM etc invested heavily in tuning specifically for benchmarks like TPC-C just so they could post ads in the Wall St Journal saying theirs was faster. I do sympathize with OP, though, their objection to measuring cold-start queries is incomplete without also describing how often cold start needs to happen. If you restart once every f…
Re: Lies, Damn Lies and Database Benchmarks
#18Earlier quoted context omitted.
The dataset they use is I don't think this is an oversight but it is just what they found to be feasible. This is explicitly written in [1]. Also the guy who setup this benchmark is very serious about benchmarking under difficult conditions [2] My personal opinion is that you need a massive amount of data and massive number of different variables to test for separately. For example you might want to monitor how many…
Yeah, the tl;dr is that benchmarking is freaking hard because what you actually care about is "does my workload today and in the future run better or worse given current setup?" but identifying what your workload actually is, what systems you are going to be allowed to run it on, what tweaks would even be possible if you know the interiors of a system and how it aligns with your hardware, and it all comes with the pr…
Benchmarking is hard, no argument from me!
Re: Lies, Damn Lies and Database Benchmarks
#19Anyone here using QuestDB in production? What is your use case? What is your experience? We want to migrate away from InfluxDB eventually (because of their 180 on OSS, and their tendency to reinvent the product every major release), and QuestDB seems like an interesting option.