Let me explain my experience with tsdb selection. in 2018 we understood that we need something for long term storage.
Selection was between thanos, elasticsearch and m3db.
M3db looks promising but after reading issues and docs i found i have to test it like database not like a drop in solution. for example that topic https://groups.google.com/forum/#!topic/m3db/6iG2NL7hJ7A
And cortex and thanos announced that tweet https://twitter.com/fredbrancz/status/1043060822988259333
Elasticsearch got disqualification because of no remote_read support. So i stopped looking for anything for at least half year just updated retention policy in Prometheus to 150d.
Also VictoriaMetrics was banned because of no source code.
Also https://github.com/akumuli/Akumuli was banned because of nobody hear about it. :(
Than after some time VictoriaMetrics appears to be open source and there was no issues with rate function and useless extrapolation.
So i test it on a small setup at about 7k metrics per second on single server. And it was amazing.
Than 14k/s and 20k/s
Previously i have the volume for Prometheus data and it was about 30 gigs on smallest install to 70 gigs on largest
Moving from storing 30days in Prometheus to 90 days in vm was the huge benefit.
On every of three instances with 7, 14 and 20k metrics per second i can extend retention from 3 to 5x on the same volume. With same dashboards. With same alerts. Just added remote read and remote write.
Than i decide to take it on a real life web scale production.
So i started from 11 servers f2-micro on gcp.
* 3 storage
* 2 insert nodes
* 2 select nodes
* 2 promxy
* 1 grafana
* 1 selfmon prometheus
Got lots of expected ooms on 60k per second. Than i move to n1-standart-1 for storage and insert.
It can handle at about 650k per second insert load for several weeks without ooms or any unaxpected behaviors.
That was real life data from prometheus-operator from one of our rc clusters. node-exporters, application metrics and kubemetrics.
Tuning it to n1-highmem-2 for storage nodes so get enough room for background merges and so on.
Also i copy my Prometheus rules from prod to promxy (at about 200 in sum).
That makes some noise for read. So i got at about 70 reads per second and pretty 90% cpu utilization on promxy servers.
But almost no additional cpu load on vm servers. So i just bump all numbers in queries from seconds i moved to minutes, minutes to hours and hours to days in every query that have offset or rate or increase.
That add some load to vm but not that much i expected.
In summary i'm amazed with simplicity of scheme i got. Performance is also great. My dashboards looks the same in Prometheus and VictoriaMetrics.
Oh. by the way i have some experience asking questions in issue tracker of Prometheus and VictoriaMetrics. And honestly prefer Aliaksander style of answering - long and with good under the hood info.