Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
1–10 of 11 posts
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#2One of the top questions we received was: “How well does it perform at high throughput?”
We ran a load test and would like to share some results with you.
Summary of the test:
- Tested on 20m records
- Kafka produced 55,000 records/sec
- Processing rate of GlassFlow (deduplication): 9,000+ records/sec
- Measured on a MacBook Pro (M3 Max)
- End-to-end latency: Here is the blog post with full test results and tried with different parameters (rps, # of publishers, etc.): https://www.glassflow.dev/blog/load-test-glass-flow-for-clic...
It was important to us to set up the testing in a way that everybody could reproduce. Here are the docs: https://docs.glassflow.dev/load-test/setup
We would love to get feedback, especially from folks consuming high-throughput in ClickHouse.
Thanks for reading!
Ashish and Armend (founders)
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#3My laptop can run 70B LLMs at usable speeds.
I know. Doesn’t scale. No redundancy. No auto redeploy on failures. This is what I mean.
Do we really have to sacrifice this much efficiency for those things or are we doing it wrong? Does the ability to redeploy on failures, cluster, and scale really require order of magnitude performance penalties across the whole stack?
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#4Unless I’m missing some big numbers somewhere you could do that locally on a pi 5 with efficient code. Nothing heroic required, just a decently fast language like Go. My laptop can run 70B LLMs at usable speeds. I know. Doesn’t scale. No redundancy. No auto redeploy on failures. This is what I mean. Do we really have to sacrifice this much efficiency for those things or are we doing it wrong? Does the ability to rede…
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#5Unless I’m missing some big numbers somewhere you could do that locally on a pi 5 with efficient code. Nothing heroic required, just a decently fast language like Go. My laptop can run 70B LLMs at usable speeds. I know. Doesn’t scale. No redundancy. No auto redeploy on failures. This is what I mean. Do we really have to sacrifice this much efficiency for those things or are we doing it wrong? Does the ability to rede…
Totally fair point. For stable, known workloads, you can get really far with something lightweight on a single machine. The challenge comes when you need fault tolerance, scaling, and delivery guarantees without constantly jumping in to fix things. Often heard from data teams talking about data peaks that they cannot predict as easily. But yes, a lot of existing tools make you pay a high-efficiency cost for that. At…
20m records and 9k/sec isn’t very impressive. I would imagine most prospective customers have larger workloads, as you could throw this behind Postgres and call it a day. FWIW I was interested but your metrics made me second guess and wonder what was wrong.
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#6Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#7Earlier quoted context omitted.
Totally fair point. For stable, known workloads, you can get really far with something lightweight on a single machine. The challenge comes when you need fault tolerance, scaling, and delivery guarantees without constantly jumping in to fix things. Often heard from data teams talking about data peaks that they cannot predict as easily. But yes, a lot of existing tools make you pay a high-efficiency cost for that. At…
I think your benchmark may miss the mark a bit if this is your angle. 20m records and 9k/sec isn’t very impressive. I would imagine most prospective customers have larger workloads, as you could throw this behind Postgres and call it a day. FWIW I was interested but your metrics made me second guess and wonder what was wrong.
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#8That site has no scrollbars so I can't read it. Any alternative?
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#9Hi HN, A few weeks ago, we shared GlassFlow: Open Source streaming ETL to dedup and join streams from Kafka for ClickHouse ( https://news.ycombinator.com/item?id=43953722 ). One of the top questions we received was: “How well does it perform at high throughput?” We ran a load test and would like to share some results with you. Summary of the test: - Tested on 20m records - Kafka produced 55,000 records/sec - Processi…
Everything was running on the same machine?
Re: Load Test GlassFlow for ClickHouse: Real-Time Dedup at Scale
#10Hi HN, A few weeks ago, we shared GlassFlow: Open Source streaming ETL to dedup and join streams from Kafka for ClickHouse ( https://news.ycombinator.com/item?id=43953722 ). One of the top questions we received was: “How well does it perform at high throughput?” We ran a load test and would like to share some results with you. Summary of the test: - Tested on 20m records - Kafka produced 55,000 records/sec - Processi…
> - Measured on a MacBook Pro (M3 Max) Everything was running on the same machine?