Earlier quoted context omitted.
Building the inverted index is quite CPU-intensive, and we are also merging index files called "splits".
I never being able to understand why log indexing has to build inverted index. Decent columnar store with partitioning by date should be enough to quickly filter gigabytes of logs.
Quickwit 0.8: Indexing and Search at Petabyte Scale
11–20 of 31 posts
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#12Amazing to see how far Tantiviy has come. Remember using and making some smaller contributions to this 3 years ago - slop to phrase queries for example. Curious how the design has changed to enable large scale production usage.
PS: it’s tantivy!!!
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#1313.4GB/s with 200x6 vcpus, gives 11MB/s per core, it is good but hard to say impressive.
What is your frame of reference?
It's very healthy to take maximum bandwidth limits into consideration when reasoning about performance. For instance, for temporal stores, the bottlenecks you see are due to RAM latency and memory parallelism, because of the write-allocate. The load/store uarch can actually retire way more data from SIMD registers.
So there's already some headroom for CPU-bound tasks. For instance 11MB/s is very slow for JIT baseline compiler. But if your particular problem demands arbitrary random access that exceed L3 regularly, maybe that speed is justified.
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#14Never had the chance to use Quickwit at a $DAYJOB (yet?), but I really appreciate the fact that it scales down quite well too. Currently running it on my homelab, after a number of small annoyances using Loki in a single-node cluster, and it's been working very well with very reasonable resource usage. I also decide to use Tantivy (the rust library powering/written by Quickwit) for my own bookmarking search tool by e…
Ah Loki, I wanted to try it at my homelab bit it wasn't as simple as it says. Now I wanted to try Zincsearch or Openobserve. Have you tried that?
mind elaborating? we built loki for some pretty massive scale but I've always tried to make it work at super small scale to. what went wrong?
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#15Amazing to see how far Tantiviy has come. Remember using and making some smaller contributions to this 3 years ago - slop to phrase queries for example. Curious how the design has changed to enable large scale production usage.
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#16Earlier quoted context omitted.
Building the inverted index is quite CPU-intensive, and we are also merging index files called "splits".
I never being able to understand why log indexing has to build inverted index. Decent columnar store with partitioning by date should be enough to quickly filter gigabytes of logs.
After all, it does not matter much if a log search query answers in 300ms or 1s. However, there are use cases where a few GB just does not cut it.
The tale saying that you can always prune your dataset using timestamp and tags is simply not always valid.
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#17Never had the chance to use Quickwit at a $DAYJOB (yet?), but I really appreciate the fact that it scales down quite well too. Currently running it on my homelab, after a number of small annoyances using Loki in a single-node cluster, and it's been working very well with very reasonable resource usage. I also decide to use Tantivy (the rust library powering/written by Quickwit) for my own bookmarking search tool by e…
Ah Loki, I wanted to try it at my homelab bit it wasn't as simple as it says. Now I wanted to try Zincsearch or Openobserve. Have you tried that?
[1] https://github.com/signoz/signoz [2] https://signoz.io/blog/logs-performance-benchmark/
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#18Earlier quoted context omitted.
What is your frame of reference?
Per-core store bandwidth is at least 14GB/s on Zen3, 35GB/s for non-temporal stores. Parsing JSON can be done at +2GB/s. It's very healthy to take maximum bandwidth limits into consideration when reasoning about performance. For instance, for temporal stores, the bottlenecks you see are due to RAM latency and memory parallelism, because of the write-allocate. The load/store uarch can actually retire way more data fro…
The largest work we do is building an inverted index. Oversimplified, it is equivalent to this:
inverted_index = defaultdict(list)
for (doc_id, doc_json) in enumerate(doc_jsons):
c = json.loads(payload)
for (field, field_text) in c.items():
for (position, token) in enumerate():
inverted_index[token].push((doc, position))
serialize_in_compressed_way_that_allows_lookup(inverted_index)You can implement it in a couple of hours in the language of your choice to get a proper baseline.
I am sure we can still improve our indexing throughput... but I have never seen any search engine indexing as fast as tantivy.
If someone knows a project I should know of, I'd be genuinely keen on learning from it.
Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#19Re: Quickwit 0.8: Indexing and Search at Petabyte Scale
#20Earlier quoted context omitted.
Per-core store bandwidth is at least 14GB/s on Zen3, 35GB/s for non-temporal stores. Parsing JSON can be done at +2GB/s. It's very healthy to take maximum bandwidth limits into consideration when reasoning about performance. For instance, for temporal stores, the bottlenecks you see are due to RAM latency and memory parallelism, because of the write-allocate. The load/store uarch can actually retire way more data fro…
What we do is CPU bound and we are not just parsing JSON here. The largest work we do is building an inverted index. Oversimplified, it is equivalent to this: inverted_index = defaultdict(list) for (doc_id, doc_json) in enumerate(doc_jsons): c = json.loads(payload) for (field, field_text) in c.items(): for (position, token) in enumerate(): inverted_index[token].push((doc, position)) serialize_in_compressed_way_that_a…