Earlier quoted context omitted.
They are useful for categorical variables. For example, is a record in the "Likes motorcycles" category? They are fast because (well, one reason) bitwise logical operations are very fast for CPUs to do. Adtech is an example of a sector that benefits from this...they slice and dice datasets a lot to target ad campaigns and such. Being able to do that quickly is useful.
So are you saying that the data is stored in categories which allows for those types of lookups to run faster? Do you have specifics on how the design of a bitmap based database achieves this? How does it maintain these relationships? Just through 0 and 1's? I guess it's easy for me to visualize both row and column based. Im struggling with the bitmaps concept.
FeatureBase: Open-Source, Real-Time Database Built on Roaring Bitmaps
11–14 of 14 posts
Re: FeatureBase: Open-Source, Real-Time Database Built on Roaring Bitmaps
#12any ideas on real life Use Cases?
FeatureBase could be the "feature store" in the middle of the batch prediction section's diagram, or simply be a drop-in replacement for the model's registry.
Re: FeatureBase: Open-Source, Real-Time Database Built on Roaring Bitmaps
#13Why is the bigmap database faster than other distributed database? and what's the differences?
000 - whatever needs to associate with animals, but has no associations currently
001 - whatever it is is associated with having a "mouse" included
111 - whatever it is is associated with having a "dog", a "cat" and a "mouse" included
In the past, high cardinality data sets weren't good for storing in binary form, or a binary index, but nowadays there are ways around this. So, that list of animals could be quite large.
The primary reason it's so much faster is that many CPUs nowadays can do 10s of lookups in a single instruction cycle. That makes them extremely fast.
Re: FeatureBase: Open-Source, Real-Time Database Built on Roaring Bitmaps
#14Earlier quoted context omitted.
They are useful for categorical variables. For example, is a record in the "Likes motorcycles" category? They are fast because (well, one reason) bitwise logical operations are very fast for CPUs to do. Adtech is an example of a sector that benefits from this...they slice and dice datasets a lot to target ad campaigns and such. Being able to do that quickly is useful.
So are you saying that the data is stored in categories which allows for those types of lookups to run faster? Do you have specifics on how the design of a bitmap based database achieves this? How does it maintain these relationships? Just through 0 and 1's? I guess it's easy for me to visualize both row and column based. Im struggling with the bitmaps concept.