Live data from Hacker News

Comparing Database Types

prisma.io

161–170 of 176 posts

Re: Comparing Database Types

#161
post #136

Earlier quoted context omitted.

Also Druid, HBase, Vertipaq (engine behind PowerBI), Redshift, Azure SQL DW, etc Columnar compression is a really interesting engineering problem

ClickHouse is another favourite

Clickhouse is a columnar system, yes, but is not a full-fledged DBMS. Specifically, I don't think it can join tables.

Re: Comparing Database Types

#162

Earlier quoted context omitted.

Also Druid, HBase, Vertipaq (engine behind PowerBI), Redshift, Azure SQL DW, etc Columnar compression is a really interesting engineering problem

memsql (which was in there for new sql).

Eh... not quite. It's in-memory representation is row-based. It seems it uses columnar secondary storage. At least - that's what it says here: https://en.wikipedia.org/wiki/MemSQL

Re: Comparing Database Types

#163
post #140

The document completely overlooks Columnar databases, which are focused on analytics and are much faster than (most, not all) general-purpose DBMSes. See: https://en.wikipedia.org/wiki/Column-oriented_DBMS and https://www.slideshare.net/arangodb/introduction-to-column-o... or get: http://www.nowpublishers.com/article/Details/DBS-024 Examples: * MonetDB * SAP Hana * Actian Vector (formerly Vectorwise) * Oracle In-Memo…

They are much faster for analytical queries, not for transactional ones. If they were overall faster they would be the default relational databases.

Yes, I did say "focused on Analytics" but I didn't clarify that they're only/mostly fast for analytic workloads.

Re: Comparing Database Types

#164

Earlier quoted context omitted.

I know very, very little about any of this, but would this be akin to entity component systems in video games? Forgive me if I'm way off.

Sorry, I don't know what "entity component systems" are. If you mean saving the same component for all entities, rather than saving a bag of components for each of the entities, then sort of.

Close enough to be helpful. Thank you.

Re: Comparing Database Types

#165

Earlier quoted context omitted.

Those are still relational databases, just with column-oriented/column-store tables. I don't see how the storage layer changes the database type. For example, MemSQL has both rowstores and columnstores. Postgres 12 has pluggable storage with column-store (zedstore).

> I don't see how the storage layer changes the database type. It does, because it leads to other types of optimization. LittleTable [0], for example, keeps adjacent data in time domain adjacent in disk. So querying large amount of data that are close to each other is efficient even on slow (spinning) disk. Vertica [1] does column compression which allows it to work with denormalized data (common in analytics workloa…

The underlying database type hasn't changed. LittleTable is a relational database (it's the first sentence in the paper). Vertica is also a relational database.

Stored is an implementation detail. Optimizations are improvements to performance. Neither affects the fundamental data model, which in relational databases is relational algebra over tuple-sets.

Re: Comparing Database Types

#166

Earlier quoted context omitted.

Time-series is more about a specific use-case about data that has a primary time component (like sensor metrics). You can store it in any database, although the common ones are usually some sort of key/value or relational with specific features for time-based queries. Hbase/Bigtable/DynamoDB/Cassandra are key/value. InfluxDB is key/value. Timescale is an extension to Postgres.

If you could store timeseries data "in any database", kdb wouldn't be a thing. Just go and ask a quant trader replace his kdb instance with postgres. (Be prepared to be laughed out of the room.)

I'm not sure what your point is. Time-series describes the data, not the database.

You can store timeseries data in Postgres if you want to (and optionally adding extensions like Timescale). You can store it in key/value like Redis or Cassandra. You can store it in bigtable. You can store it in MongoDB. Obviously different scenarios require different solutions.

KDB is a relational database with row and columnstore with features for time series and advanced numerical analytics along with a programming language. KDB is a thing because of those abilities, whether you use it for time-series data or not.

Re: Comparing Database Types

#167

Earlier quoted context omitted.

If you could store timeseries data "in any database", kdb wouldn't be a thing. Just go and ask a quant trader replace his kdb instance with postgres. (Be prepared to be laughed out of the room.)

I'm not sure what your point is. Time-series describes the data, not the database. You can store timeseries data in Postgres if you want to (and optionally adding extensions like Timescale). You can store it in key/value like Redis or Cassandra. You can store it in bigtable. You can store it in MongoDB. Obviously different scenarios require different solutions. KDB is a relational database with row and columnstore wi…

It is very deceptive to say that you can _store_ timeseries data in "Postgres ... Redis or Cassandra" so the nature of the data should not be used to categorize databases. You can "store" data in /dev/null if you never have to do anything with the data.

> I'm not sure what your point is.

My point is very simple - there is a category of databases widely accepted as "timeseries database", and they deserve a place in any conversation about "types of databases".

Re: Comparing Database Types

#168

Earlier quoted context omitted.

> I don't see how the storage layer changes the database type. It does, because it leads to other types of optimization. LittleTable [0], for example, keeps adjacent data in time domain adjacent in disk. So querying large amount of data that are close to each other is efficient even on slow (spinning) disk. Vertica [1] does column compression which allows it to work with denormalized data (common in analytics workloa…

The underlying database type hasn't changed. LittleTable is a relational database (it's the first sentence in the paper). Vertica is also a relational database. Stored is an implementation detail. Optimizations are improvements to performance. Neither affects the fundamental data model, which in relational databases is relational algebra over tuple-sets.

Timeseries is the data model and that is, for the upper end, synonymous with column-oriented. In my top comment, I mean timeseries/column-oriented (there are other series besiudes time, byt they fit the same data model).

The top TS databases are more than just storage too. You need a query language that can exploit the ordering column-oriented gives you that the row-oriented relational doesn't.

On the lower end (eg, Timescale db) trying to fit a timeseries model on a row-oriented architecture which is a poor fit.

Re: Comparing Database Types

#169

All of them are in fact graph databases, they just didn't realize about it and got lost giving the implementation the category of design for many reasons specific to the context in which they were created. I think we should think more often as mathematicians and a little bit less as "hackers"

I think this is a mischaracterization. The relational model which motivated relational DMBSs is based on predicate logic. Mappings to graphs are obvious, but are not the organizing principle. This was one of the strengths of the relational model, encouraging a more flexible view of the data than graph databases had previously offered. In a complex relational schema, you can discover and work with all kinds of implicit graphs that were not originally intended by the schema design.

Re: Comparing Database Types

#170

Earlier quoted context omitted.

I'm not sure what your point is. Time-series describes the data, not the database. You can store timeseries data in Postgres if you want to (and optionally adding extensions like Timescale). You can store it in key/value like Redis or Cassandra. You can store it in bigtable. You can store it in MongoDB. Obviously different scenarios require different solutions. KDB is a relational database with row and columnstore wi…

It is very deceptive to say that you can _store_ timeseries data in "Postgres ... Redis or Cassandra" so the nature of the data should not be used to categorize databases. You can "store" data in /dev/null if you never have to do anything with the data. > I'm not sure what your point is. My point is very simple - there is a category of databases widely accepted as "timeseries database", and they deserve a place in an…

What category? The examples you used are formally relational databases. We can certainly talk about common use-cases for certain specific database vendors and products but that's not the same as the underlying type.

For example, here's Pinterest handling time-series data on Hbase: https://medium.com/pinterest-engineering/pinalyticsdb-a-time...

There's a big difference and muddying the definitions with marketing jargon ends up causing too much confusion in this industry.

Post reply on HN