Earlier quoted context omitted.
All conclusions are only valid for similar workloads, but each of MapD and GPUs, Q/kdb+ and Xeon Phi, Redshift, Athena, Big Query, Presto, and Elasticsearch claim to be fast, inexpensive, easy to work with, and otherwise great for Big Data. Which ones really are fast? How fast is fast? How much is this going to cost? Do I need 5 nodes or 50? A few examples of some useful conclusions: - Just because a relatively well-…
These conclusions don't seem very useful because either they are already well established or are not valid. Some examples: Just because a relatively well-optimized PostgreSQL database on a regular workstation takes 5 minutes to run a query doesn't mean you can't get special hardware to run that query faster than you can type. Already well established for years with systems like redis, and more recently with gpu datab…
> I don't see how it's possible for 104GB of csv text data to decompress into only 125GB. For cvs to compress only ~20%...doesn't make sense.
The CSV file itself is around 500 GB. The internal representation, which might use binary formats for numbers, or compress text, uses 125 GB. Redshift expands it to 2TB for all the indexing and mapping.
> Bigger problem: The MapD software itself will be $50,000.
Ouch. That's a rather large oversight. Is the author affiliated with MapD, perhaps?