Live data from Hacker News

MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

developers.memsql.com

11–20 of 30 posts

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#11
post #6

Is this one of those, "If you have to ask how much it costs, you can't afford it?" situations? Because "Try it now Free!" is the only thing I can see on their site related to a cost, and "First one's free!" rarely means it will always be free. :|

Their resource estimator starts off at 1640 cores [1] and even the bottom most tick on the scale represents 500 cores. For people who can dedicate that amount of hardware to their database the licensing cost is probably not a major component. The numbers look impressive, though. http://www.memsql.com/why-memsql#scale-out

Try again - I found the bottom ticks were 8 cores and 256GB.

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#12
post #10

How does this compare to kdb+? This seems like a much less arcane competitor.

kdb+ is compressed columnar in memory on a single box with a very exotic language called Q. memsql is row-based in memory across n-machines using SQL.

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#15
post #9

the problem with scaling out to multi cores with a focus on ram is that in larger datasets you end up trading disk latency for network and protocol latency. I a not sure that is a great trade even if we are talking about fiber channel as a medium.

I have to disagree; disk is ancient - it's mechanical egads! - while 10GigE is pretty commonplace now and infiniband and fiber channel are even faster. back from my CS 101 takeaways: there are only 3 bottlenecks in a computer system: CPU, network, and IO. looks like MemSQL is fixing the CPU and IO bottlenecks, but physics is physics so network is pure hardware solution haha

The problem is that you can end up with larger latency over the network because it still takes a fixed amount of time for nodes to communicate. Even with a 1TB/s link between nodes you can still have a good 30ms between them all adding even more latency. That can be mitigated somewhat by a good protocol that can manage that latency properly (e.g. not blocking while waiting on ACKs and such), it can still end up with far more latency than a few large disks would be (even better now with SSDs). That said I do imagine that some datasets will benefit from this kind of topology (I can imagine that geospatial stuff will do well with that, since you can locate physically close things on a single machine and reduce the amount of talking needed).

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#16
post #10

How does this compare to kdb+? This seems like a much less arcane competitor.

kdb+ is compressed columnar in memory on a single box with a very exotic language called Q. memsql is row-based in memory across n-machines using SQL.

kdb+ is also very fast for real-time time series analysis and signal generation. Seeing Morgan Stanley and Credit Suisse in the customer list made me wonder if memsql could become a competitor in the niche that kdb+ currently dominates?

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#17
post #11
post #6

Earlier quoted context omitted.

Their resource estimator starts off at 1640 cores [1] and even the bottom most tick on the scale represents 500 cores. For people who can dedicate that amount of hardware to their database the licensing cost is probably not a major component. The numbers look impressive, though. http://www.memsql.com/why-memsql#scale-out

Try again - I found the bottom ticks were 8 cores and 256GB.

Fair enough - I had considered that to be on the axis since it's the minimum value, and thus not a tick, but visually it certainly is represented as a tick.

At any rate, the point I was trying to make is that they clearly expect you to throw a lot of hardware at these systems.

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#18
post #9

Earlier quoted context omitted.

I have to disagree; disk is ancient - it's mechanical egads! - while 10GigE is pretty commonplace now and infiniband and fiber channel are even faster. back from my CS 101 takeaways: there are only 3 bottlenecks in a computer system: CPU, network, and IO. looks like MemSQL is fixing the CPU and IO bottlenecks, but physics is physics so network is pure hardware solution haha

The problem is that you can end up with larger latency over the network because it still takes a fixed amount of time for nodes to communicate. Even with a 1TB/s link between nodes you can still have a good 30ms between them all adding even more latency. That can be mitigated somewhat by a good protocol that can manage that latency properly (e.g. not blocking while waiting on ACKs and such), it can still end up with…

He was joking.

Re: MemSQL ships 2.0, scales across hundreds of nodes, thousands of cores

#20
post #9

Earlier quoted context omitted.

I have to disagree; disk is ancient - it's mechanical egads! - while 10GigE is pretty commonplace now and infiniband and fiber channel are even faster. back from my CS 101 takeaways: there are only 3 bottlenecks in a computer system: CPU, network, and IO. looks like MemSQL is fixing the CPU and IO bottlenecks, but physics is physics so network is pure hardware solution haha

The problem is that you can end up with larger latency over the network because it still takes a fixed amount of time for nodes to communicate. Even with a 1TB/s link between nodes you can still have a good 30ms between them all adding even more latency. That can be mitigated somewhat by a good protocol that can manage that latency properly (e.g. not blocking while waiting on ACKs and such), it can still end up with…

30ms? In anything resembling a modern datacenter? 0.3-0.5ms is more typical these days.
Post reply on HN