Live data from Hacker News

TigerBeetle Core System Architecture: Deconstructing Performance Engineering

ixuvo.com

11–20 of 43 posts

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#12
post #3

I really wish they turned it into a dependency or database framework where users could define their own business logic to swap out the double entry accounting, while reusing all the system architecture and networking features, consensus etc. Sort of like a new paradigm where opinionated custom databases could be created with arbitrary entry logic built on this stack.

llvm for databases?

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#13
Amazing writing, very accessible too.

So batching requests is always something I think should increase performance by a lot, but most server implementations make this pretty difficult, but the thing I struggle the most to understand is how to keep the latency down if you have multiple clients request all batched together? The total amount of latency for all clients is always the latency for the slowest.

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#14

Amazing writing, very accessible too. So batching requests is always something I think should increase performance by a lot, but most server implementations make this pretty difficult, but the thing I struggle the most to understand is how to keep the latency down if you have multiple clients request all batched together? The total amount of latency for all clients is always the latency for the slowest.

If you give the TigerBeetle client a single transfer, it sends it off immediately to the cluster. There's no delay. No Nagle!

But if your application then creates another transfer against the client, and another, while the first request is inflight, then the client will autobatch under the hood and send these off as a batch when the first request returns.

You get this sweetspot then between latency and throughput. And your latency is not spiking as your load increases, since your throughput is now able to keep up.

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#15

Joran from TigerBeetle here! I created TB. Happy to answer questions!

If one has a TigerBeetle cluster with high inter node latency, are there any easy wins to lower the latency of the whole cluster left? My head hurts when I think of latency in large clusters, so your work this year with latency was inspiring.

EDIT: I guess part of the question is about the problems with clusters with >130ms latency and if there are challenges you consider easy.

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#16
post #3

I really wish they turned it into a dependency or database framework where users could define their own business logic to swap out the double entry accounting, while reusing all the system architecture and networking features, consensus etc. Sort of like a new paradigm where opinionated custom databases could be created with arbitrary entry logic built on this stack.

Emphasis on: opinionated custom databases. One database might be SQL-based for periodic report generation. Another database might be a noSQL key-val store designed to effortless grow with the number of end-users. General-purpose programming languages already cater to the 'own business logic' part - it's their whole job. What's left? System architecture, networking features, consensus. That's Kafka.

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#17

Joran from TigerBeetle here! I created TB. Happy to answer questions!

Hey Joran, awesome stuff. I see a TB post every now and then and it seems like such an interesting problem space to work in. I'm all the way at the other end of the stack, most software we write day-to-day is in JVM-based languages where you don't have to think about these things at all (we get by with ms latencies instead of ns). Reading this post inspires me explore low-level engineering more and I figure zig or rust would be a good place to start.

What are some interesting problems or things you can think of to work on that would give someone new a nice amount of exposure to this kind of programming?

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#18

Amazing writing, very accessible too. So batching requests is always something I think should increase performance by a lot, but most server implementations make this pretty difficult, but the thing I struggle the most to understand is how to keep the latency down if you have multiple clients request all batched together? The total amount of latency for all clients is always the latency for the slowest.

I think you design in layers, frontends that work as clients to TigerBeetle for work in batches (as mentioned by the sibling comment), but the whole idea of removing latency differences by removing unpredictability means that you don't get the jitter of latency differences that can cause backing up in normal scenarios.

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#19
post #17

Joran from TigerBeetle here! I created TB. Happy to answer questions!

Hey Joran, awesome stuff. I see a TB post every now and then and it seems like such an interesting problem space to work in. I'm all the way at the other end of the stack, most software we write day-to-day is in JVM-based languages where you don't have to think about these things at all (we get by with ms latencies instead of ns). Reading this post inspires me explore low-level engineering more and I figure zig or ru…

Hey Joos, ah appreciated!

I'd check out tigerstyle.dev, pick up Zig, and then make an HTTP server or file format parser. Those are great ways to learn and experience this kind of programming. At some point, you start to realize that it's just easier to build API services in this way.

But for sure, you can learn so much in JVM-based languages. They make you appreciate low-level techniques all the more!

Re: TigerBeetle Core System Architecture: Deconstructing Performance Engineering

#20
post #15

Joran from TigerBeetle here! I created TB. Happy to answer questions!

If one has a TigerBeetle cluster with high inter node latency, are there any easy wins to lower the latency of the whole cluster left? My head hurts when I think of latency in large clusters, so your work this year with latency was inspiring. EDIT: I guess part of the question is about the problems with clusters with >130ms latency and if there are challenges you consider easy.

Hi, Tobi here from TB. Great question! Generally, 130 ms of network latency is challenging, and there often isn't an easy way around it as you're ultimately constrained by the speed of light (e.g. cross region deployments).

That said, network latency usually follows a distribution. For example, the median might be 130 ms while p99 is 200 ms. So one important goal is to avoid being affected by the high-latency tail.

In consensus and replication systems such as TigerBeetle, you can reduce the impact quite a bit by taking advantage of the fact that you only need a quorum. We have six replicas, and under normal operation we only need acknowledgements from three (including the primary, since we use flexible quorums). That means the primary only has to wait for the two fastest replicas to respond. This is very effective at reducing tail latency.

Then, to get as close as possible to speed-of-light latency, you want to avoid adding unnecessary latency inside the system itself. We've done quite a few algorithmic optimizations there over the past year. For example, introducing radix sort and tournament trees to make CPU processing more efficient.

Post reply on HN