Live data from Hacker News

On latency, measurement, and optimization in algorithmic trading systems

architect.co

21–30 of 31 posts

Re: On latency, measurement, and optimization in algorithmic trading systems

#21
post #18
post #3

I remember in 2017, I was trying to benchmark some highly concurrent code in F# using the async monad. I was using timers, and I was getting insanely different times for the same code, going anywhere from 0ms to 20ms without any obvious changes to the environment or anything. I was banging my head against it for hours, until I realized that async code is weird. Async code isn’t directly “run”, it’s “scheduled” and th…

Isn't that part of the point? If the code runs in the scheduler then its performance is relevant. Same with garbage collection, if the garbage collector slows your algorithm down then you usually want to know, you can try to avoid allocations and such to improve performance, and measure it using your benchmarks. Maybe you don't always want to include this, I can see how it might be challenging to isolate just the cod…

> Isn't that part of the point? If the code runs in the scheduler then its performance is relevant.

That's the entire point.

Finding out you have tens of milliseconds of slop because of TPL should instantly send you down a warpath to use threads directly, not encourage you to find a way to cheat the benchmarking figures.

Async/await for mostly CPU-bound workloads can be measured in terms of 100-1000x latency overhead. Accepting the harsh reality at face value is the best way to proceed most of the time.

Async/await can work on the producer side of an MPSC queue, but it is pretty awful on the consumer side. There's really no point in yielding every time you finish a batch. Your whole job is to crank through things as fast as possible, usually at the expense of energy efficiency and other factors.

Re: On latency, measurement, and optimization in algorithmic trading systems

#22
post #20

You generally want the following in trading: * mean/med/p99/p999/p9999/max over day, minute, second, 10ms * software timestamps of rdtsc counter for interval measurements - am17 says why below * all of that not just on a timer - but also for each event - order triggered for send, cancel sent, etc - for ease of correlation to markouts. * hw timestamps off some sort of port replicator that has under 3ns jitter - and a…

How does rdtsc behave in the presence of multiple cores? As in: first time sample is taken on core P, process is pre-empted, then picked up again by core Q, second time sample is taken on core Q. Assume x64 Intel/AMD etc.

In trading all threads are pinned to a core, so that scenario doesn't happen.

Re: On latency, measurement, and optimization in algorithmic trading systems

#23
post #16

In HFT context (as in the article) measurement is quite easy: you tap incoming and outgoing network fibers and measure this time. Also you can do this in production, as this kind of measurement does not impact latency at all

The article also touches on some reasons this isn't enough. You might want to test outside of production, you might want to measure the latency when you decide to send no order, and you might want to profile your code at a more granular level than the full lifecycle of market data to order.

All internal parts are usually measured by low-overhead logger (which materializes log messages in a separate thread, and uses rdtsc in the hot path to record timestamps)

Re: On latency, measurement, and optimization in algorithmic trading systems

#24
post #20

Earlier quoted context omitted.

How does rdtsc behave in the presence of multiple cores? As in: first time sample is taken on core P, process is pre-empted, then picked up again by core Q, second time sample is taken on core Q. Assume x64 Intel/AMD etc.

In trading all threads are pinned to a core, so that scenario doesn't happen.

Yeap!

Also - all modern systems in active use have invariant tsc. So even for migrating threads - one is ok.

Re: On latency, measurement, and optimization in algorithmic trading systems

#25

You generally want the following in trading: * mean/med/p99/p999/p9999/max over day, minute, second, 10ms * software timestamps of rdtsc counter for interval measurements - am17 says why below * all of that not just on a timer - but also for each event - order triggered for send, cancel sent, etc - for ease of correlation to markouts. * hw timestamps off some sort of port replicator that has under 3ns jitter - and a…

> mean/med/p99/p999/p9999/max over day, minute, second, 10ms

So basically you’re taking 18 measurements and checking if they’re Not in finance, but I operate a high-load, low latency service and I’ve always been curious about how y’all think about latency.

Re: On latency, measurement, and optimization in algorithmic trading systems

#26
post #18
post #3

I remember in 2017, I was trying to benchmark some highly concurrent code in F# using the async monad. I was using timers, and I was getting insanely different times for the same code, going anywhere from 0ms to 20ms without any obvious changes to the environment or anything. I was banging my head against it for hours, until I realized that async code is weird. Async code isn’t directly “run”, it’s “scheduled” and th…

Isn't that part of the point? If the code runs in the scheduler then its performance is relevant. Same with garbage collection, if the garbage collector slows your algorithm down then you usually want to know, you can try to avoid allocations and such to improve performance, and measure it using your benchmarks. Maybe you don't always want to include this, I can see how it might be challenging to isolate just the cod…

Yes, but it was initially leading to some incorrect conclusions on my end, about certain things being “slow”.

For example, because I was trying to use fine-grained timers for everything async, I thought the JSON parsing library we were using was a bottleneck, because I saw some numbers like 30ms to parse a simple thing. I wasn’t measuring total throughput, I was measuring individual items for parts of the flow and incorrectly assumed that that applied to everything.

You just have to be a bit more careful than I was with using timers. Either make sure than your timer isn’t going across any kind of yield points, or only use timers in a more “macro” sense (e.g. measure total throughput). Otherwise you risk misleading numbers and bad conclusions.

Re: On latency, measurement, and optimization in algorithmic trading systems

#27
At a past job, I was in charge of investigations as to why latency had changed in a high frequency trading system.

There was dedicated team of folks that had built a random forest model to predict what the latency SHOULD be based on features like:

- trading team

- exchange

- order volume

- time of day

- etc

If that system detected change/unexpected spike etc, it would fire off an alert and then it was my job, as part of the trade support desk, to go investigate why e.g. was it a different trading pattern, did the exchange modify something etc etc

One day, we get an alert for IEX (of Flash Boys fame). I end up on the phone with one of our network engineers and, from IEX, one of their engineers and their sales rep for our company.

We are describing the change in latency and the sales rep drops his voice and says:

"Bro, I've worked at other firms and totally get why you care about latency. Other exchanges also track their internal latencies for just this type of scenario so we can compare and figure out the the issue with the client firm. That being said, given who we are and our 'founding story', we actually don't track out latencies so I have to just go with your numbers."

Re: On latency, measurement, and optimization in algorithmic trading systems

#28
post #26
post #18

Earlier quoted context omitted.

Isn't that part of the point? If the code runs in the scheduler then its performance is relevant. Same with garbage collection, if the garbage collector slows your algorithm down then you usually want to know, you can try to avoid allocations and such to improve performance, and measure it using your benchmarks. Maybe you don't always want to include this, I can see how it might be challenging to isolate just the cod…

Yes, but it was initially leading to some incorrect conclusions on my end, about certain things being “slow”. For example, because I was trying to use fine-grained timers for everything async, I thought the JSON parsing library we were using was a bottleneck, because I saw some numbers like 30ms to parse a simple thing. I wasn’t measuring total throughput, I was measuring individual items for parts of the flow and in…

I would highly recommend using a specialized library like BenchmarkDotNet. It's more relevant for microbenchmarks but can be used for less micro benchmarks as well.

It will do things like force you to build in Release mode to avoid debug overhead, do warmup cycles and other measures to avoid various pitfalls related to how .NET works - JIT, runtime optimization and stuff like that, and it will output nicely formatted statistics at the end. Rolling your own benchmarks with simple timers and stuff can be very unreliable for many reasons.

Re: On latency, measurement, and optimization in algorithmic trading systems

#29
post #28
post #26

Earlier quoted context omitted.

Yes, but it was initially leading to some incorrect conclusions on my end, about certain things being “slow”. For example, because I was trying to use fine-grained timers for everything async, I thought the JSON parsing library we were using was a bottleneck, because I saw some numbers like 30ms to parse a simple thing. I wasn’t measuring total throughput, I was measuring individual items for parts of the flow and in…

I would highly recommend using a specialized library like BenchmarkDotNet. It's more relevant for microbenchmarks but can be used for less micro benchmarks as well. It will do things like force you to build in Release mode to avoid debug overhead, do warmup cycles and other measures to avoid various pitfalls related to how .NET works - JIT, runtime optimization and stuff like that, and it will output nicely formatted…

Oh no argument on this at all, though I haven't touched .NET in several years, since I no longer have a job doing F# (though if anyone here is hiring for it please contact me!).

Even still I don't know that a benchmarking tool would be helpful in this particular case, at least at a micro level; I think you'd mostly be benchmarking the scheduler more than your actual code. At a more macro scale, however, like benchmarking the processing of 10,000 items it would probably still be useful.

Re: On latency, measurement, and optimization in algorithmic trading systems

#30
post #25

You generally want the following in trading: * mean/med/p99/p999/p9999/max over day, minute, second, 10ms * software timestamps of rdtsc counter for interval measurements - am17 says why below * all of that not just on a timer - but also for each event - order triggered for send, cancel sent, etc - for ease of correlation to markouts. * hw timestamps off some sort of port replicator that has under 3ns jitter - and a…

> mean/med/p99/p999/p9999/max over day, minute, second, 10ms So basically you’re taking 18 measurements and checking if they’re Not in finance, but I operate a high-load, low latency service and I’ve always been curious about how y’all think about latency.

[deleted]
Post reply on HN