The article never really explains what eBPF is -- AFAIU, it’s a kernel feature that lets you trace syscalls and network events without touching your app code. Low overhead, good for metrics, but not exactly transparent. It’s the umpteenth OTEL-critical article on the front page of HN this month alone... I have to say I share the sentiment but probably for different reasons. My take is quite the opposite: most value i…
OpenTelemetry for Go: Measuring overhead costs
11–20 of 51 posts
Re: OpenTelemetry for Go: Measuring overhead costs
#12The article never really explains what eBPF is -- AFAIU, it’s a kernel feature that lets you trace syscalls and network events without touching your app code. Low overhead, good for metrics, but not exactly transparent. It’s the umpteenth OTEL-critical article on the front page of HN this month alone... I have to say I share the sentiment but probably for different reasons. My take is quite the opposite: most value i…
Re: OpenTelemetry for Go: Measuring overhead costs
#13Logging, metrics and traces are not free, especially if you turn them on at every requests. Tracing every http 200 at 10k req/sec is not something you should be doing, at that rate you should sample 200 ( 1% or so ) and trace all the errors.
100% traces are a mess. I didn’t see where he setup sampling.
Re: OpenTelemetry for Go: Measuring overhead costs
#14The article never really explains what eBPF is -- AFAIU, it’s a kernel feature that lets you trace syscalls and network events without touching your app code. Low overhead, good for metrics, but not exactly transparent. It’s the umpteenth OTEL-critical article on the front page of HN this month alone... I have to say I share the sentiment but probably for different reasons. My take is quite the opposite: most value i…
I don't want to take away from your point, and yet... if anyone lacks background knowledge these days the relevant context is just an LLM prompt away.
Re: OpenTelemetry for Go: Measuring overhead costs
#15The nice thing about Go is that you don't need an eBPF module to get decent profiling.
Also, CPU and memory instrumentation is built into the Linux kernel already.
Re: OpenTelemetry for Go: Measuring overhead costs
#16Funny timing—I tried optimizing the Otel Go SDK a few weeks ago ( https://github.com/open-telemetry/opentelemetry-go/issues/67... ). I suspect you could make the tracing SDK 2x faster with some cleverness. The main tricks are: - Use a faster time.Now(). Go does a fair bit of work to convert to the Go epoch. - Use atomics instead of a mutex. I sent a PR, but the reviewer caught correctness issues. Atomics are subtle a…
Re: OpenTelemetry for Go: Measuring overhead costs
#17Logging, metrics and traces are not free, especially if you turn them on at every requests. Tracing every http 200 at 10k req/sec is not something you should be doing, at that rate you should sample 200 ( 1% or so ) and trace all the errors.
Metrics are usually minimal overheard. Traces need to be sampled. Logs need to be sampled at error/critical levels. You also need to be able to dynamically change sampling and log levels. 100% traces are a mess. I didn’t see where he setup sampling.
FWIW at my former employer we had some fairly loose guidelines for folks around sampling: https://docs.honeycomb.io/manage-data-volume/sample/guidelin...
There's outliers, but the general idea is that there's also a high cost to implementing sampling (especially for nontrivial stuff), and if your volume isn't terribly high then you'll probably eat a lot more in time than paying for the extra data you may not necessarily need.
Re: OpenTelemetry for Go: Measuring overhead costs
#18I feel like this is a lesson that unfortunately did not escape Google, even though a lot of these open systems came from Google or ex-Googlers. The overhead of tracing, logs, and metrics needs to be ultra-low. But the (mis)feature whereby a trace span can be sampled post hoc means that you cannot have a nil tracer that does nothing on unsampled traces, because it could become sampled later. And the idea that if a met…
Re: OpenTelemetry for Go: Measuring overhead costs
#19Logging, metrics and traces are not free, especially if you turn them on at every requests. Tracing every http 200 at 10k req/sec is not something you should be doing, at that rate you should sample 200 ( 1% or so ) and trace all the errors.
How is logging in OTel?
Re: OpenTelemetry for Go: Measuring overhead costs
#20I feel like this is a lesson that unfortunately did not escape Google, even though a lot of these open systems came from Google or ex-Googlers. The overhead of tracing, logs, and metrics needs to be ultra-low. But the (mis)feature whereby a trace span can be sampled post hoc means that you cannot have a nil tracer that does nothing on unsampled traces, because it could become sampled later. And the idea that if a met…
How would you handle the case where you want to trace 100% of errors? Presumably you don't know a trace is an error until after you've executed the thing and paid the price.