Live data from Hacker News

Application Traffic with eBPF

thebsdbox.co.uk

11–20 of 20 posts

Re: Application Traffic with eBPF

#11
post #7
post #2

Isn't this how tcpdump/ngrep/gopacket work? For parsing the HTTP protocol, I find netpeek effective [1] https://github.com/darshanime/netpeek

tcpdump only uses BPF, not eBPF. BPF is a simpler language that, among other things, is guaranteed to run in finite time because it doesn't have backward jumps, and has limitations on program size (4096 instructions). (The "e" in eBPF stands for "extended", as it extends BPF to remove those limitations, among other changes.) It compiles your filter expression into a series of instructions, using libpcap. For instance…

eBPF are still verified for completion, not just BPF. This is not relaxed in eBPF.

Re: Application Traffic with eBPF

#12
post #9

For those interested, you can also take a look at our open-source project DeepFlow - https://deepflow.io - https://github.com/deepflowio/deepflow We use eBPF to achieve non-intrusive (we call it `zero-code`) observability without modifying any application code, and have implemented three core features: Universal Map, Distributed Tracing, and Continuous Profiling. Yes, we have implemented *Distributed* tracing using e…

I have to say I find projects that talk about generic concepts (observability, tracing, eBPF), but then when you dig in the docs it's 100% Kubernetes-specific, to be highly misleading. Not everyone uses Kubernetes, and distributed tracing is a good thing to have regardless of the underlying platform.

"Cloud Native" == Kubernetes

As in https://www.cncf.io/

Re: Application Traffic with eBPF

#13
post #12
post #9

Earlier quoted context omitted.

I have to say I find projects that talk about generic concepts (observability, tracing, eBPF), but then when you dig in the docs it's 100% Kubernetes-specific, to be highly misleading. Not everyone uses Kubernetes, and distributed tracing is a good thing to have regardless of the underlying platform.

"Cloud Native" == Kubernetes As in https://www.cncf.io/

Yeah, no. Cloud-native used to mean something even before Kubernetes became mainstream, and technically the CNCF isn't about Kubernetes only, which is why KubeCon and CloudNativeCon are separate events (held together, but separate). Just going to their website shows me two case studies, one is around Kubernetes (Spotify), the other around Vitess (Slack) and has nothing to do with k8s.

Re: Application Traffic with eBPF

#14
post #13
post #12

Earlier quoted context omitted.

"Cloud Native" == Kubernetes As in https://www.cncf.io/

Yeah, no. Cloud-native used to mean something even before Kubernetes became mainstream, and technically the CNCF isn't about Kubernetes only, which is why KubeCon and CloudNativeCon are separate events (held together, but separate). Just going to their website shows me two case studies, one is around Kubernetes (Spotify), the other around Vitess (Slack) and has nothing to do with k8s.

Kubernetes heavily dominates there though.

I don't recall people throwing Cloud Native as a term before k8s or outside that space. Google trends seems to confirm they emerged at the same time:

https://trends.google.com/trends/explore?date=all&q=Cloud%20...

Re: Application Traffic with eBPF

#15
post #11
post #7

Earlier quoted context omitted.

tcpdump only uses BPF, not eBPF. BPF is a simpler language that, among other things, is guaranteed to run in finite time because it doesn't have backward jumps, and has limitations on program size (4096 instructions). (The "e" in eBPF stands for "extended", as it extends BPF to remove those limitations, among other changes.) It compiles your filter expression into a series of instructions, using libpcap. For instance…

eBPF are still verified for completion, not just BPF. This is not relaxed in eBPF.

That's a good point. I should clarify that no such verification is necessary in classic BPF due to the absence of backward jumps -- you can trivially show that the maximum steps executed by a BPF program is 4096, since you'll execute each instruction at most once, and there are at most 4096 instructions.

Meanwhile, the verification that an eBPF program terminates is dependent on the correctness of the verifier, and similarly there's no guarantee that a program with appropriately-bounded complexity will be accepted by the verifier.

To be clear: I'm not trying to throw shade at the verifier; to the contrary, I think it's an impressive piece of software. But there's a difference between being able to prove in one sentence that a program always terminates, and needing to rely on the correctness of some verification software.

Re: Application Traffic with eBPF

#16
post #6

For those interested, you can also take a look at our open-source project DeepFlow - https://deepflow.io - https://github.com/deepflowio/deepflow We use eBPF to achieve non-intrusive (we call it `zero-code`) observability without modifying any application code, and have implemented three core features: Universal Map, Distributed Tracing, and Continuous Profiling. Yes, we have implemented *Distributed* tracing using e…

For distributed tracing, how is deepflow able to correlate an inbound request (eg. client call) with an outbound request (eg. 3rd party API call required to service client call) without being inside the business logic?

Our paper provides some explanations: https://dl.acm.org/doi/10.1145/3603269.3604823

Re: Application Traffic with eBPF

#17
post #7
post #2

Isn't this how tcpdump/ngrep/gopacket work? For parsing the HTTP protocol, I find netpeek effective [1] https://github.com/darshanime/netpeek

tcpdump only uses BPF, not eBPF. BPF is a simpler language that, among other things, is guaranteed to run in finite time because it doesn't have backward jumps, and has limitations on program size (4096 instructions). (The "e" in eBPF stands for "extended", as it extends BPF to remove those limitations, among other changes.) It compiles your filter expression into a series of instructions, using libpcap. For instance…

Could you recommend any book/article/video about how eBPF works? - I got a bit interested in this topic but couldn't find anything technical.

Re: Application Traffic with eBPF

#18
post #14
post #13

Earlier quoted context omitted.

Yeah, no. Cloud-native used to mean something even before Kubernetes became mainstream, and technically the CNCF isn't about Kubernetes only, which is why KubeCon and CloudNativeCon are separate events (held together, but separate). Just going to their website shows me two case studies, one is around Kubernetes (Spotify), the other around Vitess (Slack) and has nothing to do with k8s.

Kubernetes heavily dominates there though. I don't recall people throwing Cloud Native as a term before k8s or outside that space. Google trends seems to confirm they emerged at the same time: https://trends.google.com/trends/explore?date=all&q=Cloud%20...

If you remove Kubernetes which is way more popular and hiding the numbers of cloud-native, you'll see that cloud-native started being talked about around 2011, with steady small growth untill it explodes alongside Kubernetes later on.

I recall hearing cloud native compared to lift and shift regarding migrating to AWS ~2012-2013.

Re: Application Traffic with eBPF

#19
post #17
post #7

Earlier quoted context omitted.

tcpdump only uses BPF, not eBPF. BPF is a simpler language that, among other things, is guaranteed to run in finite time because it doesn't have backward jumps, and has limitations on program size (4096 instructions). (The "e" in eBPF stands for "extended", as it extends BPF to remove those limitations, among other changes.) It compiles your filter expression into a series of instructions, using libpcap. For instance…

Could you recommend any book/article/video about how eBPF works? - I got a bit interested in this topic but couldn't find anything technical.

There are a number of good resources at https://ebf.io, including a couple of links to books. I haven't read those books personally, but I would be surprised if "BPF Performance Tools" by Brendan Gregg isn't worthwhile.

Re: Application Traffic with eBPF

#20
post #6

Earlier quoted context omitted.

For distributed tracing, how is deepflow able to correlate an inbound request (eg. client call) with an outbound request (eg. 3rd party API call required to service client call) without being inside the business logic?

Our paper provides some explanations: https://dl.acm.org/doi/10.1145/3603269.3604823

For the curious, the relevant section is 3.3 and the tl;dr appears to be: "heuristics".
Post reply on HN