Live data from Hacker News

Benchmarking OpenTelemetry: Can AI trace your failed login?

quesma.com

1–10 of 85 posts

Re: Benchmarking OpenTelemetry: Can AI trace your failed login?

#3
I've been building an 'sre agent' with LangGraph for the past couple of weeks and honestly I've been incredibly impressed with the ability for frontier models, when properly equipped with useful tools and context, to quickly diagnose issues and suggest reasonable steps to remediate. Primary tooling for me is access to source code, cicd environment and infrastructure control plane. Some cues in the context to inform basic conventions really helps.

Even when it's not particularly effective, the additional information provided tends to be quite useful.

Re: Benchmarking OpenTelemetry: Can AI trace your failed login?

#4

If everyone else is the problem... maybe you are the problem. To me this says more about OTel than AI.

Can you help me understand where you are coming from? Is it that you think the benchmark is flawed or overly harsh? Or that you interpret the tone as blaming AI for failing a task that is inherently tricky or poorly specified?

My takeaway was more "maybe AI coding assistants today aren’t yet good at this specific, realistic engineering task"....

Re: Benchmarking OpenTelemetry: Can AI trace your failed login?

#9
post #6

Our humans struggle with them too. It’s the only domain where you need actually to know everything . I wouldn’t touch this with a pole if our MTTR was dependent on it being successful though.

I can say that as someone that does this for a job for a while, it's starting to be useful in many domains related to SRE that make parts of the job easier.

MCP servers for monitoring tools are making our developers more competent at finding metrics and issues.

It'll get there but nobody is going to type "fix my incident" in production and have a nice time today outside of the most simple things that if they are possible to fix like this, could've been automated already anyway. But between writing a runbook and automating sometimes takes time so those use cases will grow.

Post reply on HN