Live data from Hacker News

Traceway: MIT-licensed observability stack you can self-host in ~90s

github.com

81–90 of 93 posts

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#81

Earlier quoted context omitted.

The implication is that by virtue of using NixOS, you're already a self selected power user. The people that would find setting this thing up in production difficult and the people who would use NixOS are a very small overlap, if any, on that venn diagram.

NixOS is an additional thing on top of prometheus, not a replacement. Not sure why it'd dictate how easy/hard it is to run prometheus & co, you still have to know the same stuff as without it.

You miss the point I’m making which is that if you’re smart enough to use NixOS, you are not representative of gen pop at all. so you saying something is easy doesn’t really say much about the struggles gen pop has with these kinds of things.

In other words, I would absolutely expect a NixOS user to understand how to stand up prom in prod. But I wouldn’t value a statement from that person saying that it’s easy to do so because of course it is. You use NixOS. You might argue that the skill sets are entirely different, and I might accept that argument, but then still say, if you are smart enough and masochistic enough to beat your head against the wall to use NixOS, then deploying prom in prod is obviously going to be trivial to you.

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#82

Earlier quoted context omitted.

If I can ask a separate question: what scalability problems did you run into with Victoria{Metrics|Logs|Traces}, and at what scale did you hit them? VictoriaMetrics and Logs have worked fine in my quiet homelab, and VictoriaMetrics appeared to work great for the infrastructure team of an open source online video game I contribute to (say about 10 physical nodes and 20 applications/services ) ... I was going to sugges…

I honestly think you are a bot. When ever I see Victoria mentioned it is always the same, always asking about hitting a scaling problem + promoting it, never responding to any comments. Hope I'm wrong, but it's been one too many. I refuse to use a product that is this dishonest.

I work at VictoriaMetrics.

Just to clarify: VictoriaMetrics doesn't use bots for HN or for any other media for promotion.

I don't know the person who you responded to. Most of the activity you see is coming from community members who genuinely use the project or from the core engineering team trying to answer user's questions or address misunderstandings.

> never responding to any comments

Could you please share examples like this? I can't say for community members, but our internal policy for engineers is very much focused on great support. You can check our slack/github to see that every question is answered and well explained.

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#83

Earlier quoted context omitted.

I honestly think you are a bot. When ever I see Victoria mentioned it is always the same, always asking about hitting a scaling problem + promoting it, never responding to any comments. Hope I'm wrong, but it's been one too many. I refuse to use a product that is this dishonest.

I work at VictoriaMetrics. Just to clarify: VictoriaMetrics doesn't use bots for HN or for any other media for promotion. I don't know the person who you responded to. Most of the activity you see is coming from community members who genuinely use the project or from the core engineering team trying to answer user's questions or address misunderstandings. > never responding to any comments Could you please share exam…

Hi, first of all, thank you for your response, I really appreciate actual Victoria team commenting.

This is probably the 5th comment (almost identical) I have seen about VictoriaMetrics, mostly on Reddit. I engaged with a few trying to learn more about your product and eventually just gave up. If you really want you can comb through my reddit comments, but be warned, I have commented on a lot of things... a lot...

You should be proud of what you have built, I've looked a bit more and your product looks incredible. I personally think that a sales person might have been testing an automation tool, but if it was actual customers that just shows how good the product is!

Traceway is not working on addressing the problems y'all are solving, it is more focused on having an out of the box experience with preconfigured dashboards, SLOs, integrations, automatic endpoint ranking, frontend session replays/RUM, symbolication etc.

Again, thank you for your comment.

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#84

Earlier quoted context omitted.

There are ways to scale Prometheus (look at Thanos), but none of the solutions is really bug free. See this PR for example ( https://github.com/prometheus/prometheus/pull/18364 ) - this used to impact a production deployment I worked on. Prometheus, Thanos and even OpenTelemetry are full of those kind of problems - but at the same time it's the best we have and we should be grateful they're free and open source. I'd…

> It's ironically cheaper to build a team of engineers that will maintain a sane observability stack instead of feeding the monster(s). Can you show the math here? This is a very bold claim, and I’m super curious. A shared Google Sheet would work well.

[flagged]

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#85

Earlier quoted context omitted.

If I can ask a separate question: what scalability problems did you run into with Victoria{Metrics|Logs|Traces}, and at what scale did you hit them? VictoriaMetrics and Logs have worked fine in my quiet homelab, and VictoriaMetrics appeared to work great for the infrastructure team of an open source online video game I contribute to (say about 10 physical nodes and 20 applications/services ) ... I was going to sugges…

I honestly think you are a bot. When ever I see Victoria mentioned it is always the same, always asking about hitting a scaling problem + promoting it, never responding to any comments. Hope I'm wrong, but it's been one too many. I refuse to use a product that is this dishonest.

Hi, I am not a bot. Also I do not work for VictoriaMetrics.

Please feel free to go through my post history and observe I comment on things I am interested in, like databases, servers, and video games.

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#87
post #8

There's a few contenders in self-hostable otel: - ClickStack (ex HyperDX) - SigNoz - Traceway - a few more does someone has enough feedback on those to be able to tell which one works best?

Hi, creator of Traceway here. I have not used SigNoz or ClickStack. I believe both are very good products that focus on slightly different things. With Traceway I am trying to focus on providing a pre configured system that works out of the box, tells you whats wrong and what to fix. It comes with a great issue tracker, session replays/RUM, preconfigured Dashboards and it's easy to host. It has an alerting integratio…

I keep seeing you pop up... good to meet you. I'm new to the platform and find all this very interesting... and to think - I was making a better logger for a godot game...

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#88

Earlier quoted context omitted.

I honestly think you are a bot. When ever I see Victoria mentioned it is always the same, always asking about hitting a scaling problem + promoting it, never responding to any comments. Hope I'm wrong, but it's been one too many. I refuse to use a product that is this dishonest.

Hi, I am not a bot. Also I do not work for VictoriaMetrics. Please feel free to go through my post history and observe I comment on things I am interested in, like databases, servers, and video games.

Got it, sorry, your comment just looked like a bunch of others and felt extremely out of place as nobody mentioned hitting any limits, especially with the Victoria stack (that I could see).

The comment read out of place/generic and given my previous experience I incorrectly assumed it was another generic bot - my bad.

Hopefully no hard feelings and Victoria looks great.

What do you like the most about it and how was your experience scaling it?

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#89

At KubeCon Europe a very good chunk of booths were observability stacks. Everyone was claiming they're better than the competitors (with some of the just justifying themselves by saying "it's written in Rust). Having dealt with Prometheus (+Thanos) / Grafana / OTEL and other stacks (e.g: custom solution on ClickHouse, Victoria{Metrics,Logs}, Jaeger/Tempo, Loki, ...) and even cloud ones (Google's Monarch rebranded as…

FWIW we've also tried all sorts of different things, and honestly the very vanilla (prometheus -> central thanos, fluentbit -> central loki, grafana) ends up on top. The resource consumption is surprisingly minimal (for a sense of scale, we run about 200k eps for metrics and 1k eps for logs). For all these solutions, I find myself asking the same question as you.. what problem are you trying to solve? Is there anythi…

Hi, sorry for not responding sooner this one slipped through the cracks. I've tried to explain my reasons for starting traceway and what I've been building with it. It's not aimed to be the fastest ingestion tool out there, but it's backed by clickhouse and does minimal processing of the data, you can expect the perf to be as close to clickhouse as possible.

I'm working on a comprehensive benchmark of Traceway performance on different hardware configurations. The most I've tested with was the smallest managed ch instance with 250k traces per sec, handled it without a hiccup (but that's empirical). You can checkout the traceway git, there is an issue I've opened for benchmarking and you can subscribe/comment on it if you're interested. I'm benchmarking across sqlite, self hosted clickhouse and managed clickhouse. I am a huge fan of systematic, realistic and most of all reproducible benchmarks, so I am really excited about the progress on that.

Anyhow, you can checkout traceway and see what it offers, it's aimed at providing SLOs out of the box, session replays, alerting, configurable dashboards and great exception tracking (automatic symbolication) etc...

Re: Traceway: MIT-licensed observability stack you can self-host in ~90s

#90

Earlier quoted context omitted.

Hi, I am not a bot. Also I do not work for VictoriaMetrics. Please feel free to go through my post history and observe I comment on things I am interested in, like databases, servers, and video games.

Got it, sorry, your comment just looked like a bunch of others and felt extremely out of place as nobody mentioned hitting any limits, especially with the Victoria stack (that I could see). The comment read out of place/generic and given my previous experience I incorrectly assumed it was another generic bot - my bad. Hopefully no hard feelings and Victoria looks great. What do you like the most about it and how was…

No worries, no hard feelings, I was just surprised that what I thought was a specific response was assumed to be a generic-ish bot response. (then again, I didn't spell out that the game I contribute to is Beyond All Reason.) I do totally sympathize with the feeling of being overwhelmed by AI slop.

After digging into Traceway documentation, it looks like you were looking to primarily use OTEL for ingestion? Or would you say that's a misreading of the documentation and you actually support metrics, logging etc easily? It looks easy to setup via docker, I might try the SQLite version just to get a taste for how it works and how easily data can be ingested.

For myself, I was initially interested in the Loki/Prometheus/Grafana stack but it wasn't going to fit in the 4GB of RAM I had available on a Raspberry Pi that was already hosting two services that consumed a GB of RAM each. So when I found VictoriaMetrics (a) happily ran in 200MB of RAM (b) was used by CERN (c) had excellent, comprehensive documentation with plenty of examples (d) supported so many different ingestion and export/reporting APIs that I would be able to set up everything I wanted for my homelab without any shim scripts or one-off API converters and (e) offered a basic reporting UI with sane defaults (auto-detecting rate vs sum for a graph) even without having to set up Grafana, I was blown away and grateful that such a useful thing existed. Same for VictoriaLogs, it was just easy to set up once I put my mind to it, because the documentation for everything was very clear, and they clearly had "sane defaults + configurable options" once you needed something slightly different. Having sane support for backfills and tolerating duplicates was also nice. "Throw us your data in one of these shapes , we'll sort it out" was just nice to finally see rather than digging through pages of Prometheus documentation for what the edge cases could be if I sent duplicates or the data was from a month ago.

I just have a homelab of random docker container across a few nodes thrown together with underpowered hardware, but VictoriaMetrics met me where I was and made it trivial to experiment using the nodes I had rather than have to migrate to bigger nodes, and it was very well behaved at idle, steady-state, and "I want to trickle-feed a million data points via http calls" loads. I don't yet need OTEL, I don't have cattle, I have homelab pets and very little time to play with them. I just want to either scrape metrics or fire metrics at some sort of endpoint that can figure out what I meant if I get close enough.

But VictoriaMetrics was so easy to get working because the documentation was laid out as "here's the starter command line options, here's how you ingest data in a variety of input URLs, here's how you retrieve your data via a variety of output URLs, if you want specialty stuff that's described farther down the page..." it was about as hard as falling off a log. It just became the obvious place to base anything else around because it had so many connectors and sane defaults.

So when the Beyond All Reason infrastructure team asked "is there a infrastructure and application metrics solution for a handful of nodes that is self-hosted, easy to set up and won't break the bank or require babysitting?" I had one recommendation: VictoriaMetrics (+ Grafana)

Admittedly I do sort of wish for unified metrics and logs and traces, but that's merely a platonic ideal dream state for me. In reality I can see that both I and an organization generally sets up metrics, or logs, or traces, in a piecemeal fashion. An organization (in my limited experience) generally doesn't think about all three at once, and so the "do one thing and do it well" becomes a nice simplification of scope rather than a mark against VictoriaMetrics or VictoriaLogs not having the whole enchilada under one common roof.

I have not personally worked on scaling it horizontally yet, and I didn't set it up myself, but (a) I observe the Beyond All Reason VictoriaMetrics server has 8 GB of RAM, 3 vCPU and appears to serve 75k active time series (14.5 billion data points, ingest about 5 thousand data points per second) without complaint. The resource usage graphs are flat, humming quietly and (b) I did appreciate that the vmagent and vlagent do send to multiple targets easily (tested this with vlagent) , making "active -> standby fail-over" easy to setup -- all ingestion agents would multiplex to all sinks and you were done, any sink "should" have the same data.

Post reply on HN