Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
blog.thecloudside.com
Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
1–10 of 31 posts
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#2The industrial historians solve the same problems - collect data at nodes that might have intermittent connectivity, send to a centralized server/service that can handle lots of data, and allow users to plot it.
I wonder if we’ll start to see more open source monitoring on the factory floor. While it will be easy for a product to work as well as industrial offerings, maybe their value is in the long term support (usually close to a decade) and supporter upgrade paths.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#3You don't need to host InfluxDb and Grafana yourself. I would also consider gathering logs and traces to troubleshoot problems. Straightforward with top tier observability vendors, harder to do it on your own.
Disclaimer: I'm employee of Sumo Logic.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#4Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#5It is strange that there isn’t more overlap between tech software monitoring and metrics products and industrial historians and HMI products. Osisoft was purchased by Aveva/Schneider for 5 billion despite them already owning citect, wonderware, and probably 6 other historian products. The industrial historians solve the same problems - collect data at nodes that might have intermittent connectivity, send to a central…
We've also updated all our tender documents for future projects to include a requirement that we can query metrics and logs through an API or direct DB access.
A recent project I worked on identified over 500 applications whose only use is to provide monitoring to a bespoke system or tool. This isn't uncommon at a university as different faculties and departments will buy "the best tool for XYZ" without ever asking IT if perhaps there is a tool that is almost as good that we already have.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#6I bought in to the TICK stack and planned on using an enterprise support contract when going to production, but every interaction with InfluxData the company has felt a bit sleazy. Trying to push very hard to the cloud offering for example.
That’s bad enough, but the documentation and observability of the database is quite poor, and it’s trivially easy to “vanish” all your data and lock your instance up for hours or days by changing the retention policy of a database. (Not making it much different).
Now of course it’s not TICK at all. More like “TI” as kapacitor and chonograph (dashboarding and alerting respectively) are deprecated products and rolled in to the main offering.
Added to that they completely changed the query language.
I have to say; pick something better if you can. TimescaleDB or Prometheus (which uses openTSDB) are promising.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#7If I have a regret in my observability stack I think it’s got to be influxdb. I bought in to the TICK stack and planned on using an enterprise support contract when going to production, but every interaction with InfluxData the company has felt a bit sleazy. Trying to push very hard to the cloud offering for example. That’s bad enough, but the documentation and observability of the database is quite poor, and it’s tr…
I set up 2.x for myself recently, and they have really done a lot of work. The OSS offering has most of the features that cloud/enterprise would. It was easy to set up -- they don't have any instructions for installing it in Kubernetes, and haven't updated their Helm charts for 2.x, but it was like 3 minutes to write a manifest (https://github.com/jrockway/jrock.us/tree/master/production/...) myself, which I prefer 99.9% of the time anyway. The new query language is incredibly verbose, but I see the steps that I remember having with Google's internal system, align, delta, aggregate... all possible. (I had to scratch my head a lot, though, to make it work. And I really am not able to reason about what operations it's doing, what's indexed or not indexed, why I ingest my data as rows but process it as columns, etc.) The performance is good, and it worked well for my use case of pushing data from my Intranet of Stuff. Generally I like it and I don't think they are being shady in any way. It's on my list of something to set up at work to collect various pieces of time series data outside of the Prometheus ecosystem (CI runtimes, etc.).
The reason I picked InfluxDB over TimescaleDB for my personal stuff is because InfluxDB has an HTTP API with built-in authentication. I already a ton of HTTP services exposed to the Internet, and I understand them well. (Yup, I have SSO and rate limiting and all that stuff for my personal projects ;) I can give each of my devices an API key from their web interface, and I make an HTTP request to write data. Very simple. (They have a client library, but honestly my main target is a Beaglebone, and it doesn't have enough memory to compile their client library. I've never seen "go build" run out of memory, but their client makes that happen. I shouldn't develop on my IoT device, of course, but it's just easier because it has Emacs and gopls, and all the sensors connected to the right bus. Was easier to just manually make the API calls than to cross-compile on my workstation and push the release build to the actual device.) TimescaleDB doesn't have that, because it's just Postgres. So I'd basically have to expose port 5432 to the world, create Postgres users for every device, generate a password, store that somewhere, etc. Then to ingest data, I'd connect to the database, tune my connection pool, retry failed requests manually, etc. Using HTTP gets me all that for free; I can just configure retries in Envoy.
But... SQL queries are a lot easier to figure out than FluxQL queries, and I already have good tools for manipulating raw data in Postgres (DataGrip is my preferred method), so I think I will likely be revisiting TimescaleDB. Honestly, I'd pay for a managed offering right now if they had a button in Google Cloud Console that was "Create Instance and by the way this just gets added to your GCP bill for 10% more than a normal Cloud SQL instance".
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#8SaaS version of that would be use Telegraf, but send data to Sumo Logic, Data Dog or other observability vendor. I would also You don't need to host InfluxDb and Grafana yourself. I would also consider gathering logs and traces to troubleshoot problems. Straightforward with top tier observability vendors, harder to do it on your own. Disclaimer: I'm employee of Sumo Logic.
The one thing I would add to this guide is enabling HTTPS for the whole stack, if you are transmitting over the public internet. Fortunately, it is quite straightforward (and free) with Let's Encrypt.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#9Maybe not in this specific case, but in general Prometheus my preferred TSDB sitting between Telegraf and Grafana
If you need further scale-out there are options for federating Prometheus instances as well.
Re: Monitoring Raspberry Pi Devices Using Telegraf, InfluxDB and Grafana
#10It is strange that there isn’t more overlap between tech software monitoring and metrics products and industrial historians and HMI products. Osisoft was purchased by Aveva/Schneider for 5 billion despite them already owning citect, wonderware, and probably 6 other historian products. The industrial historians solve the same problems - collect data at nodes that might have intermittent connectivity, send to a central…
As well as the things you identified, I suspect that there's just a lot of mistrust of open source in the industrial world - there's that whole thing of perceived value being directly proportional to product cost, plus commercial vendors also tend to at least offer training and tech support, even if they're not always the most helpful.