Request: Alerting based on metrics
“+1 there's no option right now to email/page alert for failing healthchecks and other key metrics like memory pressure. will be willing to pay for this and it would allow us to move off of third-party monitoring providers that are only aware of app-level issues and not infrastructure-level like h...”
flyio-community · lev1ty · 5/28/2025
Open original →Self-managed monitoring and alerting on Fly.io with Prometheus + Alertmanager
“Alerts on http://Fly.io Fly.io ? While currently http://Fly.io Fly.io does not have built-in alerting functionality ( not entirely accurate, we alert via e-mail when an application OOMs ), the good news is that you can set up your own monitoring and alerting cluster! You can configure it to alert...”
flyio-community · roadmr · 5/13/2024
Open original →Reliability: It's Not Great
“The last four months have been rough. We've had more issues than we're OK with. I've hesitated to share this because, well, I'm fighting a debilitating feeling of failure. Fear, too. If we don't improve, our company ceases to exist, and I really like working on this company. One interesting probl...”
flyio-community · kurt · 3/6/2023
Open original →Grafana host Monitoring. What to do if Grafana host goes down?
“We are so much relying on Grafana for our entire infrastructure monitoring. But suppose, if something happens to our grafana host, how we can overcome this situation. Suppose data not coming to any particular data-source(As we are using multiple data-sources). Grafana host goes down. We don’t rec...”
grafana-community · pinchu · 10/12/2022
Open original →Detect hosts that stop sending metrics
“Hello, we are monitoring our infrastructure using telegraf, InfluxDB and grafana. We also use the alerts from grafana. One graph that we also use for alerts shows the memory usage of our systems grouped by host using a query like this: SELECT mean("used_percent") FROM "mem" WHERE $timeFilter GROU...”
grafana-community · dpn2go · 6/29/2017
Open original →