DevOps and infrastructure

Monitoring and Alerting Gaps

Buying intentTrend: NewConfidence: MediumFirst seen 7/29/2026 · last seen 7/29/2026

Opportunity score

47

Mentions

5

Communities

2

Growth since last check

new

What's happening

Users are struggling with the lack of built-in alerting features in their monitoring tools, particularly for infrastructure-level metrics. Many express a willingness to pay for solutions that can provide alerts based on critical metrics like health checks and memory pressure, which are currently not well supported.

Why this score: The posts indicate a strong desire for a solution, with users explicitly stating they are willing to pay for it, and the problem is common across different platforms.

Who's affected

DevOps teamsInfrastructure engineersSite reliability engineers

What people try to do

  • Trying to set up custom alerting using existing tools like Grafana and Prometheus.
  • Using third-party monitoring providers for infrastructure-level alerts.

Why current solutions fail

  • Current tools do not support alerting for infrastructure-level metrics.
  • Existing solutions are too focused on application-level issues.

What to build

  • Develop a monitoring tool that integrates with existing infrastructure and provides customizable alerting based on key metrics.
  • Create a plugin for popular monitoring tools that adds alerting capabilities for infrastructure-level metrics.
  • Offer a consulting service to help teams set up their own monitoring and alerting systems tailored to their specific needs.

Reasons to be careful

  • People already tried: Third-party monitoring providers — this space isn't empty.

Evidence

Request: Alerting based on metrics

+1 there's no option right now to email/page alert for failing healthchecks and other key metrics like memory pressure. will be willing to pay for this and it would allow us to move off of third-party monitoring providers that are only aware of app-level issues and not infrastructure-level like h...

flyio-community · lev1ty · 5/28/2025

Open original →

Self-managed monitoring and alerting on Fly.io with Prometheus + Alertmanager

Alerts on http://Fly.io Fly.io ? While currently http://Fly.io Fly.io does not have built-in alerting functionality ( not entirely accurate, we alert via e-mail when an application OOMs ), the good news is that you can set up your own monitoring and alerting cluster! You can configure it to alert...

flyio-community · roadmr · 5/13/2024

Open original →

Reliability: It's Not Great

The last four months have been rough. We've had more issues than we're OK with. I've hesitated to share this because, well, I'm fighting a debilitating feeling of failure. Fear, too. If we don't improve, our company ceases to exist, and I really like working on this company. One interesting probl...

flyio-community · kurt · 3/6/2023

Open original →

Grafana host Monitoring. What to do if Grafana host goes down?

We are so much relying on Grafana for our entire infrastructure monitoring. But suppose, if something happens to our grafana host, how we can overcome this situation. Suppose data not coming to any particular data-source(As we are using multiple data-sources). Grafana host goes down. We don’t rec...

grafana-community · pinchu · 10/12/2022

Open original →

Detect hosts that stop sending metrics

Hello, we are monitoring our infrastructure using telegraf, InfluxDB and grafana. We also use the alerts from grafana. One graph that we also use for alerts shows the memory usage of our systems grouped by host using a query like this: SELECT mean("used_percent") FROM "mem" WHERE $timeFilter GROU...

grafana-community · dpn2go · 6/29/2017

Open original →

Track signals like this one

Run your own research, save signals you care about, and see how they change over time.