← Insights

DevOps and infrastructure pain points

Not one complaint — the same problem flagged by different people, in different places: the kind of thing developers and users describe in forum threads and issue trackers, not a single loud post. Pulled from public discussions, not from us guessing what might annoy someone, and updated as fresh evidence comes in.

Updated 7/29/2026

Users are struggling with the lack of built-in alerting features in their monitoring tools, particularly for infrastructure-level metrics. Many express a willingness to pay for solutions that can provide alerts based on critical metrics like health checks and memory pressure, which are currently not well supported.

Why this score: The posts indicate a strong desire for a solution, with users explicitly stating they are willing to pay for it, and the problem is common across different platforms.

Already tried: Trying to set up custom alerting using existing tools like Grafana and Prometheus., Using third-party monitoring providers for infrastructure-level alerts.

Why that fell short: Current tools do not support alerting for infrastructure-level metrics., Existing solutions are too focused on application-level issues.

Buying intentTrend: NewConfidence: Medium5 mentions · 2 communities · new

+1 there's no option right now to email/page alert for failing healthchecks and other key metrics like memory pressure. will be willing to pay for this and it would allow us to move off of third-party monitoring providers that are only aware of app-level issues and not infrastructure-level like h...” — flyio-community

Alerts on http://Fly.io Fly.io ? While currently http://Fly.io Fly.io does not have built-in alerting functionality ( not entirely accurate, we alert via e-mail when an application OOMs ), the good news is that you can set up your own monitoring and alerting cluster! You can configure it to alert...” — flyio-community

Users are experiencing various issues with Grafana and Prometheus integration, including timeouts, data not populating, and alert misconfigurations. Many are trying to set up monitoring dashboards but face challenges in getting accurate data and alerts.

Why this score: The issues are specific to integration and configuration, which can be addressed with targeted tools or guides.

Already tried: Trying to restart Prometheus server, Configuring Grafana dashboards with various data sources

Why that fell short: Timeout errors when querying Prometheus, Dashboards showing 'not available' data

Pain pointTrend: NewConfidence: Low6 mentions · 1 communities · new

<p>Metrics collection with Docker Deploy Replica</p> <p>I am a developer, but in my new job, the company doesn't have a DevOps team. So, we don't have any type of metrics collection or proper CI/CD flows. Because of that, I am trying to implement a few things around here, but I am no expert.</p> <p>” — stackexchange

<p>I use Grafana to monitor my company's infrastructure. Everything worked fine until this week, I started to see alerts on Grafana with an error message :</p> <pre><code>request handler error: Post &quot;http://prometheus-ip:9090/api/v1/query_range&quot;: dial tcp prometheus-ip:9090: i/o timeout </” — stackexchange

Users are experiencing performance issues with their monitoring setups, particularly with tools like Grafana and Logstash. Complaints include slow processing times and timeouts, which hinder their ability to effectively monitor their infrastructure.

Why this score: The performance issues are causing frustration and operational challenges, but there is no immediate indication of users looking to pay for a solution.

Already tried: Adjusting configurations to improve performance., Seeking community advice on optimizing their setups.

Why that fell short: Current tools are not optimized for high-performance environments., Complex configurations lead to slow processing times.

Pain pointTrend: NewConfidence: Low3 mentions · 2 communities · new

Hi everyone, I've deployed the following Grafana-based observability stack on AWS ECS: Cloudflare Logpush → S3 S3 → Logstash Logstash → Grafana Alloy Alloy → Loki Loki → Grafana All components are running in ECS containers. Issue 1: Logstash Processing Too Slowly Logstash is processing logs from ...” — grafana-community

Hi everyone, I'm having some issues with my Django Channels WebSocket application scaling configuration. I've set up connection-based scaling with a softlimit of 20 and hardlimit of 25 connections. However, the limits don't seem to be working as expected, and I'm noticing some strange behavior. T...” — flyio-community

Users are experiencing various deployment issues when using Cloudflare, including problems with data synchronization, caching, and MIME type errors. Many are struggling to debug these issues due to a lack of clear error messages or documentation.

Why this score: The issues are specific to Cloudflare deployments and users are actively seeking solutions, indicating a significant pain point.

Already tried: Trying to debug deployment issues with Cloudflare, Seeking help in community forums for Cloudflare-related problems

Why that fell short: Lack of clear error messages during deployment, Difficulty in understanding Cloudflare's caching mechanisms

Pain pointTrend: NewConfidence: Low5 mentions · 1 communities · new

<p>I built a small writing tool (extracts domain vocabulary, organizes learning notes) using AI-assisted coding. I'm not a traditional developer — I use AI to generate the scripts and deploy them, mostly via CLI.</p> <p>I deployed to Cloudflare, and the app runs fine on both mobile and web. No error” — stackexchange

<p>How do I have Cloudflare respect Cache-Control directives for a FastAPI App deployed on Digitalocean App Platform?</p> <h2>Problem</h2> <p>I'm experiencing an issue with caching in my FastAPI application when it's deployed in production on DigitalOcean App Platform, which uses Cloudflare. Locally” — stackexchange

Research this further

These are the strongest signals found so far. Run your own search to add fresh evidence, or check the ideas this points to.