False positive server down messages sent

Incident Report for CloudAMQP

Postmortem

Summary

Between June 25 and June 26, our monitoring system incorrectly reported some dedicated servers as unreachable and sent "server down" notifications for servers that were in fact operating normally. This was a monitoring-side issue only — no customer servers or services were actually down, and no data or message delivery was affected.

Impact

Affected customers received one or more alerts indicating their dedicated server was down. These alerts were false positives. The servers themselves remained healthy and fully available throughout the incident.

Root cause

Before connecting to a server, our monitoring system performs a network authorization step. A defect in how one component handled a transient network hiccup could leave a monitoring process unable to complete that step, after which it failed to connect to the servers it was responsible for checking. The servers themselves stayed healthy and available throughout — the monitoring process had simply lost its ability to reach them, and reported them as down.

Resolution

We identified the affected monitoring processes, confirmed the root cause in production, and deployed a fix that makes this authorization step resilient to such transient failures and able to recover automatically. After deploying, we monitored the system to confirm that notifications returned to normal.

We apologize for the confusion these false notifications may have caused.

Posted Jun 29, 2026 - 11:35 UTC

Resolved

This incident has been resolved.
Posted Jun 26, 2026 - 14:44 UTC

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Jun 26, 2026 - 14:22 UTC

Identified

The issue has been identified and a fix is being implemented.
Posted Jun 26, 2026 - 09:17 UTC

Investigating

We are currently investigating this issue.
Posted Jun 26, 2026 - 08:38 UTC

Identified

The issue has been identified and a fix is being implemented.
Posted Jun 25, 2026 - 14:21 UTC

Investigating

We are currently investigating this issue.
Posted Jun 25, 2026 - 14:10 UTC
This incident affected: Dedicated servers and Backend.