Skip to content
2026-09-22

An intermittent hardware fault on a link between two Paris availability zones (AZ) degraded access logs, the API, Cellar, the Console, deployments, Materia, metrics, and some managed databases.

On September 22, 2026, between 10:06 and 11:41 CEST, an intermittent failure on a link between two of our Paris AZ degraded several services in the Paris region. Status updates were communicated through:

The following table lists the affected products, the impact observed on each, and when it occurred:

ProductImpactWindow (CEST)
Cellar (object storage, Paris)Requests returned HTTP 50310:06 to 10:17
Materia (key-value, time series)Errors and packet loss on requests crossing the two AZ10:10 to 10:32
MetricsIngestion stopped, then the backlog was cleared by 11:0310:10 to 11:03
Console, API, deploymentsConsole errors, then deployments not progressing until around 11:1410:10 to 11:14
Access logsDelivery delayed10:10 to 11:41
Managed databasesSome add-ons unreachable; one custom instance until 11:2710:13 to 11:27

If you observed effects on your own workloads outside what is described here, the support team can look at your specific case.

Timeline (CEST)

TimeDescription
2026-09-22 09:57One high-capacity link between two Paris AZ starts losing and recovering its signal, hundreds of times over the following minutes
2026-09-22 10:06Internal services report that they can no longer reach each other. Cellar starts returning HTTP 503.
2026-09-22 10:07An external probe raises an alert
2026-09-22 10:10A customer reports the issue. Errors appear on Materia, the console, and the API. Metrics ingestion stops and access log delivery is delayed.
2026-09-22 10:13The routing layer reacts to the link for the first time. Some managed database add-ons become unreachable.
2026-09-22 10:13Traffic keeps being routed over the failing link until 10:23
2026-09-22 10:16First network alert, raised at low priority
2026-09-22 10:17Cellar recovers
2026-09-22 10:24An engineer takes the link out of service manually. Traffic moves onto the remaining paths immediately.
2026-09-22 10:29Metrics ingestion, saturated by client reconnections, is restarted
2026-09-22 10:32Impacted Materia services recover
2026-09-22 10:42An internal control plane component that did not recover on its own is restored
2026-09-22 11:03The metrics backlog is cleared
2026-09-22 11:14Deployments resume
2026-09-22 11:27The last affected managed database instance recovers
2026-09-22 11:41Access log delivery is back to normal
2026-09-22Customers whose add-ons require follow-up work are contacted directly

Analysis

Root cause

The link between the two AZ failed intermittently from 09:57. Our network is designed to handle a link failing outright: traffic is rerouted over the other available paths within seconds. In this case, the link repeatedly recovered for short periods, and the routing layer continued to consider it usable. It first reacted at 10:13, sixteen minutes after the first signal losses, and traffic was still sent over the link until it was taken out of service manually at 10:24. The reason for this sixteen-minute delay is still under investigation.

Diagnostics on both ends of the link were normal throughout the incident. This rules out a cut fiber or a failing laser, and points to a hardware connection issue: a cable, connector, or transceiver. The investigation is ongoing, and the faulty part has not been identified yet.

Three factors extended the impact beyond the network fault itself:

  • The two AZ host complementary halves of several of our clusters. A single link therefore affected object storage, metrics, messaging, and part of our control plane at the same time.
  • When connectivity returned, the volume of client reconnections saturated our metrics ingestion, which had to be restarted.
  • One internal control plane component did not recover on its own. Deployments resumed around 11:14, after it was restored, while the network had been stable since 10:24.

Detection and response

The first signals came from the services themselves at 10:06, followed by an external probe at 10:07 and a customer report at 10:10. The first network alert was raised at 10:16, at low priority. The low-level signal counters on these links, which reflected the fault from 09:57, are not currently monitored.

The link was taken out of service manually at 10:24. The following hour was spent restoring the services that did not recover on their own once the network was stable: metrics ingestion, the control plane component blocking deployments, and the remaining managed database instances.

Actions

Done:

  • Audit the managed database fleet for a service configuration mismatch found during the incident

In progress:

  • Keep the link out of service until the hardware fault is identified and fixed, then restore full capacity between the two AZ
  • Identify the faulty hardware (cable, connector, or transceiver)
  • Determine why the routing layer took sixteen minutes to react to the failing link

Planned:

  • Alert on the low-level signal counters of inter-room links
  • Review whether automated recovery should hold back when a machine’s monitoring is unreachable because of the network rather than the machine
  • Add connection limits to metrics ingestion so that a reconnection storm cannot saturate it

Conclusion

An intermittent hardware fault on a single link between two Paris AZ degraded several services in the Paris region. The link was taken out of service manually 27 minutes after the first signal losses; the rest of the incident was spent recovering services affected by the resulting reconnections. The incident also shows that several of our clusters depend on the connectivity between a single pair of AZ.

The cause of the routing delay is still under investigation. We will update this page once we have the answer.

Last updated on