Prove the Remaining Cooling Path Before Planned Maintenance
Before planned maintenance changes cooling redundancy, prove the exact maintenance-state topology, remaining cooling path and accepted capacity for every affected critical load area.
Before this cooling maintenance starts, show the exact maintenance-state topology, every affected critical load area, the complete remaining cooling path and accepted capacity, all shared dependencies that must remain available, and the executable recovery or rollback path if the state is not maintained.
Issue a controlled maintenance-state diagram or state matrix showing the normal state, each consequential maintenance step, unavailable components, remaining cooling path, shared utilities, control modes, alarms, and recovery state
Map every affected room, rack zone, process area, or owner-designated critical load to the cooling resources that remain available in each maintenance step
Establish the project-approved capacity and environmental acceptance basis for the remaining path using current load, expected conditions, equipment availability, control limits, and any temporary cooling; do not substitute installed nameplate redundancy for maintenance-state proof
Identify shared dependencies that can defeat both nominally redundant paths, including power, controls, compressed air, water, valves, sensors, networks, heat rejection, and maintenance access where applicable
Define the MOP hold points, stop conditions, monitoring, communications, escalation, maximum permitted degraded duration if the project uses one, recovery or rollback actions, and responsible authorities for each step
Immediately before maintenance begins, reconcile the approved state to the field and confirm the required remaining equipment, controls, alarms, support utilities, and monitoring are available
Where practicable and authorized, prove the maintenance-state cooling response through the project's functional or integrated test route before relying on it for consequential live maintenance
On August 27, 2026, Proton experienced a widespread service outage after the cooling system in its Frankfurt data center failed. Proton reported that room temperature rose rapidly after cooling loss and that servers and networking equipment began to fail. Its subsequent investigation traced the root cause to air-filter replacement on both redundant air compressors powering the cooling system. Proton also stated that the operator performed the work without prior notice and did not communicate the cooling failure when it occurred. Cooling was restored after roughly an hour and a half, but some servers suffered heat damage and recovery continued with reduced redundancy. The public report does not disclose the lower-level mechanism by which filter work caused both compressors to become unavailable, the cooling one-line, equipment capacities, or the maintenance procedure.
Evidence to confirm
Controlled cooling maintenance-state topology and step matrix
controls integrator · Before the relevant work begins
Every consequential step shows unavailable cooling components, remaining paths, shared dependencies, control modes, alarms, temporary provisions, and the defined recovery or rollback state.
Critical-load to remaining-cooling-path crosswalk
controls integrator · Before the relevant work begins
Every affected critical room, zone, rack area, or process load is mapped to the cooling resources that remain available in each maintenance state.
controls integrator · Before the relevant work begins
The accepted basis reconciles current or bounded load, expected environmental conditions, available equipment, controls, support utilities, and any temporary cooling to project-defined acceptance criteria.
Shared cooling dependency verification
controls integrator · Before the relevant work begins
Power, controls, compressed air, water, valves, sensors, networks, heat rejection, and other applicable shared dependencies are identified and verified in the state required for the maintenance step.
State-specific maintenance MOP release
controls integrator · Before the relevant work begins
The approved MOP identifies each state, hold point, monitoring requirement, stop criterion, responsible role, communications path, and executable recovery or rollback action.
Immediate pre-maintenance field confirmation
controls integrator · Before the relevant work begins
Named personnel reconcile the actual equipment, load, controls, alarms, support utilities, monitoring, and environmental conditions to the accepted maintenance-state basis immediately before work begins.
Required maintenance-state functional or integrated proof
Conditions to resolve before proceeding
The maintenance-state topology or affected-load mapping does not match the field installation or current operating state
The remaining cooling path or accepted capacity has not been established for the actual load and environmental condition
Two nominally redundant components or paths will be affected by the same maintenance action or shared dependency without an explicitly accepted maintenance-state basis
A required shared support utility, control path, alarm, sensor, power source, valve state, or heat-rejection path is unavailable, bypassed, inhibited, or unexplained
The recovery or rollback sequence cannot be executed from the current maintenance step
The actual load, weather, equipment availability, maintenance duration, simultaneous work, temporary cooling, or control state differs materially from the approved basis
Temperature, pressure, flow, equipment status, or other project-defined monitoring crosses the accepted hold or abort criterion
Where the lesson comes from
Sources
Use the original material to understand the evidence, scope, and context behind this Pearl. Suggested project actions are Build Pearls’ interpretation.
Close Coupled Cooling and Reliability
Uptime Institute · Cooling redundancy and worst-case maintenance/failure discussion · Accessed: 2026-09-03
controls integrator · Before the relevant work begins
Where required by the project, the remaining cooling path demonstrates the specified functional response under the accepted maintenance state and any exceptions are closed or formally accepted before live reliance.
Return-to-normal and as-left record
controls integrator · Before the relevant work begins
All maintenance-isolated components, controls, alarms, valves, support utilities, temporary provisions, and outstanding deficiencies are reconciled to the accepted normal state after work.
The status page is maintained by the same operator as the incident report and is not an independent forensic source.
This article addresses cooling-design reliability generally and does not establish the configuration or root cause of the Proton facility.
This package does not infer a Tier classification, a universal one-at-a-time maintenance rule, or an equipment-specific capacity margin.
This source is general infrastructure guidance and does not describe the Proton incident.
Uptime Institute guidance supports the principle of maintenance-state availability but does not establish the historical facility's Tier classification or prove the Proton root cause.
The project can obtain current cooling drawings, controls information, load data, equipment availability, OEM requirements, maintenance procedures, and operational acceptance criteria.
Qualified project authorities can define the accepted maintenance-state capacity and environmental basis for the specific cooling architecture.
The Pearl will be implemented through existing commissioning, MOP, management-of-change, operations, and owner authorization routes.
Proton · Source date: 2026-08-27 · Status history: August 27, 2026 - Global Outage · Accessed: 2026-09-03