Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Power System Maintenance & Troubleshooting in Data Centers
Rating: 4.5 out of 5(9 ratings)
47 students

Power System Maintenance & Troubleshooting in Data Centers

Master 10 failure scenarios, from UPS breakdowns to generator faults — and build expert-level MOPs for maintenance
Created byEyob Jagema
Last updated 4/2025
English

What you'll learn

  • Identify and troubleshoot the top 10 power-related failures in Tier 3 data centers, including UPS faults, generator issues, and PDU overloads.
  • Perform routine and preventive maintenance on critical power systems such as ATS, STS, UPS, and generator infrastructure.
  • Analyze real-time alarms and interpret BMS/BEMS feedback to make safe, accurate decisions during live incidents.
  • Apply industry best practices and method-of-procedure (MOP) principles to reduce risk and ensure electrical reliability during switching and maintenance.

Course content

2 sections12 lectures35m total length
  • Scenario 1: Utility Power Failure2:29

    Scenario 1: Utility Power Failure

    This scenario covers the sequence of events during a complete utility power loss in a Tier 3 data center. You’ll learn how the UPS systems bridge the gap, how standby generators start automatically, and the role of the catcher system as a last line of defense. The focus is on how a shift engineer should respond in real time to protect critical loads.


  • Scenario 2: Generator Start Failure2:00

    Scenario 2: Generator Start Failure

    In this scenario, the utility supply fails—but the generator doesn’t start. You’ll explore common causes such as battery failure or fuel system faults, and how to respond using manual start procedures, alarm diagnostics, and escalation protocols. You’ll also review best practices for preventing this failure through routine testing and inspections.

  • Scenario 3: UPS Battery Failure1:45

    Scenario 3: UPS Battery Failure

    In this scenario, you’ll face a situation where a brief utility flicker occurs, but instead of seamless backup, the UPS drops the load. The cause? A degraded or open battery string. This module walks through how to diagnose UPS battery issues using control panel readings, runtime data, and voltage checks. You’ll also learn how to assess risk in real time and trigger escalation or load shedding if backup runtime is compromised.

  • Scenario 4: ATS/STS Transfer Failure3:27

    Scenario 4: ATS/STS Transfer Failure

    This scenario explores what happens when an Automatic Transfer Switch (ATS) or Static Transfer Switch (STS) fails to complete a transfer between power sources — leaving the load in limbo. You’ll learn how to identify mid-transfer faults, use touchscreen interfaces or mimic panels to diagnose the failure, and safely engage manual bypass procedures if authorized. The focus is on minimizing downtime and ensuring safe switching during partial outages or switching delays.

  • Scenario 5: Overloaded PDUs2:17

    Scenario 5: Overloaded PDUs

    This scenario addresses what happens when new server loads push a PDU circuit or phase over capacity — causing a breaker trip and unexpected outages. You’ll learn how to use real-time monitoring tools to detect overload trends, balance loads across phases, and coordinate with IT to redistribute or shed load. The focus is on proactive capacity management and fast incident response.

  • Scenario 6: Human Error During Switching1:43

    Scenario 6: Human Error During Switching

    This scenario highlights the risk of operator error — such as isolating the wrong breaker during live maintenance. You’ll learn how to apply the two-person rule, verify labels using mimic diagrams, follow LOTO procedures, and perform full MOP walk-throughs. The key takeaway is that human error is avoidable with the right culture and process.

  • Scenario 7: Fuel Contamination or Shortage1:33

    Scenario 7: Fuel Contamination or Shortage

    In this scenario, the generator fails mid-test despite showing a full tank — the cause is fuel contamination. You’ll learn how to identify signs of water, sludge, or microbial buildup in diesel tanks, compare sender readings to physical levels, and take immediate action by switching tanks or calling for fuel polishing. This scenario emphasizes proactive inspection and fuel quality management.

  • Scenario 8: Harmonics and Electrical Noise1:38

    Scenario 8: Harmonics and Electrical Noise

    This scenario introduces harmonic distortion caused by non-linear loads like blade servers and switching power supplies. You’ll learn how to detect excessive THD (Total Harmonic Distortion) using power analyzers, interpret UPS alarms related to waveform distortion, and implement fixes like harmonic filters or load balancing. Clean power is the goal — and this module shows how to maintain it.

  • Scenario 9: Cooling System Power Loss1:11

    Scenario 9: Cooling System Power Loss

    In this scenario, a power feed issue causes CRAC or AHU units to go offline, and the temperature in the data hall begins to rise rapidly. You’ll learn how to identify failed cooling units using the BMS, safely restart them, monitor environmental trends, and escalate when N+1 redundancy is compromised. This module emphasizes fast response to thermal threats and the importance of UPS-backed controls.

  • Scenario 10: False Alarms Triggering Escalation2:25

    Scenario 10: False Alarms Triggering Escalation

    This scenario explores a situation where a BMS alarm — like “GENERATOR NOT IN AUTO” — appears during scheduled maintenance. You’ll learn how to differentiate false alarms from real threats, verify against MOPs or schedules, and communicate calmly with your team. The focus is on using BMS Watch Mode, alarm suppression, and proper coordination to prevent unnecessary panic and wasted response efforts.

  • Recap1:31

    Recap:

    In this course, you explored the 10 most critical power-related issues that challenge uptime in Tier 3 data centers. From system failures to human mistakes, each scenario was designed to prepare you for real-world incidents and responses.

    You learned how to manage utility outages, respond to generator failures, and recover from UPS breakdowns. You tackled phase imbalance, overloaded PDUs, and transfer switch malfunctions. You saw how human error during switching can trigger major outages — and how careful planning, verification, and Method of Procedure (MOP) usage prevents it.

    You diagnosed fuel contamination, tracked down harmonic distortion, reacted to cooling system losses, and navigated the confusion of false alarms.

    Every scenario was more than just a lesson — it was a reminder that in a live critical environment, power isn’t just electrical — it’s operational responsibility.

    Now, you’re equipped to troubleshoot faster, maintain smarter, and lead with clarity. These 10 issues represent the real risks on shift — and you now know how to face them.


Requirements

  • Basic understanding of electrical systems (e.g., voltage, current, circuit breakers) • Familiarity with standard safety practices in electrical or facility environments • Some experience in data center operations, electrical maintenance, or technical fieldwork is helpful but not required • Comfort reading one-line diagrams or equipment schematics is a plus

Description

This course is designed for data center professionals, electrical operations technicians (EOTs), and facility engineers responsible for maintaining and troubleshooting power infrastructure in mission-critical environments. You’ll explore 10 real-world power failure scenarios that challenge even the most experienced technicians — from UPS battery faults and generator start failures to overloaded PDUs, human error, false alarms, and cooling system dropouts.


Each scenario breaks down the cause, shows the correct response, and explains the preventive measures that should be in place. You’ll gain the skills to work under pressure, respond to alarms, investigate faults through BMS or EPMS systems, and coordinate effectively with your operations team to protect uptime.


You’ll also learn how to create a professional, compliant Method of Procedure (MOP) — an essential document in Tier 3 and Tier 4 environments. This includes defining the scope, writing step-by-step procedures, identifying risk, integrating rollback plans, and ensuring all work is executed safely, consistently, and with proper sign-offs.


Whether you’re just entering the field or already working in data center operations, this course will strengthen your technical understanding, situational awareness, and confidence.


By the end of the course, you’ll be able to troubleshoot power failures, maintain critical systems, develop risk-based MOPs, and document your actions like a true shift leader — prepared, reliable, and accountable in any critical environment.


Who this course is for:

  • Electrical Operations Technicians (EOTs) working in Tier 3 or Tier 4 data centers • Facility engineers and technicians responsible for power maintenance and fault response • New hires or junior staff preparing to take on power-related responsibilities in mission-critical environments • IT infrastructure professionals who want a better understanding of power dependencies and risk points • Anyone preparing for a role in data center operations or critical facility management