Building performance

Condition monitoring from electrical data

What energy-monitoring data can tell you about motors, pumps, fans and compressors, how to build a baseline, and when you need vibration or thermal measurement instead.

Condition monitoring measures the state of a machine while it runs, so that maintenance follows evidence and not a calendar. ISO 17359 sets out the procedure: rank the equipment by criticality, identify the failure modes that matter, select measurements that respond to those failure modes, set a baseline and act on the changes.

Many sites already measure their large motors, pumps, fans and compressors for energy reporting. The current sensors and the meter are installed, and the data is logged. The extra cost of condition monitoring from that data is the baseline work and the alert rules, not new hardware.

When condition monitoring pays

Run-to-failure is correct for a cheap, non-critical item. A fixed service interval is correct when wear is predictable and the service is cheap. Condition monitoring pays on the assets where one unplanned stop costs more than several years of monitoring: a process chiller in a food plant, the duty air compressor on a production line, a wastewater lift pump with no standby, or a large motor with a rewind lead time of several weeks. Rank the assets with the criticality step in ISO 17359, and start at the top of the list.

What electrical data can show

A three-phase monitor such as the ZEM records RMS current and voltage, active and apparent power and power factor on each phase, plus frequency and cumulative energy. The useful condition indicators come from these readings, compared at the same operating state:

IndicatorDirectionLikely causeConfirm with
Power at the same flow (pump, fan)Rises over weeksFouled strainer or heat exchanger, impeller wear, bearing drag, misalignmentDifferential pressure, bearing temperature
Power at the same speed reference (belt-driven fan)Steps downBelt slip or a broken belt, a closed damperFan speed or airflow
Energy per unit of output (kWh/m³, kW per m³/min)Rises over weeksLoss of efficiency anywhere in the systemThe output meter and its calibration
Compressor load with production stoppedAbove zeroCompressed air leaksLoad and unload times during a shutdown
Starts per hourRisesShort cycling, a failed sensor, a control faultController set points and dead bands
Current unbalance with balanced voltageRisesA high-resistance connection, a winding faultThermal image of the terminals, insulation test
Voltage unbalanceAbove 1%Uneven single-phase loads, a supply faultSingle-phase loads on the same board
Power factor (direct-on-line motor)FallsThe motor now runs more lightly loaded, for example after a process changeActive power against the rated input power

Winding and rotor faults rarely show clearly in interval power factor. A falling power factor on a direct-on-line motor usually means a load change, because the magnetising current stays almost constant while the working current falls.

Current is a poor measure of motor load below about 50% of full load. The DOE motor load fact sheet shows that the current curve becomes non-linear in that region, again because of the magnetising current. Use active power, not current, to estimate load on a lightly loaded motor.

Speed changes the numbers more than any fault does. The affinity laws put pump power in proportion to the cube of speed. A reduction from 1500 to 1200 rpm (80% speed) cuts ideal power to 0.8³, or 51%. At 60% speed the pump draws about 22%. A raw power trend from a variable-speed pump mixes duty changes with condition changes, so group the readings by speed or flow first:

Voltage unbalance: a worked example

Voltage unbalance is the easiest electrical check to act on. Under unbalanced voltage the motor currents are unbalanced by about 6 to 10 times the voltage unbalance. NEMA MG 1-14.35 gives the effect on heat: in the phase with the highest current, the percentage increase in temperature rise is about twice the square of the percentage voltage unbalance. At 3.5% unbalance that is about 25%. The DOE tip sheet recommends that unbalance at the motor terminals stays below 1%. NEMA MG 1 derates the motor above 1% and advises against operation above 5%.

The calculation uses the three line-to-line voltages. For 460 V, 467 V and 453 V, the average is 460 V and the largest deviation is 7 V:

The result is 7 ÷ 460 = 1.52%. That is above the 1% limit. The currents can be 9 to 15% unbalanced, and the hottest winding runs about 2 × 1.52² = 4.6% hotter than on a balanced supply. The DOE tip sheet gives the efficiency cost for the motor in its table, at full load: 94.4% on a balanced supply, 93.0% at 2.5% unbalance.

If the voltage unbalance is below 1% but the current unbalance on one motor is high, the cause is in that circuit. Check the terminations and the windings before you check the supply.

Motors on variable-speed drives

Most of the pumps and fans in this guide run on variable-speed drives. The drive changes what the monitor can see.

Fit the monitor on the line side of the drive. The ZEM is specified for 50 Hz or 60 Hz supplies, and the drive output runs at a variable frequency with a switched waveform. On the line side, active power is the motor input power plus the drive losses of a few percent. A power-against-flow baseline still works.

The other electrical checks change. Line-side power factor describes the drive's rectifier, not the motor. Voltage unbalance at the drive input affects the drive, but the motor receives a balanced voltage that the drive makes. A winding fault can therefore hide behind the drive. Read the drive's own output current, speed and fault log over Modbus RTU and use the speed as the state signal for the baseline.

Measurement quality and data events

A baseline compares the measurement chain with itself. A fixed calibration error is the same in the baseline and in today's reading, so it cancels. Two types of error do not cancel.

The first is error that changes with the operating point. Meter error limits widen at low current and at low power factor, and CT phase error has most effect on power at low power factor. Choose a current sensor range that suits the running current of the motor. A 3000 A Rogowski coil on a 40 A motor circuit puts every reading at the bottom of its range.

The second is a step change in the measurement chain. A replaced CT with a different ratio, a changed scaling factor or a meter swap moves every reading by a fixed percentage at one moment. A model reports this as a fault. A CT that is fitted the wrong way round gives negative active power on that phase. Two CTs swapped between phases give implausible power factors on both phases. After every electrical intervention, check the sign of the active power and the power factor on each phase, and record the intervention as an event against the asset.

Build a baseline

A baseline is the normal range of each indicator in each operating state of the asset.

  1. List the operating states. For example: off, standby, duty at low, medium and high load, and a cleaning cycle.
  2. Record a state signal. Use a run signal, a speed reference, a flow or a production count, so that each reading belongs to one state. A ZIO reads a 4 to 20 mA flow meter, and a ZMB reads a drive or a flow computer over Modbus.
  3. Collect normal data. Collect at least two weeks of each state and at least 50 run cycles for an asset that starts and stops. A weather-dependent load needs a heating season and a cooling season.
  4. Record events. Note maintenance, repairs, set-point changes and process changes with their dates. Each one starts a new baseline.
  5. Calculate the normal range. Use the median and the 5th and 95th percentiles of each indicator in each state.

Keep the baseline with the asset record. When somebody asks why an alert fired, the baseline is the evidence.

From baseline to alert

Use four types of rule, in this order:

  1. Fixed limits for conditions that are never normal: voltage unbalance above 1%, current above the nameplate full-load current, or a motor that runs outside its schedule.
  2. Baseline limits for each state, as in the pump example: power above the 95th percentile plus a margin for a set time. Start with a 3 to 5% margin and 30 to 60 minutes, then tune.
  3. Rate of change for slow degradation: specific energy that rises for three or more consecutive weeks.
  4. Models for assets where normal power depends on several variables together (see the next section).

Give each rule an owner and a defined action. For each asset, record the number of alerts each month and how many led to a work order. If a rule gives three false alarms in a row, widen the margin or lengthen the hold time, or remove the rule.

Where machine learning helps

A fixed limit fails when normal power depends on several variables at once. A chiller is the usual case. Its power depends on the cooling load, the condenser water or outdoor air temperature and the chilled water set point. A limit that is correct in July is wrong in January.

For these assets, fit a regression of power against the variables that drive it, from the baseline period. Alert on the residual, the measured power minus the predicted power, with the same limit-and-hold rule as above. For example, alert when the residual is more than three standard deviations above zero for 30 minutes. A linear regression with a few terms is usually enough, and an engineer can read its coefficients and check them against the plant.

Faults are rare, so there are few labelled examples to train a more complex model. The NIST AI Risk Management Framework asks for documented validation and for monitoring of a model's performance in use. For a maintenance model, keep a hold-out period from the baseline to test the model, log every alert with its outcome, and fit the model again after each recorded event.

Other measurements

Electrical data does not replace the measurement that responds first to a failure mode. Vibration responds first to bearing defects, imbalance, misalignment and looseness. Temperature responds to bearing, winding and connection faults. Oil analysis covers gearboxes, hydraulics and transformers. Flow, pressure and differential pressure turn power into efficiency.

The most useful pairing on a critical motor is power with a probe temperature on the drive-end bearing housing. A TES-2X probe temperature sensor on the housing reports over the same Zigbee network as the ZEM. Compare the bearing temperature rise above ambient at the same load. A power rise with a normal bearing temperature points to the driven load or the process. A power rise with a bearing temperature rise points to the bearing or the alignment.

Start with a pilot

  1. Choose three to five critical assets from the criticality ranking, each with a known failure mode.
  2. Fit a ZEM on each motor circuit and bring in one process variable that describes the output, such as flow, pressure or production count.
  3. Collect the baseline and record every event.
  4. Configure fixed limits first, then baseline limits.
  5. Review every alert for three months. Keep the rules that found real problems and remove the rules that did not.
  6. Extend to more assets with the rules that worked.

Define success before the pilot starts: for example, one fault found before it caused an unplanned stop, or a verified efficiency loss that pays for the repair.

Common questions

Can an energy monitor detect bearing faults?

Not reliably. ISO 20958 lists rolling-element bearing defects among the faults that electrical signature analysis can find, but the bearing components in the motor current are small and need waveform capture at kilohertz rates. An accelerometer on the bearing housing finds them earlier. Interval power data shows only late secondary effects, such as more power at the same output, and a probe temperature sensor on the housing shows the heat.

How much normal data do I need for a baseline?

At least two weeks of each operating state, and at least 50 run cycles for an asset that starts and stops. A weather-dependent load, such as an air handling unit fan, needs a heating season and a cooling season. Record the date of every repair or set-point change, because each one starts a new baseline.

What is the difference between predictive and condition-based maintenance?

Condition-based maintenance acts when a measured condition crosses a limit. Predictive maintenance also extrapolates the trend to estimate when the limit will be reached. Both use the same data and the same baseline. The estimate is only as good as the evidence that the trend continues at the same rate.