A demand response rule starts a chiller pre-cool at 16:00 and stops it at 16:30. The gateway restarts at 16:10. If the stop existed only in memory, the chiller runs until somebody notices. Moving control to a local gateway takes the internet out of that chain, but the sensors, the site network, the equipment and the gateway's own restarts are still in it. A failure matrix lists each of these faults with its expected response, and every row becomes a test.
Draw the chain of authority
Draw the path from a request to the equipment. A typical chain is the aggregator's platform, the gateway's automation, the BMS, the equipment's own controller and a local hand switch. For each layer, mark whether it may request a change, whether it may block a change, and what it does when it loses contact with the layer above.
Give each writable point one owner at a time. When a DR request, a BMS schedule and an operator can all write the same set point, write down the order. A common order is: the local hand switch beats everything, a BMS operating limit beats the DR request, and the DR request beats the normal schedule. BACnet enforces an order like this with the 16-level priority array on commandable objects. Modbus has no equivalent, so the last write wins and neither writer knows. The BACnet and Modbus comparison and the priority array guide cover both.
Put each function in the layer that survives the failure of the layers above it. The schedule for a DR event can live in the gateway. A maximum-demand limit or a minimum supply temperature belongs in the BMS or the drive, so it still works when the gateway is off. Protection belongs in the relay, the drive's trip settings or a safety controller.
Then decide what the equipment does when the gateway stays off. With a relay output, wire the load to the contact whose de-energised position gives the state you want. With a Modbus or BACnet write, the equipment needs its own communication-loss timer. ABB's ACS580-family drives, for example, can take no action, trip, hold the last speed or run at a preset safe speed once fieldbus messages stop for longer than a set time. Set that time longer than the longest normal gap between the gateway's writes, or the drive trips in normal service. On the drive's embedded Modbus interface, also check which messages reset the timer. If any message does, a gateway that still polls while its control rule has stopped keeps the drive satisfied. Where the equipment has no such timer, the gateway can increment a heartbeat register every cycle, and the BMS can take over when the value stops changing.
The failure matrix
Write the expected response for each row before the test. The right response depends on the load. Suppose the gateway rule is what lowers a chiller's leaving-water set point. If the rule holds that set point on a stale supply temperature, the chiller's own freeze protection is the only remaining limit. Holding a lighting circuit's last state usually costs only energy, unless the circuit serves an occupied space that must not go dark.
The owner and time limit columns below are example values for the chiller site, with 60 s reporting and 60 s rule evaluation. Replace them with your own.
| Condition | Detected by | Response | Owner | Time limit | Return to normal |
|---|---|---|---|---|---|
| Link to the platform lost | No successful exchange for 3 heartbeat intervals (180 s at 60 s) | Local rules continue. Events already started run to their saved end. No new remote starts. | Site engineer | Next evaluation | 3 consecutive good exchanges |
| A measurement that a rule uses goes stale | Point age above 3 reporting intervals, or a bad quality flag | Block new starts that depend on it. End commands already owed still run. | Site engineer | Next evaluation | 2 consecutive fresh readings |
| A command times out | No response within the protocol timeout | Treat the state as unknown. Read it back before any retry. Resend only the same absolute value. | Controls contractor | Read-back within 1 poll cycle | Read-back shows the requested state |
| The gateway restarts during a cycle | Saved cycle state at start-up | Run owed end commands. Hold anything past its recovery window for review. | Site engineer | End command within 5 min of its deadline | End state confirmed by device feedback |
| The gateway stays off | The equipment's communication-loss timer or the BMS heartbeat check | The equipment takes its chosen fail state | Equipment supplier | The communication-loss time | Gateway writing again, and any drive fault reset |
| An equipment interlock refuses an action | Mode, limit or timer status from the equipment | Show the block and its source. Do not retry into it. | Equipment supplier | Next evaluation | Interlock clear, from the next cycle |
| A second system writes the same point | Read-back differs from the last value written | Stop writing. Report the conflict. | Controls contractor | Next poll | One owner agreed for the point |
| The configuration changes during a cycle | Configuration version | The version that started the cycle also ends it | Whoever edits the rule | Not time-bound | Cycle complete |
| Time sync lost, or the clock steps | Sync status and size of the step | Hold time-based starts while the offset exceeds half the evaluation interval (30 s here) | Site engineer | Next evaluation | Sync restored |
| Daylight saving change | Timezone-aware scheduling, not clock status | A written rule for the skipped hour and the repeated hour | Whoever writes the schedule | Not time-bound | Next ordinary day |
A row that says only "raise an alarm" is incomplete if the action must also be stopped or handed over.
Here is the stale row worked for the chiller above. The supply temperature reports every 60 s. The point is stale after 180 s without a report. The rule then refuses new DR starts on that chiller within one evaluation, so within 60 s. The 16:30 end command still runs. Normal operation resumes after 2 consecutive fresh readings. The stale data guide defines the data states that a rule must separate.
The conflicting-writer row is harder, because nothing fails. The gateway writes 6 °C to the chiller's leaving-water set point register over Modbus at 16:00. At 16:15 the BMS schedule rewrites its normal 7 °C. The gateway's next read-back, at 16:16, shows 7.0 °C. If the gateway writes 6 °C again, the two systems take turns every 15 minutes and the log shows only successful writes. The matrix response is to stop writing, report the point with both values, and let the site agree an owner. Usually the BMS suspends its schedule write while a DR event flag is set.
The interlock row often catches minimum run and off times. Copeland recommends at least 3 minutes from start to stop for its scroll compressors, so that oil returns to the sump. A controller that enforces this will defer a DR stop sent 2 minutes after the BMS started the compressor. Many chiller controllers also add an anti-recycle delay between starts. Read both values from the controller's parameter list and put them in the matrix, so that the rule expects the refusal and does not report a fault.
The daylight saving row is not a clock fault. In Ireland and the UK, local time jumps from 01:00 to 02:00 on the last Sunday in March and repeats 01:00 to 02:00 on the last Sunday in October. A daily rule at 01:30 local time has no start in March and two possible starts in October. Store schedules in UTC, or define what happens in both cases.
Command design
Send absolute values: "set 21 °C" and "off", not "raise by 2 °C" or "toggle". A repeated absolute command is harmless. A late one is not. An "on" that arrives after the event window has closed still turns the load on. Give each command an expiry that matches its purpose. A DR start is valid until its event window closes. An owed end command is valid for a short recovery window after its deadline. After that, a person checks the equipment before anything else is sent. The end command itself is saved before the start is sent, so it survives a restart and a change to the rule.
A timeout means the state is unknown. The command may have run while its response was lost. A normal response to a Modbus write (function code 05, 06, 15 or 16) proves that the device accepted the frame. It does not prove that the equipment moved. Read the state back, as the BESS command verification guide shows. The Zigbee On/Off cluster has a Toggle command as well as On and Off. Use On and Off, and treat the command as done only when a fresh report of the OnOff attribute shows the new state. Telecontrol protocols make the stages explicit, as the IEC 60870 command guide shows.
Test a complete cycle
Test in a simulator or on a non-critical load before live equipment. Use a rule that switches a load on at minute 0 and off at minute 10. For each step, record the time and value of every command, its feedback and the observed equipment state. Mark each step pass or fail against the time limit in its matrix row.
- Run the normal cycle. Pass: one start and one end command, each confirmed by feedback.
- Restart the gateway at minute 4. Pass: the end command runs at minute 10, on the same equipment, and the start is not sent again.
- Power the gateway off at minute 8 and back on at minute 11. Pass: the owed end command runs as soon as the gateway is back. A second "off" after a crash is acceptable. A second start is a failure.
- Keep the gateway off from minute 8 to minute 20. Pass: the end command is not replayed at minute 20. The rule is held and the operator is asked to check the equipment, as the restart row requires. If the load has a communication-loss timer shorter than 12 minutes, the record also shows the load reaching its fail state.
- Withhold the end feedback. Pass: the operator sees an unresolved result, and the next start is blocked until someone confirms the equipment state.
- Edit or delete the rule at minute 5. Pass: the cycle already started ends at minute 10 as planned.
- Stop the input measurement at minute 2. Pass: new starts are blocked within the stale limit plus one evaluation.
- In the simulator, set the site timezone and run a daily rule across both daylight saving changes. Pass: the behaviour matches the rule you wrote for the skipped and repeated hour.
Keep the records. They are the evidence that each row was tested, and they are the baseline for a repeat test after a firmware or configuration change.
Local control in Edge
Edge runs automation on the Gateway, triggered by data, device status, data quality or a schedule. Rule evaluation ignores data older than a configurable limit, 5 minutes by default. Set it to the stale limit in your matrix.
Rules that start an action and end it after a set time save the exact end commands before the start is sent. If that save fails, the start is not sent. Editing or deleting the rule does not cancel an end command already owed. After a restart, Edge recovers owed end commands for up to five minutes after their deadline. Later than that, it does not replay the command. It blocks automation and asks for the equipment to be checked, and turning automation back on does not clear that block. End actions must set an explicit state, so Edge refuses a toggle or a placeholder value as an end action. Turning automation off stops new actions only. Owed end commands still run. It is not an emergency stop.
For Zigbee controls, Edge treats a command as done only when a fresh device report shows the requested value. By default it waits 5 s for that report and retries, up to three attempts, then raises a critical alert. An unconfirmed end action blocks the same rule's next start until a fresh report shows the end state. That report proves the device's logical state. Use a power measurement to prove that the load actually stopped. Edge does not arbitrate between two rules that write the same actuator, so give each actuator one rule.
Daily on/off rules use the Gateway's configured timezone. If a daylight saving change removes the end time, Edge does not send the start. A repeated hour in October does not start a daily cycle a second time. Plain schedule triggers behave differently: a time skipped in March does not fire, and a time repeated in October fires twice unless the rule is rate-limited.
In the test above, step 3 falls inside Edge's five-minute recovery window. Step 4 falls outside it, so the pass condition is a blocked recovery and a request to check the equipment.
Common questions
What is a failure matrix for control?
A table with one row for each thing that can fail: a link, a measurement, a command, a restart, a clock. Each row says how the fault is detected, what the control does, who owns that response, how quickly it must happen and what must be true before normal operation resumes. Each row is also one commissioning test.
How old can a measurement be before a rule stops using it?
Tie the limit to the reporting interval. Three missed reports is a reasonable starting point: 180 s for a point that reports every 60 s. A shorter limit raises false alarms from one lost message. A much longer limit lets a rule act on a value that no longer describes the equipment.
Why send absolute values instead of changes?
A command can arrive twice, for example after a retry. "Set point 21 °C" sent twice leaves the set point at 21 °C. "Raise by 2 °C" sent twice raises it by 4 °C. The same applies to "off" compared with "toggle".