Flexibility and grid codes

Local control: interlocks and a failure matrix

How to plan local automation on a gateway: who has authority, a failure matrix for lost links, stale data, timeouts and restarts, and the tests to run before control goes live.

A demand response rule starts a chiller pre-cool at 16:00 and stops it at 16:30. The gateway restarts at 16:10. If the stop existed only in memory, the chiller runs until somebody notices. Moving control to a local gateway takes the internet out of that chain, but the sensors, the site network, the equipment and the gateway's own restarts are still in it. A failure matrix lists each of these faults with its expected response, and every row becomes a test.

Draw the chain of authority

Draw the path from a request to the equipment. A typical chain is the aggregator's platform, the gateway's automation, the BMS, the equipment's own controller and a local hand switch. For each layer, mark whether it may request a change, whether it may block a change, and what it does when it loses contact with the layer above.

Give each writable point one owner at a time. When a DR request, a BMS schedule and an operator can all write the same set point, write down the order. A common order is: the local hand switch beats everything, a BMS operating limit beats the DR request, and the DR request beats the normal schedule. BACnet enforces an order like this with the 16-level priority array on commandable objects. Modbus has no equivalent, so the last write wins and neither writer knows. The BACnet and Modbus comparison and the priority array guide cover both.

Put each function in the layer that survives the failure of the layers above it. The schedule for a DR event can live in the gateway. A maximum-demand limit or a minimum supply temperature belongs in the BMS or the drive, so it still works when the gateway is off. Protection belongs in the relay, the drive's trip settings or a safety controller.

Then decide what the equipment does when the gateway stays off. With a relay output, wire the load to the contact whose de-energised position gives the state you want. With a Modbus or BACnet write, the equipment needs its own communication-loss timer. ABB's ACS580-family drives, for example, can take no action, trip, hold the last speed or run at a preset safe speed once fieldbus messages stop for longer than a set time. Set that time longer than the longest normal gap between the gateway's writes, or the drive trips in normal service. On the drive's embedded Modbus interface, also check which messages reset the timer. If any message does, a gateway that still polls while its control rule has stopped keeps the drive satisfied. Where the equipment has no such timer, the gateway can increment a heartbeat register every cycle, and the BMS can take over when the value stops changing.

The failure matrix

Write the expected response for each row before the test. The right response depends on the load. Suppose the gateway rule is what lowers a chiller's leaving-water set point. If the rule holds that set point on a stale supply temperature, the chiller's own freeze protection is the only remaining limit. Holding a lighting circuit's last state usually costs only energy, unless the circuit serves an occupied space that must not go dark.

The owner and time limit columns below are example values for the chiller site, with 60 s reporting and 60 s rule evaluation. Replace them with your own.

ConditionDetected byResponseOwnerTime limitReturn to normal
Link to the platform lostNo successful exchange for 3 heartbeat intervals (180 s at 60 s)Local rules continue. Events already started run to their saved end. No new remote starts.Site engineerNext evaluation3 consecutive good exchanges
A measurement that a rule uses goes stalePoint age above 3 reporting intervals, or a bad quality flagBlock new starts that depend on it. End commands already owed still run.Site engineerNext evaluation2 consecutive fresh readings
A command times outNo response within the protocol timeoutTreat the state as unknown. Read it back before any retry. Resend only the same absolute value.Controls contractorRead-back within 1 poll cycleRead-back shows the requested state
The gateway restarts during a cycleSaved cycle state at start-upRun owed end commands. Hold anything past its recovery window for review.Site engineerEnd command within 5 min of its deadlineEnd state confirmed by device feedback
The gateway stays offThe equipment's communication-loss timer or the BMS heartbeat checkThe equipment takes its chosen fail stateEquipment supplierThe communication-loss timeGateway writing again, and any drive fault reset
An equipment interlock refuses an actionMode, limit or timer status from the equipmentShow the block and its source. Do not retry into it.Equipment supplierNext evaluationInterlock clear, from the next cycle
A second system writes the same pointRead-back differs from the last value writtenStop writing. Report the conflict.Controls contractorNext pollOne owner agreed for the point
The configuration changes during a cycleConfiguration versionThe version that started the cycle also ends itWhoever edits the ruleNot time-boundCycle complete
Time sync lost, or the clock stepsSync status and size of the stepHold time-based starts while the offset exceeds half the evaluation interval (30 s here)Site engineerNext evaluationSync restored
Daylight saving changeTimezone-aware scheduling, not clock statusA written rule for the skipped hour and the repeated hourWhoever writes the scheduleNot time-boundNext ordinary day

A row that says only "raise an alarm" is incomplete if the action must also be stopped or handed over.

Here is the stale row worked for the chiller above. The supply temperature reports every 60 s. The point is stale after 180 s without a report. The rule then refuses new DR starts on that chiller within one evaluation, so within 60 s. The 16:30 end command still runs. Normal operation resumes after 2 consecutive fresh readings. The stale data guide defines the data states that a rule must separate.

The conflicting-writer row is harder, because nothing fails. The gateway writes 6 °C to the chiller's leaving-water set point register over Modbus at 16:00. At 16:15 the BMS schedule rewrites its normal 7 °C. The gateway's next read-back, at 16:16, shows 7.0 °C. If the gateway writes 6 °C again, the two systems take turns every 15 minutes and the log shows only successful writes. The matrix response is to stop writing, report the point with both values, and let the site agree an owner. Usually the BMS suspends its schedule write while a DR event flag is set.

The interlock row often catches minimum run and off times. Copeland recommends at least 3 minutes from start to stop for its scroll compressors, so that oil returns to the sump. A controller that enforces this will defer a DR stop sent 2 minutes after the BMS started the compressor. Many chiller controllers also add an anti-recycle delay between starts. Read both values from the controller's parameter list and put them in the matrix, so that the rule expects the refusal and does not report a fault.

The daylight saving row is not a clock fault. In Ireland and the UK, local time jumps from 01:00 to 02:00 on the last Sunday in March and repeats 01:00 to 02:00 on the last Sunday in October. A daily rule at 01:30 local time has no start in March and two possible starts in October. Store schedules in UTC, or define what happens in both cases.

Command design

Send absolute values: "set 21 °C" and "off", not "raise by 2 °C" or "toggle". A repeated absolute command is harmless. A late one is not. An "on" that arrives after the event window has closed still turns the load on. Give each command an expiry that matches its purpose. A DR start is valid until its event window closes. An owed end command is valid for a short recovery window after its deadline. After that, a person checks the equipment before anything else is sent. The end command itself is saved before the start is sent, so it survives a restart and a change to the rule.

A timeout means the state is unknown. The command may have run while its response was lost. A normal response to a Modbus write (function code 05, 06, 15 or 16) proves that the device accepted the frame. It does not prove that the equipment moved. Read the state back, as the BESS command verification guide shows. The Zigbee On/Off cluster has a Toggle command as well as On and Off. Use On and Off, and treat the command as done only when a fresh report of the OnOff attribute shows the new state. Telecontrol protocols make the stages explicit, as the IEC 60870 command guide shows.

Test a complete cycle

Test in a simulator or on a non-critical load before live equipment. Use a rule that switches a load on at minute 0 and off at minute 10. For each step, record the time and value of every command, its feedback and the observed equipment state. Mark each step pass or fail against the time limit in its matrix row.

  1. Run the normal cycle. Pass: one start and one end command, each confirmed by feedback.
  2. Restart the gateway at minute 4. Pass: the end command runs at minute 10, on the same equipment, and the start is not sent again.
  3. Power the gateway off at minute 8 and back on at minute 11. Pass: the owed end command runs as soon as the gateway is back. A second "off" after a crash is acceptable. A second start is a failure.
  4. Keep the gateway off from minute 8 to minute 20. Pass: the end command is not replayed at minute 20. The rule is held and the operator is asked to check the equipment, as the restart row requires. If the load has a communication-loss timer shorter than 12 minutes, the record also shows the load reaching its fail state.
  5. Withhold the end feedback. Pass: the operator sees an unresolved result, and the next start is blocked until someone confirms the equipment state.
  6. Edit or delete the rule at minute 5. Pass: the cycle already started ends at minute 10 as planned.
  7. Stop the input measurement at minute 2. Pass: new starts are blocked within the stale limit plus one evaluation.
  8. In the simulator, set the site timezone and run a daily rule across both daylight saving changes. Pass: the behaviour matches the rule you wrote for the skipped and repeated hour.

Keep the records. They are the evidence that each row was tested, and they are the baseline for a repeat test after a firmware or configuration change.

Local control in Edge

Edge runs automation on the Gateway, triggered by data, device status, data quality or a schedule. Rule evaluation ignores data older than a configurable limit, 5 minutes by default. Set it to the stale limit in your matrix.

Rules that start an action and end it after a set time save the exact end commands before the start is sent. If that save fails, the start is not sent. Editing or deleting the rule does not cancel an end command already owed. After a restart, Edge recovers owed end commands for up to five minutes after their deadline. Later than that, it does not replay the command. It blocks automation and asks for the equipment to be checked, and turning automation back on does not clear that block. End actions must set an explicit state, so Edge refuses a toggle or a placeholder value as an end action. Turning automation off stops new actions only. Owed end commands still run. It is not an emergency stop.

For Zigbee controls, Edge treats a command as done only when a fresh device report shows the requested value. By default it waits 5 s for that report and retries, up to three attempts, then raises a critical alert. An unconfirmed end action blocks the same rule's next start until a fresh report shows the end state. That report proves the device's logical state. Use a power measurement to prove that the load actually stopped. Edge does not arbitrate between two rules that write the same actuator, so give each actuator one rule.

Daily on/off rules use the Gateway's configured timezone. If a daylight saving change removes the end time, Edge does not send the start. A repeated hour in October does not start a daily cycle a second time. Plain schedule triggers behave differently: a time skipped in March does not fire, and a time repeated in October fires twice unless the rule is rate-limited.

In the test above, step 3 falls inside Edge's five-minute recovery window. Step 4 falls outside it, so the pass condition is a blocked recovery and a request to check the equipment.

Common questions

What is a failure matrix for control?

A table with one row for each thing that can fail: a link, a measurement, a command, a restart, a clock. Each row says how the fault is detected, what the control does, who owns that response, how quickly it must happen and what must be true before normal operation resumes. Each row is also one commissioning test.

How old can a measurement be before a rule stops using it?

Tie the limit to the reporting interval. Three missed reports is a reasonable starting point: 180 s for a point that reports every 60 s. A shorter limit raises false alarms from one lost message. A much longer limit lets a rule act on a value that no longer describes the equipment.

Why send absolute values instead of changes?

A command can arrive twice, for example after a retry. "Set point 21 °C" sent twice leaves the set point at 21 °C. "Raise by 2 °C" sent twice raises it by 4 °C. The same applies to "off" compared with "toggle".