Protocols and data

Industrial polling, backpressure and retries

How to budget polling on an RS-485 bus or a Modbus TCP gateway: transactions per cycle, the cost of an offline device, overdue reads and safe retries.

A Modbus RTU master sends one request, waits for the reply or the timeout, and only then sends the next, so an RS-485 bus has a ceiling that no setting raises. A Modbus TCP gateway in front of the bus adds no capacity. If the poller asks for more, the queue of waiting requests grows, and the value on the dashboard falls behind the clock. Count the transactions before you set the polling intervals.

This guide is part of the data quality and delivery series.

Count transactions at the shared connection

Count at the narrowest point. Ten Modbus TCP clients that read through one serial gateway share its single RS-485 port, however fast the Ethernet side is. The gateway queues their requests and sends them to the bus one at a time.

A worked example on one RS-485 bus at 9600 bit/s, 8 data bits, even parity and 1 stop bit:

QuantityValue
Devices8 meters
Requests per device per cycle5 blocks of 20 registers
One transaction: 8-byte request, 20 ms device delay, 45-byte response, 3.5-character gap85 ms
One full cycle8 × 5 × 85 ms = 3.4 s

This bus refreshes every point about every 3.4 s, and no faster. A 1 s polling interval asks for 3.4 times its capacity. The Modbus RTU timing calculator below uses the same inputs. It gives 85 ms per exchange and 0.68 s for one request to each of the eight meters. Five blocks per meter make that 3.4 s.

With 11 bits per character, one character takes 1.15 ms at 9600 bit/s, and the 3.5-character frame gap takes 4.0 ms. Above 19,200 bit/s the serial line guide fixes the frame gap at 1.75 ms and the inter-character timeout at 0.75 ms (section 2.5.1.1). A faster baud rate shortens the wire time but not the device's response delay. At 38,400 bit/s the same exchange takes about 37 ms, and 20 ms of that is the meter.

Read larger blocks. Each transaction pays for a request, a device delay and a frame gap, however many registers it returns. At 9600 bit/s with a 20 ms device delay, 20 reads of 2 registers take about 870 ms. One read of the same 40 consecutive registers takes about 130 ms. Function codes 03 and 04 read up to 125 registers in one request (Modbus application protocol V1.1b3, sections 6.3 and 6.4). A block that includes an address the device does not implement fails with exception 02, illegal data address. Merge blocks only across ranges that the register map lists as readable, and check the map for a lower per-request limit.

BACnet MS/TP has the same limit in a different form. A controller transmits only while it holds the token, and a router sends at most Max_Info_Frames requests each time it holds it. Edge polls BACnet over BACnet/IP, so it reaches MS/TP controllers through a router, and every read waits for that router's turn. The BACnet/IP vs MS/TP guide works through the token rotation time.

The cost of an offline device

Each request to a device that does not answer waits for the full response timeout. In the example, with a 1 s timeout and five requests per meter, one offline meter adds 5 s to every cycle. The cycle grows from 3.4 s to 8.0 s, and the seven healthy meters are read less than half as often.

The serial line guide gives 1 s to several seconds at 9600 bit/s as a typical response timeout (section 2.4.1). That is safe for a slow device and expensive on a busy bus. Set the timeout from measurement instead. Record the slowest reply from each device over a day of normal operation, and set the timeout to about twice that value. A meter whose slowest reply is 150 ms gets a 300 ms timeout. The offline penalty in the example then falls from 5 s to 1.5 s.

Then back off per device. For example, after three consecutive timeouts, take the device off its normal schedule and mark its points stale. Send it one probe request every 60 s. When a probe succeeds, restore the schedule. The healthy meters return to their 3.4 s cycle, and the offline meter costs 0.3 s a minute.

The bus still needs headroom for the timeouts it sees before the backoff starts. With a 5 s interval, the example bus has 1.6 s to spare. One offline meter with a 300 ms timeout costs 1.5 s and fits. The same meter with a 1 s timeout costs 5 s and does not. When one device's timeouts do not fit in the spare time, split the bus.

When requests arrive faster than they complete

Poll the example bus every 1 s, and the poller generates 40 requests a second. The bus completes about 12. The queue grows by about 28 requests every second. If the queue is served in order, the request sent after ten minutes was queued about seven minutes earlier. Every status indicator can stay green: the bus is busy, the devices answer and no request fails. Only the timestamps show that the data is seven minutes old.

Decide what to do with work that is overdue:

Waiting workWhat to do
Several reads of the same current valueMerge them into one; only the newest matters
A missed historical sampleLost, unless the device stores history; record the gap
A stored reading waiting to go upstreamKeep it in a durable queue, with duplicate handling
A commandCheck that it is still valid; never replay an old command automatically
A read of a point that was removed from the mapDrop it

Monitor the age of the oldest waiting request as well as the queue length. On a bus that keeps up, no request waits longer than one polling interval. Raise an alarm when the oldest one is older than two intervals.

Separate live points from energy registers

Most points do not need the fastest interval. An energy register (kWh) is a running total, so a reading every 15 minutes loses no energy: the next reading includes everything since the last one. Power, a setpoint or a breaker status used for control needs seconds.

Put the few live points on a short interval and the energy and configuration registers on a long one. Suppose each meter in the example has one live block read every 5 s and four energy blocks read every 15 minutes. Live reads then use 8 × 85 ms = 0.68 s of every 5 s. The 32 energy reads take 2.7 s once every 15 minutes. Spread them across the interval so that they do not all fall on the quarter-hour and hold up the live reads for 2.7 s.

Retry only what can succeed

FailureRetry?
Timeout or CRC error on the serial busYes, up to the per-device limit, then back off
Modbus TCP connection droppedReconnect, with a growing delay between attempts
Exception 06, server device busyYes, after a delay
Exception 0B, gateway target device failed to respondTreat it as a timeout of that device; the gateway has already waited its own timeout
Exception 0A, gateway path unavailableNot at once. The gateway is usually misconfigured or overloaded; check its unit ID routing and load
Exception 04, server device failureOnce at most, then raise an alarm; the device reports an unrecoverable error
Exceptions 01, 02 and 03: illegal function, data address or data valueNo. The request is wrong; fix the register map

Behind a Modbus TCP gateway, two timeouts are in series: the client's and the gateway's serial timeout. Set the client's timeout longer than the gateway's serial timeout plus the time a request waits in the gateway's queue. Otherwise the client gives up on a request that the gateway is still processing, retries, and puts a second copy of the same read on the bus. The Modbus TCP vs RTU guide covers gateway timeouts and connection limits.

A timeout on a write (function 06 or 16) does not show that the write failed. The device may have applied it and the reply may have been lost. Read the register back before you retry. An absolute setpoint written twice does no harm. A write that triggers an action, such as a counter reset or a start command, can act twice, so never retry it without a read-back.

Exponential backoff and jitter matter when many clients share one server. Several Modbus TCP clients behind one gateway, or a fleet of site gateways that reconnect to one server after an outage, all retry at the same moment unless each one adds a random delay. AWS describes the pattern, with a cap on the number of attempts and on the longest wait. Without it, the retries after a recovery can keep a slow connection overloaded after the fault has cleared. On one RS-485 bus there is only one master, so jitter changes nothing there. The per-device backoff does the work.

Test recovery under load

  1. Run the full point list at the planned intervals for at least one hour. Record the cycle time. It must be shorter than the fastest polling interval by at least one device's timeouts.
  2. Switch one device off. After its backoff starts, the readings from the other devices must again be less than two intervals old.
  3. Switch it back on. Its points must be fresh again within one probe period, 60 s in the example above.
  4. Interrupt the upstream link for 30 minutes. Local readings must keep their normal timestamps, and the outbound queue must grow at the expected rate: points per minute × 30.
  5. Restore the link. Every queued reading must arrive once. Count the readings per point and timestamp at the receiver, and check that the bus cycle time returns to the step 1 value within two cycles.

Polling in Edge

Edge on the ZGW-20 Gateway has five Modbus connection slots. Each slot owns one endpoint, a serial port or one TCP host and port. Each has its own queues, rate limits, latency figures and error backoff, and a 2,000 ms pause in live reads after a write. All unit IDs on one slot share these, so a failing device behind a gateway can delay the healthy devices on the same slot. Use one slot for each independent endpoint.

The error backoff counts consecutive failed requests on the slot, not on one device. After 3, Edge stops reading the slot for 5 s or two polling intervals, whichever is longer. After 6 the pause is 15 s or three intervals, and after 10 it is 30 s or five intervals. One successful reply resets the count. Healthy devices on the same slot usually answer between the offline device's requests, so the count rarely reaches 3. The offline device's timeouts then continue on every cycle, which is why the timeout value matters.

By default a slot sends up to 20 live requests and 5 scheduled, clock-aligned requests a second. Both limits can be set from 1 to 100. The live default is above what the example bus can complete, about 12 a second at 9600 bit/s, so set it below the capacity you measured. The scheduled queue holds 5,000 requests and drops the oldest first. The live queue holds 250 and merges equivalent reads. When a slot's latency exceeds 5,000 ms, Edge cuts live polling on that slot to one request a second. Edge merges contiguous reads into blocks of up to 125 registers. Each slot reports its last and average latency and its error count, and Edge logs a warning when it drops scheduled reads. A device goes offline after it misses two expected polls, and never sooner than five minutes.

These limits protect a slow bus. They do not add capacity. A ZMB Modbus interface is the RS-485 master of its own short bus. The ZMB-31 reads up to 30 registers from the equipment beside it and sends them to the Gateway over the Zigbee mesh. Several ZMBs replace one long shared bus with several short ones, each with its own timeout budget.

Common questions

How many Modbus devices can I poll per second?

Divide one second by the time of one transaction: the request and response on the wire, the device's response delay and the 3.5-character gap between frames. At 9600 bit/s, with a 20-register read and a 20 ms device delay, one transaction takes about 85 ms. One RS-485 bus then completes about 12 transactions a second, shared by every device on it.

Why does one offline Modbus device slow down the others?

The master sends one request at a time. Each request to the offline device waits for the full response timeout before the next request goes out. The offline device adds one timeout for each request sent to it in each cycle: 5 s with a 1 s timeout and five requests per cycle.

What is backpressure?

The effect of a slow stage on the stages before it. When a poller generates requests faster than the connection completes them, the queue of waiting requests grows. The poller must then slow down, merge or drop work, or the delay grows without limit.