Protocols and data

Store-and-forward buffer sizing for edge gateways

How to size a store-and-forward buffer on an edge gateway: records per second, outage length, storage per record, drain time after recovery, retention limits and full-queue behaviour.

Store-and-forward keeps readings on a gateway while the link to the destination is down, and sends them when the link returns. It turns an outage into a delay instead of a loss, but only up to the size of the store. Size it from the data rate and the longest outage you must survive, then check how long recovery takes.

This guide is part of the data quality and delivery series.

Which store protects the data?

A gateway can hold data in several places, and they protect different things:

StoreSurvives an upstream outageSurvives a gateway restart
Batch in memory, waiting to sendFor a short timeNo
Local history databaseYes; you can view it on siteYes
Durable retry queue on diskYes, and it is sent on recoveryYes

Only a durable queue on disk, or a process that sends from local history, carries data through a long outage and a restart. A batch in memory is lost if the power fails before it is sent. Ask which store the gateway uses, and when data moves from one to the next.

Size the buffer

Use this worked example: a site with 1,200 readings a minute, which must survive a 12-hour outage. Assume 200 bytes per stored record; measure the real figure on the gateway, including indexes and overhead.

QuantityCalculationResult
Incoming rate1,200 ÷ 6020 records per second
Outage length12 × 3,60043,200 seconds
Records to hold20 × 43,200864,000 records
Storage864,000 × 200 bytes172.8 MB, before overhead

Add a margin for overhead and for a longer outage than planned. Check that the storage is not shared with anything that can fill it, such as logs.

Check the drain time

Recovery is not instant. The destination accepts records at a limited rate, and new readings keep arriving. The backlog shrinks at the net drain rate: the upload rate minus the incoming rate.

QuantityCalculationResult
Upload rate accepted by the destinationMeasured80 records per second
New readingsFrom above20 records per second
Net drain rate80 − 2060 records per second
Time to clear the backlog864,000 ÷ 6014,400 seconds, or 4 hours

During those four hours, the destination shows data that is hours old. Dashboards and alarms at the destination must use the measurement time of each reading, not its arrival time. The timestamps guide explains why.

Time limits as well as size

Some stores delete data after a time, whatever its size. Some destinations refuse data older than a limit, or accept it without recalculating totals and alarms. Check both:

  • the retention time of the gateway's store;
  • the oldest reading the destination accepts, and what it does with it.

When the buffer is full

Decide the behaviour before it happens:

PolicyResultSuits
Drop the oldestRecent data survives; the start of the outage is lostLive operations
Drop the newestThe complete early record survives; the end is lostBilling and audit trails, with an alarm
Stop collectingNothing is overwritten; new data is lostRarely a good choice

Alarm well before the buffer is full, for example at 50% and 80%, so that someone can act.

Duplicates on recovery

A gateway may send a batch, lose the connection before it records the acknowledgement, and send the batch again after recovery. The destination then receives some records twice. Give each record an identity, such as the point identifier plus the measurement time, so that the destination keeps one copy. The MQTT QoS guide explains at-least-once delivery.

Test the complete outage

  1. Disconnect the upstream link for the planned outage length, or a scaled test with a known record count.
  2. Watch the queue grow and check the alarms.
  3. Restore the link and measure the drain time.
  4. At the destination, count the records for the outage period, and check for gaps, duplicates and timestamps.
  5. Repeat with a gateway restart during the outage.

Store-and-forward with Edge

Edge on the ZGW-20 Gateway keeps a local history of every reading, for dashboards and analysis on site. When a destination fails, Edge moves the undelivered data to a queue on disk, checks the destination every minute and replays the data when it returns, with the original timestamps. A batch that is still in memory at the moment of a failure can be lost, and a replay can deliver a batch twice, so apply the duplicate and outage tests above to each destination.

Common questions

What is store and forward?

A method where a gateway keeps data in local storage while its connection to the destination is down, and sends it when the connection returns. It turns a network outage into a delay instead of a loss, up to the size of the store.

How big should a store-and-forward buffer be?

Records per second, times the longest outage in seconds, times the storage per record, plus a margin. For 1,200 readings a minute and a 12-hour outage at 200 bytes each, that is 864,000 records and about 173 MB before overhead.

How long does a buffer take to empty after an outage?

Backlog divided by the net drain rate, where the net drain rate is the upload rate minus the rate of new readings. A backlog of 864,000 records, uploaded at 80 per second while 20 per second still arrive, takes four hours.