Store-and-forward keeps readings on a gateway while the link to the destination is down, and sends them when the link returns. It turns an outage into a delay instead of a loss, but only up to the size of the store. Size it from the data rate and the longest outage you must survive, then check how long recovery takes.
This guide is part of the data quality and delivery series.
Which store protects the data?
A gateway can hold data in several places, and they protect different things:
| Store | Survives an upstream outage | Survives a gateway restart |
|---|---|---|
| Batch in memory, waiting to send | For a short time | No |
| Local history database | Yes; you can view it on site | Yes |
| Durable retry queue on disk | Yes, and it is sent on recovery | Yes |
Only a durable queue on disk, or a process that sends from local history, carries data through a long outage and a restart. A batch in memory is lost if the power fails before it is sent. Ask which store the gateway uses, and when data moves from one to the next.
Size the buffer
Use this worked example: a site with 1,200 readings a minute, which must survive a 12-hour outage. Assume 200 bytes per stored record; measure the real figure on the gateway, including indexes and overhead.
| Quantity | Calculation | Result |
|---|---|---|
| Incoming rate | 1,200 ÷ 60 | 20 records per second |
| Outage length | 12 × 3,600 | 43,200 seconds |
| Records to hold | 20 × 43,200 | 864,000 records |
| Storage | 864,000 × 200 bytes | 172.8 MB, before overhead |
Add a margin for overhead and for a longer outage than planned. Check that the storage is not shared with anything that can fill it, such as logs.
Check the drain time
Recovery is not instant. The destination accepts records at a limited rate, and new readings keep arriving. The backlog shrinks at the net drain rate: the upload rate minus the incoming rate.
| Quantity | Calculation | Result |
|---|---|---|
| Upload rate accepted by the destination | Measured | 80 records per second |
| New readings | From above | 20 records per second |
| Net drain rate | 80 − 20 | 60 records per second |
| Time to clear the backlog | 864,000 ÷ 60 | 14,400 seconds, or 4 hours |
During those four hours, the destination shows data that is hours old. Dashboards and alarms at the destination must use the measurement time of each reading, not its arrival time. The timestamps guide explains why.
Time limits as well as size
Some stores delete data after a time, whatever its size. Some destinations refuse data older than a limit, or accept it without recalculating totals and alarms. Check both:
- the retention time of the gateway's store;
- the oldest reading the destination accepts, and what it does with it.
When the buffer is full
Decide the behaviour before it happens:
| Policy | Result | Suits |
|---|---|---|
| Drop the oldest | Recent data survives; the start of the outage is lost | Live operations |
| Drop the newest | The complete early record survives; the end is lost | Billing and audit trails, with an alarm |
| Stop collecting | Nothing is overwritten; new data is lost | Rarely a good choice |
Alarm well before the buffer is full, for example at 50% and 80%, so that someone can act.
Duplicates on recovery
A gateway may send a batch, lose the connection before it records the acknowledgement, and send the batch again after recovery. The destination then receives some records twice. Give each record an identity, such as the point identifier plus the measurement time, so that the destination keeps one copy. The MQTT QoS guide explains at-least-once delivery.
Test the complete outage
- Disconnect the upstream link for the planned outage length, or a scaled test with a known record count.
- Watch the queue grow and check the alarms.
- Restore the link and measure the drain time.
- At the destination, count the records for the outage period, and check for gaps, duplicates and timestamps.
- Repeat with a gateway restart during the outage.
Store-and-forward with Edge
Edge on the ZGW-20 Gateway keeps a local history of every reading, for dashboards and analysis on site. When a destination fails, Edge moves the undelivered data to a queue on disk, checks the destination every minute and replays the data when it returns, with the original timestamps. A batch that is still in memory at the moment of a failure can be lost, and a replay can deliver a batch twice, so apply the duplicate and outage tests above to each destination.
Common questions
What is store and forward?
A method where a gateway keeps data in local storage while its connection to the destination is down, and sends it when the connection returns. It turns a network outage into a delay instead of a loss, up to the size of the store.
How big should a store-and-forward buffer be?
Records per second, times the longest outage in seconds, times the storage per record, plus a margin. For 1,200 readings a minute and a 12-hour outage at 200 bytes each, that is 864,000 records and about 173 MB before overhead.
How long does a buffer take to empty after an outage?
Backlog divided by the net drain rate, where the net drain rate is the upload rate minus the rate of new readings. A backlog of 864,000 records, uploaded at 80 per second while 20 per second still arrive, takes four hours.