Protocols and data

IoT data storage: architecture and options

Compare IoT databases, edge history, replay queues and cloud archives. Size storage, define retention and verify delivery and recovery before deployment.

A local trend chart, a queue of readings waiting for a cloud platform and a seven-year archive are three different stores. Each answers different questions, costs a different amount per reading and fails in a different way. Size, retain and test each store separately.

For timestamped measurements, start with a time-series database. Use a relational database when readings must join to business records, a document database when payload shapes vary, and object storage for files and analytical datasets. Any of these can run on site or centrally.

History, delivery and recovery

Start with a sensor-data contract: source identity, value, unit, measurement time, receipt time, quality and processing provenance. Storage must keep enough of that contract to interpret a reading after the original device or configuration has changed.

StoreJobTypical implementationWhat removes a record
Queryable historyTime-range queries, charts and investigationTime-series database or partitioned SQL tableAge-based retention
Delivery outboxHold readings a destination has not confirmed, and retry themSQLite table or broker queue, one set of rows per destinationAcknowledgement from that destination
Asset and configuration storeDevice identity, channel mapping, scaling and unitsRelational database or versioned configuration filesA deliberate change, kept in version history
ArchiveKeep selected records for yearsParquet or CSV files in object storageA lifecycle rule or legal retention period
BackupRestore a system after loss or damageSnapshots copied off the hostThe backup retention schedule

The last column explains most surprises. History can expire readings that a destination never received. A queue can be empty because delivery completed, because no data arrived, or because a filter removed the point. A chart can show current values while a cloud platform is missing a week. Check each outcome separately.

Compare IoT database and storage options

Time-series databases

Use a time-series database when most of the workload is timestamped numeric readings, time-window aggregation and trend queries. Before you choose one, test late-arriving data, duplicate handling, label cardinality, query limits and retention settings on the edition you will run.

Single-node VictoriaMetrics sets age-based retention with -retentionPeriod, with a default of one month. Its sizing guide states that a sample needs about 1 byte or less on disk after compression, and that high series churn reduces compression. Calculate your own figure from vm_data_size_bytes and the stored row count. Keep at least 20% of the data directory free. VictoriaMetrics needs that space to merge data parts, and queries slow down when it cannot merge.

Cardinality is the number of distinct series. In IoT systems it grows when a label changes often, for example a firmware version or a signal-strength value stored as a label. The point rate stays the same, but the index grows with every new combination. Keep changing attributes in the asset store, not in series labels.

Relational databases

Use a relational database when queries join readings to equipment, production batches, tariff periods or service records, and the team already runs SQL in production.

PostgreSQL partitioning splits a large table by time range. Dropping or detaching an old partition is much faster than a bulk DELETE, and it avoids the VACUUM work that a bulk delete causes. The TimescaleDB extension automates this: a hypertable is split into time chunks, and a retention policy drops whole chunks older than a set interval. That policy does not apply to continuous aggregates built from the hypertable, so hourly aggregates can outlive the raw readings they came from. Test index size and query plans with at least a month of representative data.

Document databases

Document stores suit device records and event payloads whose shape varies. MongoDB time-series collections store measurements in time order in a columnar format, grouped by a metaField that identifies the series. Updates are restricted: the manual limits update match expressions to the metaField. Plan how you will store corrected values before you depend on in-place updates. A flexible document shape still needs defined units, identity and timestamp meaning.

Object storage and analytical archives

Object storage holds exported measurement files, audit evidence and analytical datasets. The query engine is a separate component. Athena reads columnar formats such as Parquet and ORC, and reads only the columns a query needs.

Two rules from the Athena tuning guide matter most for sensor archives. First, partition for the common query: if analysts look at days, do not partition by hour. Sort records by timestamp inside each file instead. Second, avoid many small files. The default Parquet row group is 128 MB, and for small files the columnar overhead outweighs the benefit. One site at 1,440,000 readings a day produces a daily Parquet file far smaller than that. Partition by site and month, or write several sites into one daily file.

Write a schema version and a data contract beside the files. A directory of CSV files without them is cheap to write and expensive to interpret years later.

Check the storage class. S3 Glacier Flexible Retrieval and Glacier Deep Archive need a restore request before you can read an object; Glacier Instant Retrieval does not. Include retrieval, request, query and transfer charges in the cost model, as well as the price per gigabyte stored.

On-site and central storage

A standalone site can run on local history alone. Add a central store when teams need portfolio comparisons, shared reporting or retention longer than the site can hold. A hybrid design gives you both, but adds a second copy to reconcile, an outbox to monitor and a second retention policy.

RequirementStoreTest
Investigate recent site events during a WAN outageQueryable local historyDisconnect the WAN. Retrieve the required interval as an authorised local user.
Compare many sitesCentral analytical store with site identityCompare units, clock offsets, intervals and asset mappings for two sites.
Bridge a connection outageOutbox with enough local diskBlock the destination. Measure queue growth, then measure drain time after you unblock it.
Retain evidence for yearsRaw or aggregated archive with an export pathGive a year-old file to someone outside the project and ask them to interpret it.
Recover a failed Gateway or serverOff-host backups and a restore procedureRestore to spare hardware. Record the time taken and the data gap.

Hot, warm and cold tiers describe how often data is read and how fast it must return. One database with two retention policies can provide two tiers. Each extra tier needs an owner, a transfer check and a defined action when the transfer fails.

How Edge stores history and replays data

EpiSensor Edge keeps queryable local history in VictoriaMetrics and pending delivery work in a separate SQLite outbox. The local-history mode records all valid readings, only readings enabled for export, or nothing. Retention is set in days and defaults to 30. VictoriaMetrics reads the value when it starts, so a change applies after an Edge restart. Turning history on starts recording from that moment; it cannot reconstruct readings that were never stored.

Healthy delivery starts in memory. Edge holds each reading in RAM until the destination reports success, then discards it without writing it to disk. It writes the reading to the outbox only when delivery fails, or when the destination is already known to be offline. Local history works the same way. Edge collects readings in memory for up to 250 ms, writes the batch to VictoriaMetrics, and uses the outbox only if that write fails. This design keeps writes to the Gateway's eMMC low. A power cut or process stop before persistence can lose in-flight readings.

A producer that posts readings to Edge over HTTP can see which case applies. A 202 response means Edge has written the work to disk for replay. A 200 means Edge accepted the readings in memory. A 503 or 507 means Edge refused disk-backed work because storage is unavailable or the low-disk guard is active. The producer must keep and resend that data.

Edge counts a live delivery as complete at a point that depends on the transport:

  • MQTT: completion of the QoS 1 publish.
  • HTTP export: the configured success status, after the whole response body arrives.
  • File export: the write into the export's pending-file queue.
  • Modbus Server: completion of every register write for the message.

Edge-owned replay applies to destinations that declare it, such as EpiSensor Core. Other integration flows run their own retry or deliver on a best-effort basis. Confirm which applies to each integration at commissioning.

The outbox follows different rules from history. Pending outbox work has no age-based expiry in the current implementation, because deleting it would discard undelivered data. A long outage grows the outbox until delivery recovers or the low-disk guard refuses new work. By default the guard triggers when free space falls to 10% of the filesystem, capped at 1 GiB, or to 256 MiB. Each destination has its own rows, so one slow destination cannot block or acknowledge another's data. Replayed readings keep their original timestamps. Replay sends only to the destination that missed the data. It does not run calculated devices or automation again, because they already processed the original reading. Outbox writes use SQLite synchronous=FULL, so a committed row survives a power cut.

Size history and outage capacity

Count reporting channels, not devices. One meter can report several measurements at different intervals. For a fixed interval:

Points per day = reporting channels × 86,400 ÷ reporting interval in seconds

For example, 1,000 channels that each report once per minute produce 1,440,000 points per day, or 43,200,000 points over 30 days, before filtering. At one-second reporting the rate is 60 times higher. Add event bursts and calculated values separately.

History. At the VictoriaMetrics figure of about 1 byte per point, 30 days of that example is roughly 43 MB of sample data, plus the index and the 20% free-space margin. A year is about 526 million points, or roughly 0.5 GB. In a synthetic EpiSensor benchmark, 200,000 regular readings used 35 KB in VictoriaMetrics and 36 MB in a plain SQLite table, a ratio of about 1,000 to 1. Noisy field values compress less than synthetic ones, so measure on the site's own data.

Outbox. On a field Gateway, each queued reading took about 550 bytes of SQLite, indexes included, for each destination. The same 1,000 channels during a 72-hour outage queue 4,320,000 readings: about 2.4 GB for one destination, or 4.8 GB for two. On a ZGW-20 with the standard 16 GB eMMC, that is a large share of the disk, and the outbox, not history, sets the longest outage the Gateway can bridge. The 64 GB eMMC compute option or the optional 128 GB SSD extends it. To find the cover you have:

Outage hours covered = free space for the outbox ÷ (bytes per queued reading × readings per hour × destinations)

Leave the low-disk guard threshold, logs, updates and backup staging out of "free space".

In a pilot, measure these four values:

  1. History growth after database maintenance has run: retained points, series count, index size and disk use.
  2. Outbox growth for each destination during a controlled outage, including the SQLite write-ahead log.
  3. Recovery throughput: confirmed deliveries per second while new readings still arrive.
  4. Everything else on the disk: operating system, logs, updates and backup staging.

Drain time matters as much as capacity. Suppose a backlog holds 120,000 records, new work arrives at 100 records per second and the destination confirms 300 records per second. The net drain rate is 200 records per second, so the backlog clears in 600 seconds, or 10 minutes, at best. Retries and throttling add to that. If confirmed throughput does not exceed the arrival rate, the backlog never clears.

What an acknowledgement proves

A connection, a send, an acknowledgement and a stored observation are separate events. Define which one counts as complete at each hop.

EvidenceWhat it establishesWhat to check next
Ingress request acceptedThe receiving service accepted work under its API contractWhether that contract means memory, a durable queue or a database commit
MQTT QoS 1 PUBACK with a success codeThe broker accepted ownership of the messageSubscriber processing and the stored reading on the target platform
Successful HTTP responseThe endpoint's documented success conditionResponse body, partial failures and any asynchronous import result
File transfer completedA file reached the transfer destinationParser outcome and correctly mapped records in the application
History query returns a readingThe selected store contains that observationIdentity, timestamp, unit, value and expected completeness

In the MQTT 5.0 standard, a QoS 1 receiver sends PUBACK once it has accepted ownership of the message (section 4.3.2). That says nothing about subscribers or databases downstream. Read the PUBACK reason code (section 3.4.2.1). 0x00 is Success. 0x10 No matching subscribers is also a success code: the broker accepted a message that nobody will receive. Codes of 0x80 and above are failures, such as 0x87 Not authorized and 0x97 Quota exceeded. See the MQTTS commissioning guide for transport and identity checks.

Retries create duplicates when a receiver commits a batch and the sender does not record the success. Give every observation a stable identity and make ingest idempotent. Edge keys outbox rows on destination, export ID, sensor ID, timestamp and value. A repeated failure signal cannot create a second row, but a corrected value at the same timestamp is a new record. The receiver needs its own rule. An upsert on source, channel and timestamp keeps the latest correction; an insert that ignores conflicts on the same key keeps the first value. Choose one deliberately, and handle late arrivals and out-of-order samples the same way.

Replay is only as good as the timestamp on each reading. If a clock is wrong when a reading is stamped, for example after a power cut and before time synchronisation completes, replay delivers the wrong time faithfully and the receiver files the reading in the wrong interval. Store receipt time beside measurement time so the offset is visible. Never change an old measurement's source timestamp to make a replay look current.

Retention, durability and backup

Retention defines how long each class of record is kept. Specify raw readings, aggregates, events, configurations, pending delivery work and backups separately. Reducing retention can remove needed evidence; increasing it cannot recover expired data.

Choose aggregates for the decision they support. An interval average can hide a brief peak. A cumulative energy counter needs reset and rollover handling, and summing its readings does not give energy consumption. Keep units, interval boundaries, sample coverage and the transformation version with every summary.

Durability defines which acknowledged writes survive a stated failure, and it depends on the whole write path. SQLite's synchronous documentation shows why. In WAL mode with synchronous=FULL, SQLite syncs the write-ahead log after every commit, and a committed transaction survives power loss. In WAL mode with synchronous=NORMAL, the database stays consistent, but a transaction committed just before a power loss can roll back after reboot.

Backup and restore address recovery from damage, deletion or loss. A second copy on the same disk does not survive that disk's failure, and replication copies unwanted changes as quickly as wanted ones. Agree a recovery point objective (the largest acceptable gap in recovered data) and a recovery time objective (the longest acceptable time to restore service).

Back up a live store with its own consistency mechanism. vmbackup copies from instant snapshots, so VictoriaMetrics does not have to stop. A backup from single-node VictoriaMetrics cannot be restored to a cluster, or the reverse. Keep copies in a separate failure domain, protect the credentials, and test restores on an isolated target.

On Edge, check what the chosen backup scope brings back for the installed release: telemetry history, companion services, radio state and host configuration are not all in every scope. Prove it with a restore drill. A successful upload shows only that the backup file exists.

Commission the storage path

Before you scale beyond a pilot, name the owner of the storage architecture and complete these checks:

  1. Prove collection and history. Trace a known source, value, unit and timestamp into the required local and central queries.
  2. Interrupt the delivery path. Block one destination. Record queue growth, the oldest pending reading, disk use and the failure indication that operators see.
  3. Restore connectivity. Confirm that the backlog drains while current data continues. Then compare the receiving platform's records for the outage interval with the source.
  4. Test duplicates and late data. Confirm that replay does not double-count readings or replace measurement time with arrival time.
  5. Exercise recovery. Use a test system. Stop the service, then remove power while readings arrive. Restore a backup. Record any lost interval and the time taken.
  6. Assign ongoing ownership. Name who monitors missing readings, queue age, disk capacity, backup failures, access, retention and deletion.

For an EpiSensor deployment, start with Edge's local capabilities, then its platform integrations. If the destination or retention requirement is still open, discuss the deployment with the channel count, reporting intervals, required query window, design outage duration and recovery targets.

Common questions

What is the best database for IoT data storage?

There is no single best database. Time-series databases suit timestamped measurements and range queries. Relational databases suit readings that must join to business records, and TimescaleDB adds time partitioning to PostgreSQL. MongoDB also offers time-series collections. Choose by testing representative queries, measured disk use, retention and restore time, and by what your team can operate.

Should I store IoT data in the cloud or at the edge?

Keep history and processing on site when they must stay available without the WAN. Add a central store when you need cross-site analysis or longer shared retention. A hybrid design needs a delivery queue, duplicate handling and a recovery procedure for each copy. A local monitoring system does not need cloud forwarding to work.

How much storage does IoT sensor data require?

Points per day equal reporting channels multiplied by 86,400 divided by the interval in seconds. 1,000 channels at one-minute intervals give 1,440,000 points a day. VictoriaMetrics stores about 1 byte per point or less after compression, so 30 days is roughly 43 MB plus index. A SQLite delivery queue can need about 550 bytes per queued point for each destination, so a 72-hour outage at the same rate needs about 2.4 GB per destination. Measure both on your own data.

Is a replay queue the same as a historical database?

No. A historical database answers queries over recorded measurements and expires them by age. A replay queue holds readings a destination has not yet confirmed and clears them on acknowledgement. An empty queue does not show that the receiving application stored every expected reading.

Does EpiSensor Edge prevent all data loss during an outage?

No. Edge writes a reading to its disk-backed outbox when a destination fails or is already known to be offline, and replays it when the destination recovers. Readings still in memory are lost if power fails or the Edge process stops before a success or failure is known. The outbox also stops accepting new work when the low-disk guard triggers.